Model Catalog
Approved foundation models available to OneMain teams across OpenAI, AWS Bedrock, and Google Gemini.
GPT Realtime 2
OpenAIReasoning model for realtime voice interactions with stronger instruction following and reliable tool use for voice-agent workflows.
GPT Realtime 1.5
OpenAIDefault realtime voice model for natural two-way audio conversations, voice agents, and customer support.
GPT Realtime Translate
OpenAIStreaming speech-to-speech translation model for live multilingual audio, priced per minute of audio.
GPT Realtime Whisper
OpenAIStreaming speech-to-text model for low-latency realtime transcription, priced per minute of audio.
GPT-4o Transcribe
OpenAISpeech-to-text model powered by GPT-4o with improved word error rate and language recognition over original Whisper.
GPT-4o mini Transcribe
OpenAILighter, lower-cost speech-to-text model for high-volume transcription.
Amazon Nova Sonic
AmazonAmazon's speech-to-speech model enabling natural, real-time voice conversations with low latency and multilingual support.
Gemini 3.1 Pro
GoogleGoogle's most advanced reasoning model for complex problem-solving, deep reasoning, and frontier coding tasks.
Gemini 3.5 Flash
GoogleMost intelligent flash-class model for sustained frontier performance at high volume.
Gemini 3 Flash
GoogleFrontier-class flash model delivering performance rivaling larger models at lower cost.
Gemini 3.1 Flash Live
GoogleHigh-quality, low-latency Live API model for real-time bidirectional voice and video dialogue.
Gemini 2.5 Pro
GoogleAdvanced reasoning and coding model with deep thinking capabilities for complex tasks.
Gemini 2.5 Flash
GoogleBest price-performance model for low-latency, high-volume tasks with thinking support.
Gemini 2.5 Flash-Lite
GoogleFastest and most budget-friendly multimodal model for high-frequency, cost-sensitive tasks.
Gemini 2.5 Flash Native Audio
GoogleNative audio dialog model producing natural conversational speech from multimodal input.
Gemini Embedding 2
GoogleLatest multimodal embedding model generating vectors from text, image, audio, and video inputs.