LLM Guide
SkyStudio Model Guide
One of SkyStudio’s core values is offering users a broad range of AI models rather than limiting them to a single option. Each model has its own strengths, and we encourage you to explore which model best aligns with your needs.
Model Selection
When starting a new conversation in SkyStudio, you can select the AI model you want to use from the menu in the upper-left corner. You can also switch to another model from the same location after the conversation has started. For example, you can begin with GPT-5.4 and switch to GOAT, Claude 4.7 Opus, or Gemini 3.1 Pro after a few messages.
When you switch models, the existing context—including conversation history, documents, text, and websites—is transferred to the new model, allowing you to continue without losing context or data during the transition.
User data is not stored within the models. Each request sends only the context required at that moment. You can set your preferred default model from your account settings.
Skymod
GOAT Thinking
Use Cases
GOAT Thinking retains the speed and cost advantages of the previous version while supporting daily operational tasks, marketing and sales copy, human resources content such as announcements, emails, FAQs, and short reports, as well as workflows that require frequent and repetitive text generation.
As task complexity increases, adaptive thinking is activated automatically. For simpler tasks, the model responds at standard speed.
Task Suitability
GOAT Thinking performs well in scenarios that require complex reasoning, including contract and regulatory analysis, multi-step operational decision-support processes, document summarization and generation involving dense industry terminology, and workflows that handle sensitive data or require compliance with KVKK/GDPR regulations.
It is also effective in daily workflows that serve large numbers of users and prioritize speed and cost optimization, such as campaign messaging, informational emails, product descriptions, social media copy, internal communications, HR content, and basic question-and-answer scenarios.
Data remains within your organization’s boundaries, and your industry knowledge and business processes are not transferred to external systems through the model.
Our Recommendations
GOAT Thinking applies step-by-step reasoning to complex questions without compromising speed on simpler tasks. Select GOAT Thinking from the model selector to use it in your assistants and chats.
OpenAI Models
GPT-5.6
GPT-5.6 is OpenAI’s next-generation AI model family. It consists of three variants—Sol, Terra, and Luna—each optimized for different use cases. All three provide advanced reasoning, strong context management, high instruction adherence, and professional content generation, while differing in speed, cost, and depth of analysis.
Whether the task involves complex software development and technical analysis or daily content generation and enterprise workflows, the GPT-5.6 family offers options for different requirements. This makes it possible to select an appropriate model for scenarios ranging from high-accuracy tasks to high-volume operational processes.
GPT-5.6 Sol
Use Cases
- Advanced analysis and problem-solving: Designed for software development, system architecture, data analysis, academic work, and technical research that require high accuracy.
- Complex technical projects: Provides comprehensive analysis and reliable assessments across large codebases, long documents, and multi-step engineering processes.
Task Suitability
- Deep reasoning and maximum accuracy: Performs detailed analysis on complex problems, maintains long-context consistency, and is the most capable model in the GPT-5.6 family for critical decision-support processes.
GPT-5.6 Terra
Use Cases
- Balanced professional use: Offers a strong balance between speed and accuracy for software development, data analysis, technical documentation, research, reporting, and daily professional tasks.
- Versatile workflows: Produces consistent and reliable results across tasks such as code development, document preparation, project tracking, and enterprise content creation.
Task Suitability
- Balance of performance, speed, and cost: Performs well in both routine and moderately complex tasks. It is the general-purpose GPT-5.6 model for most enterprise use cases.
GPT-5.6 Luna
Use Cases
- Fast content generation and daily tasks: Produces rapid, low-latency results for email writing, text editing, summarization, translation, customer support, and daily office work.
- High-volume operations: Can be used efficiently in chatbots, customer support systems, and repetitive content generation workflows.
Task Suitability
- Speed and cost optimization: Responds quickly to short and moderately complex tasks. It is the fastest and most efficient model in the GPT-5.6 family for high-volume, low-latency workflows.
GPT-5.5
Use Cases
- Complex and multi-step tasks: Strong in software development, online research, data analysis, and document or spreadsheet creation. When given a fragmented, multi-part task, it can plan the work, use tools, review its output, and continue until the task is complete.
- Agentic coding and computer use: Can move between tools and execute multi-step tasks from end to end. It is particularly effective in agentic coding, computer use, and knowledge-work scenarios.
- Context management: Can return to previous conversations, files, and data to produce more consistent and personalized responses.
Task Suitability
- High accuracy and reduced hallucination: Designed to make fewer errors than previous versions in sensitive fields such as law, medicine, and finance.
- Strong performance without sacrificing speed: Maintains response times similar to GPT-5.4 while offering improved capabilities.
- General-purpose use: Delivers stable and reliable results across a broad range of tasks, from daily professional work to advanced research. It performs strongly in planning, tool use, and multi-step real-world workflows.
GPT-5.4
Use Cases
- Tasks requiring high accuracy: Designed to produce reliable and detailed results in scientific analysis, advanced engineering, complex system architecture, financial modeling, and critical decision-support scenarios.
- Ultra-long context and project tracking: Maintains strong context consistency across very long conversations, large documents, and extensive codebases. Information loss is minimal during long-running sessions.
- Advanced problem-solving and research: Focuses on comparing alternative scenarios to identify the most accurate solution in tasks that require multi-step reasoning. It is particularly strong in root-cause analysis.
Task Suitability
- Advanced reasoning profile (Deep Reasoning): Uses additional computation on complex tasks to produce more accurate results. It performs strongly in large-scale architecture decisions, algorithm design, and system analysis.
- Long-context stability: Preserves context accurately across long technical discussions, multi-file projects, and multi-stage planning processes. Topic drift remains low.
- Depth-first analysis approach: Examines problems step by step rather than resolving them superficially. It is effective at identifying multi-layered errors, causal chains, and indirect impacts.
- Accuracy-focused operation: Prioritizes accuracy over speed. It is a strong option for critical systems, academic work, and large-scale projects.
GPT-5.4 Mini
Use Cases
- Tasks requiring a balance of speed and accuracy: Provides fast and stable performance for everyday use, intermediate-level coding, data analysis, and long conversations.
- Context capacity: Supports approximately 256K tokens, making it suitable for medium-to-long conversations and projects.
- Long conversations and project tracking: Maintains context more effectively than standard models, although it does not perform reasoning as deeply as GPT-5.4. It remains consistent in long dialogues.
- Coding and technical work: Strong in backend, frontend, scripting, debugging, and refactoring tasks. It can analyze medium-sized codebases in a single pass.
Task Suitability
- Optimized reasoning profile: Produces fast responses while maintaining sufficient analytical depth. It provides adequate accuracy for most professional tasks.
- Stable context management: Its 256K-token context supports long conversations, documents, and multi-step tasks.
- General use: Handles coding, writing, analysis, planning, and technical support without significant performance degradation across task types.
- Performance-to-cost balance: A strong option for use cases that do not require maximum accuracy but still need stable results.
GPT-5.4 Nano
Use Cases
- Tasks requiring very fast responses: Designed for short questions, simple code, summarization, translation, and operations that require rapid generation.
- Context capacity: Supports approximately 64K tokens for short-to-medium-length conversations.
- Low latency and high speed: Operates faster than other GPT-5.4 models. It is suitable for real-time use and scenarios requiring rapid responses.
- Lightweight tasks: Appropriate for simple scripts, short analyses, note-taking, text editing, and small code snippets.
Task Suitability
- Light reasoning profile: Sufficient for tasks requiring basic logic, but not intended for deep analysis.
- Limited context management: Its 64K-token context supports short projects and conversations. Context loss may occur in very long sessions.
- Speed-focused operation: Prioritizes speed over accuracy. It is suitable for daily use, mobile use, and rapid content generation.
- Low cost and high throughput: Efficient for simple tasks but not recommended for large-scale analysis.
GPT-5.3
Use Cases
- Balance between general-purpose use and advanced reasoning: Provides high accuracy across daily use, professional analysis, software development, and technical research. It is strong at preserving context in long conversations.
- Context capacity: Supports approximately 256K tokens and performs reliably in medium-to-long projects.
- Advanced multimodal work: Produces consistent results when text, images, tables, and technical data are processed together. It can track tasks across long conversations without losing context.
- Coding and technical production: Strong in analysis, refactoring, and debugging across medium-to-large codebases. It can inspect many files in a single pass, although very large repository analysis requires careful context management.
Task Suitability
- Balanced reasoning profile: Responds quickly to standard tasks while applying deeper analysis to complex ones. It aims to reach accurate results without unnecessary computation.
- Stable context management: Reliably interprets earlier messages in long conversations.
- Supports long technical discussions and multi-stage planning with a 256K-token context window.
- Balance of depth and speed: Goes sufficiently deep into multi-layered problems without becoming excessively slow.
It is one of the more stable options for daily professional use.
GPT-5
Use Cases
- Advanced coding and agentic tasks: Supports long chains of tool calls, parallel or sequential tool orchestration, debugging, and feature development in large code repositories.
- Text and image tasks: Provides strong understanding of long documents, code, and visual content with a 400K-token context window and a 128K-token output limit.
- Enterprise workflows: Offers improved usefulness and reliability for writing, research, analysis, healthcare-related tasks, and real-world workflows compared with previous models.
Task Suitability
- Unified system and intelligent routing: The GPT-5 series is positioned as a reasoning-focused model family.
- Accuracy and reliability: Aims to produce reliable results through lower hallucination rates, more transparent responses, and efficient reasoning.
Our Recommendations
GPT-5 is a strong option for tasks requiring advanced reasoning, multimodal capabilities, and high accuracy in complex workflows.
Input: 272,000 • Output: 128,000
GPT-5 supports enterprise workloads through extended context, image understanding, and flexible response schemas.
GPT-5 Mini
Use Cases
- Speed- and cost-sensitive scenarios: Suitable for well-defined tasks, high-volume question answering, classification, summarization, and RAG-based search.
- Text and image tasks: Offers practical support for long-form content with a 400K-token context window and a 128K-token output limit.
Task Suitability
- Belongs to the same reasoning capability family as GPT-5, with parameters that allow a balance between speed and depth. For low-latency production workloads, minimal reasoning is often sufficient.
Our Recommendations
GPT-5 Mini is a faster and more cost-efficient version of GPT-5, suitable for daily professional tasks that still require strong reasoning.
Input: 272,000 • Output: 128,000
GPT-5 Mini balances performance and speed, providing many GPT-5 capabilities at a lower cost for routine content generation and analysis.
GPT-4.1
Features
- Strong intelligence and analytical capability
- Moderate response speed
- Text and image input support
- 1M-token context window
- 32,768-token output limit
Advantages
- Performs well in complex tasks that require multi-step reasoning and analysis.
- Strong capability for reading, understanding, and summarizing long documents.
- Suitable for technical content generation and detailed reporting.
- Provides a balanced combination of capability and reasonable latency.
Limitations
- Not as fast as models such as o4-mini in applications that require ultra-low-latency responses.
Recommended Use Cases
- Technical support assistants
- Scientific and analytical reporting
- Long-form and detailed content generation
- Educational and research-oriented information summarization
o4-mini
Features
- Fast response times
- Relatively low cost
- Text and image input support
- 200K-token context window
Advantages
- Suitable for applications that require fast results.
- Performs well in chatbots, rapid content generation, and customer support systems.
- Effective for producing quick code examples and handling intermediate reasoning tasks.
- Offers an economical option for large-scale use.
Limitations
- Less capable than GPT-4.1 and o3 in highly complex, multi-stage reasoning tasks.
- Limitations may become noticeable in very long texts or analysis-intensive scenarios.
Recommended Use Cases
- Fast customer support systems
- Dynamic question-and-answer applications
- Short- and medium-length content generation
- Rapid prototyping and coding examples
o3
Features
- Strong logical reasoning capability
- Slower response time
- Text and image input support
- 200K-token context window
Advantages
- Performs strongly in tasks that require detailed multi-step reasoning and multi-stage problem-solving.
- Delivers strong results in challenging mathematics, engineering, and science tasks.
- Produces reliable results in technical documentation and deep-analysis scenarios.
Limitations
- Response times are longer than those of other models.
- Pricing is higher and should be considered carefully from a cost perspective.
Recommended Use Cases
- Scientific article writing
- Code analysis and complex algorithm design
- Multi-stage legal analysis
- Advanced research
o3-mini
Use Cases
- Tasks requiring fast responses, short logical operations, and everyday chatbot interactions.
Task Suitability
- Optimized for speed and efficiency, making it suitable for real-time responses.
- Effective for structured and unstructured data classification, simple text generation, and lightweight reasoning tasks.
- Provides a cost-efficient balance for routine AI interactions.
The chat interface supports image analysis, but document uploads are not supported.
Anthropic Models
Claude 5 Sonnet
Use Cases
- Professional content generation and written communication: Produces natural, fluent, and consistent output for emails, reports, technical documents, meeting summaries, presentation copy, and corporate correspondence. It is particularly effective at preserving narrative consistency in long-form content.
- Coding and software development: Performs strongly in software engineering tasks such as code generation, debugging, refactoring, code review, and technical documentation. It can understand existing codebases and provide reliable development and improvement recommendations.
- Document analysis and research: Can extract key information from technical documents, contracts, academic papers, and comprehensive reports, while connecting information from multiple sources to produce structured and accessible summaries.
Task Suitability
- Balance of speed and accuracy: Claude 5 Sonnet provides high accuracy across a broad range of tasks, from daily professional work to complex analysis, while maintaining fast response times. It offers a strong balance between performance and quality for most enterprise workflows.
- Strong context management: Tracks prior information accurately across long conversations, large documents, and multi-stage tasks. It maintains topic consistency effectively in long-running projects.
- High instruction adherence: Follows instructions carefully and produces output in the requested format. It performs reliably in rewriting, translation, summarization, document editing, and content improvement tasks.
Claude Opus 4.8
Use Cases
Complex research, analysis, and critical decision-support processes: Evaluates information from multiple sources in tasks that require multi-step reasoning. It performs well in low-error-tolerance fields such as strategy, finance, law, and technical analysis, with particular strength in legal evaluation.
End-to-end development and autonomous agentic workflows across large codebases: Strong in reviewing large codebases, debugging, and multi-stage development. A key differentiator is reliability: it is more consistent in identifying issues in its own code and carrying tasks through to completion.
Task Suitability
Complex tasks requiring high accuracy: Designed for scenarios that require planning, analysis, validation, and multi-stage problem-solving rather than simple, single-step work.
Efficient long-context use and stronger task tracking: Remains stable across long documents, large datasets, and multi-step workflows. It uses a one-million-token context window efficiently, while Effort Control allows the same model to be adjusted from lower to higher effort depending on task complexity.
Claude 4.7 Opus
Use Cases
Complex research and critical decision-support processes:
Produces deep and consistent results by evaluating information from multiple sources in complex research and decision-support scenarios. It performs strongly in low-error-tolerance tasks such as strategy, finance, law, technical analysis, and detailed reporting. It is suitable for long-context and multi-step workflows.
Advanced development and review across large codebases:
Performs strongly in long-running software tasks, including reviewing large codebases, understanding architecture, debugging, code improvement, and multi-stage development. It is particularly suitable for agentic software workflows that require a task to be completed without being abandoned midway.
Task Suitability
Complex tasks requiring high accuracy:
Claude 4.7 Opus is suited to tasks that require planning, analysis, review, interpretation, and multi-stage problem-solving rather than simple, single-step answers.
More consistent long-context performance and stronger follow-through:
Provides stable results across long documents, large datasets, multi-step research, and complex task flows. Its long-context performance and efficiency in multi-step workflows make it appropriate for sustained analytical work.
Claude 4.6 Opus
Use Cases
- Advanced coding and software engineering: Strong in resolving complex, cross-system issues in large codebases, understanding architecture, and completing long-running engineering tasks as a coherent whole. It handles ambiguity and can reach solutions with limited guidance.
- Autonomous and multi-step tasks: Works consistently across long-horizon, multi-stage workflows involving tool use and decision support. It can complete tasks with fewer iterations and greater reliability.
- Enterprise tasks and complex planning: Performs strongly in enterprise scenarios that combine information extraction, tool use, and multi-step reasoning.
Task Suitability
- Challenging and multi-step tasks: Performs strongly in scenarios requiring planning, analysis, validation, and multi-stage problem-solving.
- Efficient token usage: Can solve similar problems using fewer tokens, creating cost advantages at scale.
- Strong alignment and security: Provides improved prompt-injection resistance and more consistent behavior compared with earlier versions.
Claude 4.6 Sonnet
Use Cases
Broad scanning and code review:
Uses a breadth-first and adversarial approach to identify logic errors, security vulnerabilities, and standards violations across codebases quickly and in parallel.
Long-document analysis:
Can analyze contracts, financial data, and reports exceeding 300 pages with a one-million-token context capacity.
Task Suitability
High performance with a focus on cost and speed:
Provides strong results in daily coding, testing, and data analysis while offering lower costs than Opus 4.6, particularly in output-intensive tasks such as browser automation.
Consistent and reliable task execution:
Performs well in iterative processes, multi-step workflows, and standards-based repetitive tasks such as smoke testing.
Claude 4.5 Opus
Use Cases
- Complex software engineering: Performs strongly in real-world software tasks. It resolves cross-system issues, reasons about trade-offs, and completes complex work with fewer iterations.
- Long-horizon autonomous tasks: Produces reliable results in workflows requiring continuous reasoning and multi-step execution, with fewer unproductive paths.
- Enterprise decision support and planning: Strong in complex enterprise scenarios that combine information extraction, tool use, and multi-step reasoning.
Task Suitability
- Adjustable effort control: Allows the effort level to be configured through the API to balance speed and depth. At medium effort, high accuracy can be achieved with lower token usage.
- Cost-to-performance balance: Can serve as a default option for many tasks where advanced capabilities are required at a lower cost than earlier high-end models.
- Strong alignment: Offers strong prompt-injection resistance and consistent, reliable behavior.
Claude 4.5 Sonnet
Use Cases
- Coding and agentic workflows: Optimized for software development, multi-step tool use, and long-running autonomous tasks. It performs strongly on coding benchmarks such as SWE-bench.
- Frontend and UI development: Effective at producing clean, accurate, and usable interface code.
- Long-running task tracking: Can preserve task continuity across sessions and track progress based on verifiable work.
Task Suitability
- Advanced agentic capabilities: Provides improved tool orchestration, parallel execution, and more efficient context and memory management.
- Balance of speed and cost: Performs well in high-volume production scenarios where cost efficiency matters.
- Instruction adherence: Follows instructions reliably and refactors existing code accurately.
Gemini Models
Gemini 3.1 Pro
Use Cases
Multi-source research and comprehensive analysis:
Can evaluate different data types—including text, images, audio, and video—together to perform comprehensive analysis. This makes it effective not only for document review, but also for workflows that combine screenshots, tables, audio recordings, video content, and web information. Gemini 3.1 Pro is suited to complex tasks requiring advanced reasoning across multiple modalities.
Real-time research and source-grounded response generation:
Through Google Search grounding, it can connect to current web content, making it a strong option for agents that research recent developments, provide sources, or generate high-confidence output. It is particularly suitable for workflows such as “research the web, verify the sources, and then summarize the findings.”
Task Suitability
Suitable for workflows requiring deep analysis and tool use:
Gemini 3.1 Pro is optimized not only to generate strong responses, but also to use tools correctly, execute multi-step tasks carefully, and produce more consistent results. It is therefore suitable for agents, research assistants, analysis bots, and decision-support systems.
Strong for broad world knowledge and accuracy-focused use:
It is more appropriate for multi-source evaluation, comparative analysis, detailed review, and reliable output than for answering a single question as quickly as possible. The model is positioned to produce grounded and consistent responses based on real-world information.
Gemini 3 Flash
Use Cases
Customer interactions and live agents requiring speed:
Its low latency makes it suitable for live chat, customer support bots, real-time guidance agents, and high-traffic applications. Gemini 3 Flash is designed to offer Flash-level speed and cost efficiency while maintaining strong reasoning capabilities.
Fast classification, data processing, and multi-step operations:
Useful for consolidating unstructured data, managing large numbers of function calls, rapidly interpreting visual and text-based content, batch processing, and real-time operational workflows. Common examples include data cleaning, function calling, real-time analysis, and rapid prototyping.
Task Suitability
Suitable where speed and cost balance are critical:
Not every task requires maximum analytical depth. Gemini 3 Flash is appropriate for systems that receive large volumes of requests, require fast responses, and must control operational costs. It is a strong, fast, and efficient option for agentic workflows.
Strong option for real-time multimodal applications:
Because it can process text, images, audio, code, and video quickly, it is suitable for media analysis, live assistants, automatic classification, content review, and rapid generation workflows.
Gemini 2.5 Pro
Use Cases
Designed for tasks requiring advanced reasoning and coding, Gemini 2.5 Pro is one of the more advanced models in the Gemini 2.5 family. It performs strongly in complex software development, multi-step logic problems, analysis, and decision-support scenarios. Its context window of up to one million tokens makes it suitable for large document collections and long conversations.
Task Suitability
It can be used for code review, refactoring, test scenario generation, mathematics, and technical problem-solving tasks that require high accuracy. It provides consistent reasoning across long legal and technical texts as well as comprehensive reports. It is a strong option for projects requiring high-quality output and advanced reasoning. For lighter, high-volume tasks, a hybrid approach with Gemini 2.5 Flash or Flash-Lite may be more appropriate.
Gemini 2.5 Flash
Use Cases
Gemini 2.5 Flash is a balanced model with a one-million-token context window and a strong price-to-performance profile. It is suited to daily tasks that require near-real-time responses and frequent requests, including conversational systems, customer support assistants, summarization, quality control, and extracting insights from data.
Task Suitability
It offers a balanced combination of performance and speed. It performs strongly in chatbots serving large user volumes, intelligent dashboard commentary, log and text analysis, document summarization, and routine coding tasks. When low latency and cost are important, its thinking mode can still provide deeper reasoning when required, making it suitable for high-volume enterprise use.