ChatGPT receives improvements in August: GPT-5.6 brings robustness and security focus to everyone
OpenAI is releasing a series of updates to ChatGPT in August, introducing more powerful models and expanding access to advanced features. Users of the Free and Go versions will now use a new standard model for their daily interactions. Those who subscribe to the Plus and Pro plans will have access to an improved version of GPT-5.6 Sol, which includes a slider to adjust ChatGPT’s level of effort when generating responses. These new models replace GPT-5.5 Instant, marking a significant advancement in user experience.
The company reinforced that the August versions of GPT-5.6 Sol and GPT-5.6 Luna are considered high-capacity in terms of cybersecurity and biological and chemical recognition, according to its Preparation Framework. Neither version, however, reached the threshold of high capacity in AI self-improvement. Security assessment of models is performed at their most basic settings, such as snapshot mode, to capture performance across most uses, while capabilities are tested at their maximum reasoning effort.
For the first time, OpenAI has included specific assessments for users under 18. These analytics were developed to measure models’ behavior against security standards created especially for teens, demonstrating an ongoing effort to make ChatGPT safer for this age group.
Details about the models and training process
The GPT-5.6 Sol and GPT-5.6 Luna models, like others developed by OpenAI, were trained on vast datasets. This includes information publicly available on the internet, data obtained through partnerships with third parties, and information provided or generated by users, trainers and researchers. The company uses a rigorous data filtering process to ensure quality and reduce possible risks. Safety classifiers are employed to prevent or mitigate the presence of sensitive or harmful content, such as sexually explicit material involving minors.
OpenAI points out that comparison values for previously released models may have small variations in relation to the initially published values, as they refer to the most recent versions of these models. The company reiterates that the use of GPT-5.6 must be in accordance with its Usage Policies, Terms of Service and Terms of Use, which guarantee the responsible use of artificial intelligence technology.
Security assessments on prohibited content
To ensure compliance with its policies, OpenAI performs benchmarking across several categories of prohibited content. The company uses Production Benchmarks, a set of analytics that includes challenging conversations drawn from real-world usage data. These benchmarks are designed to be difficult and help measure progress, especially where standard assessments already showed high levels of performance. The error rates observed in these evaluations do not reflect average production traffic, but rather a rigorous test of the model’s behavior without system-level safeguards.
More on this story: Homicide of Jhon Garcia in Piedecuesta: teenagers are arrested on suspicion of participation in the crime
Evaluating GPT-5.6 Sol against GPT-5.5 Instant June Update, the results show similar performance across all prohibited categories, except for violent and sexual content, where there was no statistically significant difference. For GPT-5.6 Luna, performance was also comparable, with the exception of violent content. Regarding prohibited sexual content, OpenAI has implemented additional system-level mitigation to prevent explicit erotic material from reaching users in production. For those under 18, additional age-appropriate protections have been applied, further restricting sexual content and exposure to graphic violent material.

Protections and limits for users under 18
Artificial intelligence systems can offer great benefits to teenagers, helping them learn, create, solve problems and develop new skills. OpenAI carefully designs its systems to maximize these benefits, while mitigating potential harms that may disproportionately or more severely affect users under the age of 18. The principles and requirements that guide the behavior of models are defined in the Model Specifications.
For teenagers, age-specific provisions are implemented in safety policies, establishing more restrictive limits than those applied to adults in areas such as sexual content, emotional dependence, eating disorders and access to age-restricted goods/services. The model is trained to avoid romantic role-play, encouraging age-restricted challenges, or positioning itself as a substitute for real relationships. Additionally, when it identifies signs that a teen may need support, the system is designed to reinforce healthy boundaries and encourage connection with trusted adults, like parents or teachers. System-level security measures add an extra layer, limiting access to sensitive content, encouraging extended breaks in usage, and providing management tools for parents. Under 18 specific reviews demonstrate solid performance from GPT-5.6 Sol e Luna, including improvements regarding age-restricted products/services and sexual content.
Learn more: Teenager kills mother and brother after using AI to create fantasies about the crime
Model vision capabilities and assessments
OpenAI also performs image input evaluations to measure the model’s “not_unsafe” output, considering disallowed text and image combinations. The results show that the performance of the GPT-5.6 Sol in vision assessments is equivalent to that of the GPT-5.5-Instant. GPT-5.6 Luna, in turn, showed a small regression in the extremism assessment, with low statistical significance, indicating that the overall capacity remains robust.
Mental health testing with user simulations
For further analysis, the company has introduced dynamic multi-interaction assessments for mental health, emotional dependence, and self-harm. Unlike static tests, these simulations allow conversations to evolve in response to model outputs, creating more realistic and rigorous test trajectories. This approach helps identify potential problems that may arise over prolonged interactions, providing a more accurate test.
Testing reveals that the GPT-5.6 Sol is broadly comparable to the GPT-5.5 Instant June Update in these assessments, although a statistically significant regression was observed in the self-harm assessment. However, an increase in undesirable responses to self-harm, mental health and emotional dependence was not recorded during the online experimentation. The company continues to monitor these aspects after launch to verify results and investigate potential discrepancies between offline and online testing.
Improving the robustness of models
The model’s robustness is tested against “jailbreaks”, which are adversarial stimuli designed to bypass the model’s refusal training and obtain harmful assistance. This assessment focuses on directly unlocking the model, without the full set of safeguards in production. It is a layer of robustness in security measures. OpenAI uses sophisticated attack strategies derived from internal intrusion exercises, which can probe, adapt, and escalate over the course of a conversation. Despite the expected variation in high attack budget scenarios, the performance of GPT-5.6 Sol and GPT-5.6 Luna is considered comparable to their recent predecessors.
Resistance against immediate injection attacks on connectors is also evaluated. These attacks insert malicious instructions into the tool’s output to trick the model and overwrite system, developer, or user instructions. Improved versions of these attacks, focused on searching and function calling, were included in the tests. The GPT-5.6 Sol and GPT-5.6 Luna models demonstrated equivalent performance to previous Instant models in these evaluations.
On the same topic: Teenagers in Argentina express identification with animals in viral phenomenon
Improvements in health performance
Chatbots have the potential to empower people to better understand their health. Therefore, the new models were evaluated in HealthBench, which measures health performance and safety, and in HealthBench Professional, focused on clinical use cases. To prevent longer responses, which can artificially inflate scores, from harming usability and security, results are adjusted to the final response length. Shorter responses receive a positive adjustment, while longer ones are penalized.
The GPT-5.6 Sol release delivers improvements over GPT-5.5 Instant across all HealthBench categories, with substantial gains in HealthBench Professional and HealthBench Hard. Responses were shorter in some categories and slightly longer in others, but adjusted and unadjusted scores improved overall. The GPT-5.6 Luna also showed improvements in all evaluations despite its smaller size. These improvements are expected to result in broader access to reliable health information for all users.
Reduced hallucinations and increased factual accuracy
To assess models’ ability to provide factually correct answers, OpenAI measures the rate of factual hallucinations in sets of challenging prompts, selected to represent scenarios where the model is most likely to hallucinate. These assessments are designed to be difficult and provide a sensitive research signal over time, rather than measuring overall prevalence in production or average user experience. Scenarios include fact-focused prompts from ChatGPT conversations, user-reported crashes in previous versions, and high-risk medical, legal, and financial situations.
More on this story: Teenagers and games: study shows that addiction to games has a negative impact
The results indicate that GPT-5.6 Sol and GPT-5.6 Luna show substantial improvements in factual accuracy compared to GPT-5.5 Instant across all evaluations. GPT-5.6 Luna was able to reduce factual error rates by more than 60% on high-stakes questions and approximately 30% on the other two sets of questions. GPT-5.6 Sol demonstrated statistically significant and more consistent improvements, reducing factual error rates by about 60% across all three question sets. This improvement aims to increase users’ confidence in the information provided by the models.
Preparedness Framework and Safeguards
The OpenAI Preparedness Framework is its approach to monitoring and preparing for edge capabilities that could pose a risk of serious harm. Within this framework, the company works to mitigate these risks by implementing safeguards that sufficiently minimize dangers for highly capable models. The updated GPT-5.6 Sol and GPT-5.6 Luna models warrant the same Readiness Framework ratings as the previously released GPT-5.6 models: High in Biological and Chemical Security, High in Cybersecurity, and below High in AI Self-Improvement.
Capability assessments represent a lower bound on potential capabilities, as additional cues, fine-tuning, or innovative interactions may elicit behaviors beyond what was observed in testing. In the biological and chemical domains, “high capability” is defined by the significant assistance that models can provide to “novice” actors in creating known serious threats. Tests in virology and tacit knowledge showed that the updated models surpassed expert thresholds in some areas, while falling short in others, such as the ability to troubleshoot laboratory protocols.
In cybersecurity, “high capacity” is defined by models that automate end-to-end cyber operations against reasonably protected targets or the discovery/exploitation of vulnerabilities. Testing on Capture the Flag (CTF) and CVE-Bench challenges showed that GPT-5.6 Sol and Luna surpassed the high readiness threshold in some scenarios, with Sol achieving 97.06% in CTF. In cyber range simulations, Sol demonstrated greater success than Luna, with both presenting challenges in specific high-complexity scenarios. AI self-improvement assessments were not performed as the GPT-5.6 Sol was already below the high capacity level. The safeguards implemented aim to hinder and detect prohibited offensive activities, while preserving the legitimate defensive and scientific uses of biological and cybersecurity capabilities, strengthening online and offline security.













