From code snippets to multi-file engineering tasks
The official DeepSWE 1.1 evaluation focuses on long-horizon software engineering tasks in real codebases, where GPT-6.1 Sol achieves performance close to Astra. For developers, it is better suited to providing the problem description, relevant modules, and constraints together, allowing the model to trace call relationships, identify defects, and develop cross-file modification plans.
