Topic

Prompt Engineering

Prompt design patterns, instruction hierarchy, and defensive prompt construction.

prompt engineeringsystem promptsinstruction hierarchyguardrailstask decomposition
Evergreen Overview

Prompt engineering is not just about better outputs. In practice it shapes reliability, scope, fallback behavior, and how well an AI system resists misuse when instructions, tools, and untrusted content collide.

What matters most
  • Instruction hierarchy and role separation
  • Clear task boundaries, fallback behavior, and refusal handling
  • Prompt structures that support monitoring and repeatable evaluation
Where teams get into trouble
  • Overloading prompts with too many responsibilities
  • Relying on wording instead of system controls
  • Treating prompts as static text instead of part of application design
Who this page is for
  • Teams operating prompt-heavy workflows
  • Builders refining assistant and agent behavior
  • Reviewers trying to connect prompt design to safety and risk
References

Current notes, events, and source material

These items are included because they add useful evidence, framing, implementation detail, or upcoming context for teams working in this area.

OpenAI News October 2, 2026 guide

OpenAI’s GPT-6 guide treats agent performance as a workflow measurement problem

OpenAI’s October 2026 guide recommends evaluating GPT-6-family models on complete tasks, balancing successful outcomes against latency and cost. It describes stable prompt prefixes for caching, explicit tool and authority boundaries, and context compaction that preserves important evidence during long work. Model selection and reasoning effort become variables to test against the application’s own acceptance criteria. The article is vendor guidance rather than an independent model comparison, and its examples do not establish universal performance or cost savings. Its useful contribution is a concrete set of workflow controls to evaluate together.

OpenAI News September 22, 2026 guide

Diagnose prompt-cache misses without widening an agent’s tool access

OpenAI’s prompt-caching guidance explains how to compare requests for changes that invalidate shared prefixes, place explicit breakpoints, and preserve tool definitions while changing which tools are callable. GPT-6 can also receive appended reasoning-effort updates without rewriting the earlier prefix. Workload savings still need measurement.

OpenAI Alignment Research March 25, 2026 tool

Model Spec Evals translates behavioral rules into prompts and tested rubrics

OpenAI’s March 2026 Model Spec Evals release links 596 prompts to specific policy clauses and scenario-level grading rubrics. Researchers reviewed prompts and checked rubric judgments against labeled example responses; the public repository notes that nine prompts are skipped by the public API harness. Coverage is broad but sparse and primarily text-only, using everyday scenarios rather than adversarial agent workflows. Aggregate compliance is not weighted by real-world frequency or importance, and comparisons with older models partly reflect policy changes. The dataset supports testing intended behavior, not certifying general safety.

OpenAI News March 24, 2026 tool

Teen-safety policy prompts: adapt labels and regression-test the classifier

OpenAI’s Teen Safety Policy Pack supplies prompt-based classification policies and matching validation datasets for gpt-oss-safeguard. The six initial areas cover risks including dangerous activities, harmful body ideals and age-restricted goods. Developers map the policy labels into filtering, review or monitoring workflows and can adapt the prompts to their application. These inspectable starting materials do not provide comprehensive protection or establish performance on a particular product’s users and content.

Skill engineering: independent reviews, edit hooks and rule-level evaluations video thumbnail Play video
AI Engineer September 21, 2026 video

Skill engineering: independent reviews, edit hooks and rule-level evaluations

AI Engineer’s notes from Paul Bakaus’s workshop explain separate visual and deterministic reviews, selective instruction loading, edit hooks and rule-by-rule ablation tests. They distinguish blocking pre-tool hooks from post-edit feedback, describe portability failures, and retain human judgment because aesthetic evaluators can reward the wrong behavior.

Wiz AI Security June 6, 2025 tool

Wiz publishes stack-specific instruction files for safer AI-assisted coding

Wiz’s 2025 article and public repository provide baseline security instructions for common language and framework combinations, formatted for several coding assistants. The repository exposes the generation prompt and script and clearly identifies the rules as AI-generated. The proposed method keeps guidance close to a project’s actual stack and coding practices. The release does not independently establish that these files prevent vulnerabilities or measure their effectiveness in a production repository. They are reviewable prompt material, whose usefulness depends on rule quality, context selection and subsequent code verification.

Choose prompts, retrieval and fine-tuning according to the failure you need to fix video thumbnail Play video
AI Engineer October 4, 2026 video

Choose prompts, retrieval and fine-tuning according to the failure you need to fix

Anant Srivastava separates stable behavioral instructions, changing factual knowledge and learned task behavior. His examples show how training on historical support tickets can preserve obsolete product facts even after a prompt update, while fine-tuning on runbooks leaves missing-document retrieval unresolved. He proposes code-aware chunking and permission metadata for retrieval, and fine-tuning only after human judgments converge on a stable task. This is an architectural diagnostic, not a measured comparison proving one storage choice always wins.

Ground coding-agent dependency reviews in current upstream evidence video thumbnail Play video
AI Engineer October 2, 2026 video

Ground coding-agent dependency reviews in current upstream evidence

Jakub Hojsan uses an API migration to show why a plausible diff can be misread when a coding agent relies on stale model knowledge. A version-change rule triggers retrieval of upstream documentation or changelogs, while query-specific excerpts and retrieval traces make the evidence inspectable without loading entire pages. The method separates tool availability from actually invoking verification. The talk does not establish a fixed knowledge-age gap for every model or the vendor’s claimed search-cost advantage.

GEPA: use execution feedback to improve prompts and agent programs video thumbnail Play video
AI Engineer September 26, 2026 video

GEPA: use execution feedback to improve prompts and agent programs

Lakshya Agrawal’s publisher notes describe GEPA’s reflective search: inspect execution traces and errors, propose text changes, score candidates, and retain alternatives that work well on different examples. Optimize Anything extends the editable object from prompts to programs, skills and policies. The reusable method depends on informative feedback and a suitable evaluator. The talk’s large benchmark gains lack complete evaluation protocols in the presentation, so they do not establish a universal advantage over reinforcement learning.