• Post author:
  • Reading time:8 mins read
You are currently viewing Prompt Engineering for Product Managers

The quality of an AI feature’s output is largely determined by how the model is prompted. Product managers can apply different prompt optimization techniques to get more reliable outputs, reduce variance across user inputs, and build features that behave predictably at scale.

Constructing the Prompt

Every prompt arrives through one of two roles: a system message or a user message. The system message is set by the application. It establishes the model’s tone, its behavior, and what it should and should not do across all interactions. It persists across the conversation and takes priority over the user message when they conflict. The user message carries the input that changes with each interaction.

The goal of a well-constructed prompt is to leave as little to inference as possible as the model will fill whatever gaps you leave.

Assign a Role

Telling the model what it is changes how it responds. A model instructed to act as a customer success manager approaches a support ticket differently than a model with no defined role. Role definition shapes tone, vocabulary, and the assumptions the model makes when information is missing.

Without a roleWith a role
“Summarize this support ticket.”“You are a customer success manager. Summarize this support ticket.”

Be Clear, Direct, and Detailed

The model responds to what you wrote, not what you meant. Be specific about the task, its conditions, and its priorities. Provide contextual information the model needs to act correctly: what the task results will be used for, what audience the output is meant for, and where this task fits in a larger workflow. Provide instructions as sequential steps when the task involves multiple actions.

VagueSpecific
“Write a product update email.”“Write a two-paragraph product update email for enterprise customers announcing a new reporting feature. Focus on time savings. Avoid technical jargon.”
“Analyze this user feedback and suggest improvements.”“First, identify the three most common complaints in this user feedback. Then, for each complaint, suggest one product improvement. Finally, rank the improvements by implementation effort.”

Anthropic’s documentation frames this as a practical test: show your prompt to a colleague with minimal context and ask them to follow it. If they would be confused, the model will be too.

Specify the Output Format

Define what a good response looks like before you ask for it. If you need JSON, say so. If you need a three-sentence summary, say so. The model defaults to formats that reflect patterns in its training data, which may not match what your feature requires.

Use Examples

The simplest prompt gives the model a task with no examples. This is zero-shot prompting and is fine for simple, well-defined outputs;. When the output pattern is complex or the model’s defaults do not match what you need, few-shot prompting closes the gap by including labeled input-output pairs before the actual task. The model infers the pattern from the examples rather than guessing from the description alone.

Use delimiters

When a prompt mixes instructions, context, examples, and user input, the model can conflate them. Delimiters create explicit boundaries that tell the model which part of the prompt is which. XML tags or quotes wrap each component clearly and prevent the model from misreading instructions as content or content as examples.

<instructions>
You are a customer success manager. Summarize support tickets for an engineering team.
</instructions>

<ticket>
{{user_input}}
</ticket>

Define constraints

Defining what the model should not do is as important as defining what it should. A customer support feature with no constraints will answer questions outside its intended domain. A code review feature with no constraints may start evaluating business logic that it was not asked to assess.

Without constraintsWith constraints
“Answer the user’s question about our product.”“Answer the user’s question about our product. Only respond to questions about pricing, features, and onboarding. If the user asks about anything else, say: ‘That is outside my area. Please contact our support team.'”

Managing Variable User Input

User-generated input is the primary source of variance in a production AI feature. A well-constructed system prompt does not guarantee consistent output when the inputs vary in type, complexity, and length.

Intent Classification

When a feature handles multiple types of user requests, a single set of instructions will underperform compared to instructions tailored to each request type. Intent classification addresses this by routing the input before processing it. The model first identifies what kind of request it is receiving, then applies the most relevant instructions for that type. A generic prompt trying to handle billing questions, technical support, and general inquiries simultaneously is more likely to apply the wrong instructions to the wrong request type, producing responses that miss the point of the query or add irrelevant caveats.

Classify this support request into one of the following categories:
Billing, Technical Support, Account Management.

<request>{{user_input}}</request>

Return the category inside <intent> tags.

For a production implementation of this pattern, see Anthropic’s ticket routing documentation.

Grounding with Reference Text

LLMs generate responses from training data. This becomes a reliability problem when the queries are about specific documents, policies, or proprietary content. Providing the relevant content directly in the prompt and instructing the model to answer from that material, rather than from its training, anchors the output in content you control. If the answer is not in the provided material, the model should say so rather than infer. Requiring the model to cite the specific part of the provided content that supports its answer also gives users a basis for verification and reduces the risk of the model drifting from the source.

Answer the following question using only the information in the provided document.
If the answer is not in the document, say "I cannot find this in the provided material."
Cite the specific passage that supports your answer.

<document>
<document_content>{{document_content}}</document_content>
</document>

Question: {{user_question}}

Task Decomposition

Some tasks are too complex for a single prompt to handle reliably. Instead of a prompt that asks the model to research, synthesize, draft, and format in one step, use a sequence of focused prompts where each handles one stage and the output feeds the next. The same principle applies to long documents that exceed what a single prompt can process well: break the document into sections, summarize each independently, then combine those summaries. Each step gets the model’s full attention on a bounded task rather than partial attention on an unbounded one.

Prompt 1:

Summarize the key findings in the following document section.

<document_content>{{section_1_content}}</document_content>

Prompt 2:

Summarize the key findings in the following document section.

<document_content>{{section_2_content}}</document_content>

Prompt 3:

Using the summaries below, identify the three most important themes
and rank them by business impact.

<summary_1>{{prompt_1_output}}</summary_1>
<summary_2>{{prompt_2_output}}</summary_2>

Chain-of-Thought and Self-Correction

On tasks involving multiple steps, conditions, or logical sequences, a model prompted to answer immediately will produce less accurate results than one instructed to reason through the problem first. This is chain-of-thought prompting: instructing the model to work through its reasoning before producing a final answer.

Step 1: Identify the key requirements in this feature request.
Step 2: List any conflicts or dependencies between requirements.
Step 3: Based on your analysis, recommend a priority order with reasoning.

Feature request: {{user_input}}

A related technique is self-correction, where, after the model produces an output, you prompt it to evaluate that output against the task requirements and identify what it missed or got wrong. This is particularly useful for tasks where completeness matters, such as extracting information from a document, where a single pass may overlook relevant content.

Review your previous response against the following criteria:
- Does it address all parts of the original request?
- Are there any requirements you missed?
- Is the reasoning consistent throughout?

Identify any gaps and provide a corrected response.

Both techniques increase tokens and latency. On high-stakes or complex tasks, the accuracy improvement is usually worth that tradeoff. On simple, well-defined tasks, it typically is not.

Testing for Consistency and Iterate

One good result from a prompt is not evidence that the prompt works. The same prompt will produce different results on different inputs, and the inputs your users bring will not match the examples you tested during development.

First, define what a correct output looks like and then test the prompt across a representative range of inputs. Include expected cases and edge cases. Check where the output format breaks. Check where the model ignores constraints. Check what happens when the user input is ambiguous or outside the intended scope. This is the QA stage for a prompt, not an optional pass after the output feels right.

Prompt engineering is empirical. The model’s behavior on inputs you did not anticipate during construction will surface gaps that no amount of upfront reasoning will catch. Expect to revise and build the conditions that make revision meaningful: defined success criteria, a consistent set of test inputs, and a way to compare versions against the same test inputs.