GetPro

Prompt Engineer

The Prompt Engineer designs and evaluates the instructions given to language models to adapt their responses to the needs of business teams.

By the GetPro teamPublished on Updated on

Definition and scope

A prompt engineer is a specialist who designs, tests and improves the instructions given to language models for a defined use. They translate a business need into instructions the model can understand, then compare its responses against explicit criteria. Their work focuses on prompts and the system instructions that govern the model’s behaviour.

This work combines writing and experimentation. An instruction may specify the objective, provide context or include examples of expected responses. Not every task needs all these components. The specialist selects those that serve the use case, observes errors and adjusts the instructions accordingly. The quality of the work therefore also depends on the testing approach, beyond the wording chosen.

To define the role, specify who sets the business expectations, who validates the responses and who integrates the instructions into the application. The Prompt Engineer needs to work with people who understand the domain and those developing the solution. Their reporting line should reflect these collaborations, without automatically making them responsible for the entire AI application.

Prompt Engineer, AI Engineer and LLMOps Engineer: who does what?

  • The Prompt Engineer focuses on instructions, their evaluation and successive improvements.
  • The AI Engineer designs and integrates AI solutions into applications, with a remit extending beyond prompts alone.
  • The LLMOps Engineer works on production operations, observability and the reliability of systems using language models.

This profile covers work on instructions, excluding AI research and general model operations. This boundary helps define a coherent role: improving a prompt cannot solve every problem with an application’s cost, latency or operation.

Why this hire matters

The key question is which responses the business can consider satisfactory for a given use. An output that seems convincing on first reading may fail when faced with an ambiguous request or an unusual input. Part of the Prompt Engineer’s work is to make these differences visible, so that instructions can be chosen on the basis of comparable results.

The first risk of a poorly defined recruitment brief is confusing writing fluency with evaluation skills. Someone may produce an elegant instruction without explaining what it improves. Conversely, adding more measurements is of little use if no one has defined the qualities expected of the response. The fundamental decision concerns the criteria: which elements must be present, which errors matter and how should debatable cases be assessed?

Fictional example: a team wants to use a model to classify incoming requests. An initial prompt works on straightforward requests but assigns a category to messages that do not contain enough information. The specialist can introduce an appropriate instruction, then compare the results on straightforward and ambiguous cases. The useful outcome is knowing which errors decrease and which persist.

The choice of tests matters as much as the choice of words. Representative situations allow typical use to be examined. Difficult cases reveal limitations that favourable demonstrations would miss. Qualitative assessment can complement numerical measurements, provided the criteria are defined and applied consistently.

Finally, the organisation must be able to recognise problems that go beyond instructions. For some cost or latency issues, choosing a different model may be more appropriate. Expecting the specialist to fix everything through wording risks prolonging testing without addressing the cause. Their contribution includes the ability to explain this limitation to decision-makers.

Salaries 2025-2026

Level and experienceAnnual gross base
Junior0-2 years45–60 k€
Mid-level2-5 years60–90 k€
Senior5-8 years90–130 k€
Lead (rare)8+ years130–200 k€

Paris market ranges, 2025-2026.

Outside the Paris region, expect 10 to 20 % less.

Key missions

  • Clarify the expected outcome with business stakeholders and translate it into evaluation criteria.
  • Design prompts and system instructions suited to the task and the context provided.
  • Select useful examples to guide the model’s responses.
  • Build representative test cases, including ambiguous or unusual inputs.
  • Compare prompt variants against a reference version using the selected criteria.
  • Analyse persistent errors to guide subsequent iterations.
  • Automate tests with Python and APIs when the assigned work requires it.
  • Document versions and their results so that the team can take over the work.

Skills

Technical skills

  • Instruction design: formulate a precise objective and organise useful context, instructions and examples.
  • Knowledge of language models: adapt instructions to their capabilities and recognise the limits of this work.
  • Response evaluation: define task-specific criteria and compare several versions on relevant cases.
  • Error analysis: relate observed results to the changes made and explain what remains uncertain.
  • Programming and APIs: use Python to automate tests when the role requires it.
  • Domain understanding: interpret responses in light of the constraints of the business activity they are intended to support.

Expected qualities

  • Clarity: reformulate a business request as an objective that project stakeholders can understand.
  • Rigour: distinguish an observed result from a hypothesis and retain a reference for comparing tests.
  • Ability to explain: describe the model’s behaviour and limitations to a non-technical stakeholder.
  • Listening: clarify differing expectations before changing instructions or evaluation criteria.

Common stack

Depends on the context, with no mandatory stackExperimentation: Jupyter notebooks and model APIs, including the OpenAI API in the DeepLearning.AI course.Test automation: Python and API calls.Output evaluation: string-matching tests, code-based or model-based scoring depending on the task.Version management: Agent Studio or Google Cloud SDK to retain prompts and their parameters.

Background and training

When assessing training for a Prompt Engineer, look at the skills it enables someone to apply: understanding a model’s limitations, writing instructions and organising tests. The title of a qualification or course does not, on its own, describe the ability to connect these activities to a business need. The expected level of programming must also match the assigned tasks.

Two practical learning routes are documented. DeepLearning.AI’s ChatGPT Prompt Engineering for Developers course provides practice in writing and iterating prompts with an API and Jupyter notebooks. It requires basic Python knowledge. Anthropic also offers interactive tutorials. These resources illustrate learning options, without constituting mandatory training or sufficient evidence of professional autonomy.

Work completed during training becomes more informative if it retains the initial need, the different instructions tested and the results obtained. It then shows what the person has learnt to modify and measure. For a technical role, a readable test program also provides concrete evidence of their ability to use an API and automate comparisons.

Finally, match the experience sought to the responsibilities. Running tests with predefined criteria and designing the entire evaluation require different degrees of autonomy. Someone with strong domain knowledge needs to be able to acquire the technical basics required for testing. Someone from a development background needs to learn to make business expectations explicit and judge responses in context. These additional skills should be determined by the role, without inferring autonomy from someone’s background alone.

Hiring this profile

When to hire

Consider hiring when designing and evaluating instructions becomes recurring work, with identified use cases and stakeholders able to specify the expected results. The need must translate into concrete responsibilities: designing instructions, preparing tests, comparing versions and explaining remaining errors. Simply wanting to use more AI is not enough to define these responsibilities.

At the exploratory stage, start by defining a task and a few useful criteria. If the need concerns a single series of tests, temporary expertise may suffice. It should help you understand the work to be continued and the skills needed within the team, without assuming that a permanent role is already justified.

When several use cases require ongoing adjustments, specify who will make decisions about the criteria and who will validate the responses from a business perspective. Also determine the expected degree of autonomy: applying an existing protocol, developing it or designing evaluations. The recruitment brief becomes more coherent when the specialist has access to the people and information needed to interpret their results.

Finally, examine the nature of the main problem. If the business needs to design and integrate a complete solution, an AI Engineer may meet a broader need. If the priority is production operations, observability or reliability, consider the LLMOps Engineer role. The deciding factor remains the ongoing work required on instructions, compared with the responsibilities already covered by the team. No universal number of tests can settle the question on its own.

Career path

The role may expand to include evaluation design, coordination of business expectations or sharing working methods with other teams. To prepare for this development, specify the additional decisions entrusted to the person: choosing criteria, organising version comparisons or supporting colleagues in analysing errors.

A move towards an AI Engineer role may be considered if the person develops the skills needed to design and integrate AI solutions into applications. Moving towards an LLMOps Engineer role requires skills in production operations, observability and reliability. These possibilities require a real broadening of technical skills and are not automatic promotions.

The person may also deepen their expertise in instructions and evaluations for a specific domain. The choice depends on the responsibilities they want to take on and the organisation’s needs, with no single path towards research or management.

How to assess this profile

GetPro’s common assessment framework

GetPro structures its assessment around a grid with an upper limit of approximately ten criteria. It distinguishes elements of a candidate’s background that can be verified from skills to explore further in an interview, and assigns an assessment method to each criterion. The interview explores key criteria through open questions and concrete examples. Reference checks are used to cross-check responsibilities held, strengths and areas to explore further, placing examples in their context.

Advice on assessing a Prompt Engineer

The following steps are suggestions to adapt to the role’s responsibilities. They propose criteria and exercises specific to the occupation, without describing a particular protocol used by GetPro.

1. Define the criteria and expected autonomy

Create a grid covering instruction wording, test selection, analysis of results and explanation of limitations. Link each criterion to a task the person will need to perform.

Distinguish between two situations. For autonomy within a defined framework, provide the criteria and test cases. For responsibility for evaluation design, ask the candidate to develop them and justify their choices.

Treat criteria precise enough to distinguish between two responses as a positive sign. A general promise of a “better result”, with no way to assess it, warrants further exploration.

2. Examine past work

Ask for a piece of work the candidate can share, including its objective, an initial version of the instructions and subsequent changes. Ask them to clarify their own role and the decisions made with others.

Invite them to explain an observed improvement and an error that remains. Look for the connection between the tests, the results and the changes retained. A presentation limited to the best responses does not allow you to judge the robustness of the approach.

If the role requires programming, also examine how calls to the model and comparisons were automated. Tailor this check to the technical level actually required.

3. Set a practical exercise relevant to the role

Prepare a clearly scoped exercise with data that can be shared and an understandable business outcome. Include typical inputs and a few ambiguous situations. Ask for an initial version, an iteration and an explanation of the differences between them.

Fictional example: give the candidate requests to classify, some of which are incomplete. Invite them to define acceptable responses, then compare two instructions on the same requests.

Observe how they choose examples and handle disagreements over scoring. Ask what they can conclude from the tests and what still needs testing. Value a conclusion limited to the available results.

One useful warning sign is the absence of a reference version, which makes improvement difficult to assess. Another is changing both criteria and instructions at the same time without explaining the consequences.

4. Assess collaboration and communication

Ask the candidate to present their conclusions to a non-technical stakeholder. Assess their ability to explain remaining errors and distinguish a limitation of the instructions from a problem requiring a different intervention.

Ask them to specify what information they would request from the domain expert and the developer. If the role includes supporting others, examine how they would help a colleague resume the tests and understand the choices the candidate made.

If there is no internal technical expertise, pair an assessor knowledgeable about models and testing with a business representative. Entrust the former with examining the technical approach and the latter with assessing the relevance of the criteria.

5. Cross-check responsibilities held

With the candidate’s consent, use references to clarify the responsibilities they actually held. Ask referees about the candidate’s contribution to the criteria, tests and communication of limitations.

Compare their responses with the work presented. Look for concrete evidence of autonomy and cooperation, taking the working context into account. Avoid treating a general opinion as evidence of a specific technical skill.

Frequently asked questions

What should a prompt engineer keep to hand over their work?

Keep a set of materials that allows another team to resume the tests: the prompt version, system instructions, variables, model and parameters used. Add the evaluation criteria and corresponding results. This record makes it possible to recover the conditions of a test and understand why a version was chosen. It is preferable to hand over these materials together rather than an isolated instruction.

How should data used in a recruitment exercise be handled?

Prepare a fictional case or remove confidential information from the case provided. Specify the permitted tools and the data that may be sent to them before the exercise begins. To make the work comparable, give candidates the same rules. This is a precaution when preparing the case and does not replace examining the rules applicable to your organisation.

What should be reassessed when the model changes?

Reuse the criteria and test cases to compare the new model’s responses with those of the reference. Examine errors that appear, persist or disappear, then adjust the instructions if the results justify it. Keep a record of the model and parameters used: an unchanged instruction is not, on its own, enough to conclude that responses are stable.

What can you conclude from this profile’s salary grid?

The grid provides gross annual fixed salary benchmarks in euros for 2025-2026, for the French market centred on Paris. It does not provide total remuneration package ranges. To set a budget, therefore, compare the level being considered with the responsibilities actually assigned and specify any other remuneration components separately. The grid’s experience labels do not replace an assessment of the candidate’s autonomy.

Sources and method

Related job profiles