> ## Documentation Index
> Fetch the complete documentation index at: https://docs.seekr.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Deploy a fine-tuned model for inference

> Deploy a completed fine-tuning job to a dedicated endpoint, run inference against it, and demote or reactivate the deployment.

After a fine-tuning job completes, deploy the resulting model to serve inference requests. A deployment hosts the model on dedicated compute and gives you a deployment ID to use as the `model` in chat completions. For more deployment operations, see [Deployments](/flow/sdk/deployments).

## Deploy your fine-tuned model

### Find your fine-tuning job ID

**Endpoint:** [`GET v1/flow/fine-tunes`](/flow/reference/list_fine_tune_v1_flow_fine_tunes_get)

List your fine-tuning jobs to find the ID of the completed job you want to deploy.

<CodeGroup>
  ```python Python theme={null}
  from seekrai import SeekrFlow

  client = SeekrFlow()
  print(client.fine_tuning.list())
  ```
</CodeGroup>

Each job in the response includes its ID, training files, training parameters, project ID, and status. Deploy a job only after its status is `completed`.

<CodeGroup>
  ```text Output theme={null}
  FinetuneResponse(id='ft-1234567890', training_files=['file-0987654321'], model='meta-llama/Meta-Llama-3-8B-Instruct', accel_type=<AcceleratorType.MI300X: 'MI300X'>, n_accel=8, n_epochs=1, batch_size=1, learning_rate=1e-05, created_at=datetime.datetime(2025, 4, 25, 14, 29, 41, 817017, tzinfo=TzInfo(UTC)), experiment_name='12345-examplev1', status=<FinetuneJobStatus.STATUS_COMPLETED: 'completed'>, events=None, inference_available=False, project_id=679, completed_at=datetime.datetime(2025, 4, 25, 14, 52, 34, 785756, tzinfo=TzInfo(UTC)), description=None),
  ```
</CodeGroup>

### Check a specific job

**Endpoint:** [`GET v1/flow/fine-tunes/{fine_tune_id}`](/flow/reference/get_fine_tune_v1_flow_fine_tunes__fine_tune_id__get)

Retrieve a single job to confirm its status before you deploy it.

<CodeGroup>
  ```python Python theme={null}
  print(client.fine_tuning.retrieve("ft-1234567890"))
  ```
</CodeGroup>

### Create a deployment

**Endpoint:** [`POST v1/flow/deployments`](/flow/reference/deploy_v1_flow_deployments_post)

Create a deployment that points at your fine-tuning job ID.

<CodeGroup>
  ```python Python theme={null}
  from seekrai.types.deployments import DeploymentType

  deployment = client.deployments.create(
      name="customer-support-deployment",                # 5–100 characters
      description="Fine-tuned model for chat support",   # 5–1000 characters
      model_type=DeploymentType.FINE_TUNED_RUN,
      model_id="ft-1234567890",                          # Your fine-tuning job ID
      n_instances=1,                                     # Dedicated instances, 1–50
  )
  print("Deployment ID:", deployment.id)
  ```
</CodeGroup>

A new deployment starts in `Pending` status and becomes `Active` when it's ready to serve requests. You don't need to promote it.

### Wait for the deployment to become active

**Endpoint:** [`GET v1/flow/deployments/{deployment_id}`](/flow/reference/deployment_by_id_v1_flow_deployments__deployment_id__get)

Check the deployment status until it reaches `Active`.

<CodeGroup>
  ```python Python theme={null}
  import time
  from seekrai.types.deployments import DeploymentStatus

  while True:
      status = client.deployments.retrieve(deployment.id).status
      if status == DeploymentStatus.ACTIVE:
          break
      if status == DeploymentStatus.FAILED:
          raise RuntimeError(f"Deployment {deployment.id} failed")
      time.sleep(15)

  print("Status:", status.value)
  ```
</CodeGroup>

**Sample response:**

<CodeGroup>
  ```text Output theme={null}
  Status: Active
  ```
</CodeGroup>

Use the deployment ID as the `model` when you run inference.

## Run inference

### Stream chat completions

Send chat completion requests to test your model's task or domain performance and get a sense of the end-user experience.

**Endpoint:** [`POST v1/inference/chat/completions`](/flow/reference/route_chat_completion_v1_inference_chat_completions_post)

<CodeGroup>
  ```python Python theme={null}
  stream = client.chat.completions.create(
      model=deployment.id,
      messages=[
          {"role": "system", "content": "You are SeekrBot, a helpful AI assistant trained to answer questions about financial products and services."},
          {"role": "user", "content": "Who are you?"},
      ],
      stream=True,
      max_tokens=1024,
  )
  for chunk in stream:
      if chunk.choices and chunk.choices[0].delta.content:
          print(chunk.choices[0].delta.content, end="", flush=True)
  ```
</CodeGroup>

**Sample response:**

<CodeGroup>
  ```text Output theme={null}
  I am SeekrBot, a knowledgeable guide to the realm of financial products and services.
  ```
</CodeGroup>

<Info>
  For `model`, you can use any supported base model or the ID of an active deployment.
</Info>

### Configure request parameters

Adjust parameters such as `temperature`, `max_tokens`, `top_p`, and `stop` to control the model's output. For every supported parameter and its allowed values, see [Create chat completion](/flow/reference/route_chat_completion_v1_inference_chat_completions_post).

### Return token log probabilities during inference

You can also return the token log probabilities, or "logprobs". Logprobs reveal the model’s certainty for each generated token.

Low-confidence predictions highlight gaps in training data. During staging, you can flag outputs with low confidence (for example, strongly negative logprobs) for manual review or retraining.

Unusually high logprobs for irrelevant tokens can signal hallucinations. During staging, this can help refine prompts or adjust temperature settings.

Logprob requests follow the OpenAI convention:

* To return the logprobs of the generated tokens, set `logprobs=True`.
* To also return the top *n* most likely tokens and their associated logprobs, set `top_logprobs=n`, where *n* > 0.

<CodeGroup>
  ```python Python theme={null}
  stream = client.chat.completions.create(
      model=deployment.id,
      messages=[{"role": "user", "content": "Tell me about New York."}],
      stream=True,
      logprobs=True,
      top_logprobs=5,  # The maximum depends on the deployment; higher values can return a validation error
  )

  for chunk in stream:
      print(chunk.choices[0].delta.content or "", end="", flush=True)
      print(chunk.choices[0].logprobs)
  ```
</CodeGroup>

## Manage your deployment

### Demote a deployment

**Endpoint:** [`PUT v1/flow/deployments/{deployment_id}/demote`](/flow/reference/demote_v1_flow_deployments__deployment_id__demote_put)

When you're done with the model, demote the deployment to stop serving requests. The deployment and its ID remain, and you can reactivate it later.

<CodeGroup>
  ```python Python theme={null}
  deployment = client.deployments.demote("deployment-1234567890")
  print("Status:", deployment.status.value)
  ```
</CodeGroup>

The deployment passes through `Pending` before it reaches `Inactive`.

### Reactivate a demoted deployment

**Endpoint:** [`PUT v1/flow/deployments/{deployment_id}/promote`](/flow/reference/promote_v1_flow_deployments__deployment_id__promote_put)

To serve requests from a demoted deployment again, promote it. You can only promote a deployment in `Inactive` status. Promoting an `Active` deployment returns a 400 error.

<CodeGroup>
  ```python Python theme={null}
  deployment = client.deployments.promote("deployment-1234567890")
  print("Status:", deployment.status.value)
  ```
</CodeGroup>

## Run validation checks

Before you use a deployed model in production, validate its performance to ensure a smooth transition. Make sure your model is ready for production deployment by running comprehensive validation checks.

### Prepare representative validation data

Start by curating diverse validation datasets that mirror real-world inputs, including edge cases and difficult examples your model encounters in production.

**Example:** For a customer service chatbot handling clothing returns, include:

* Simple queries ("How do I return this shirt?")
* Complex scenarios ("I received the wrong size in a different color than ordered")
* Edge cases ("I started a return but the tracking shows it's still at my house")
* Multi-intent queries ("I want to exchange this and add something to my order")

### Run comprehensive checks

Next, evaluate prediction quality and system performance to ensure all production requirements are satisfied.

## Track critical metrics

Statistics are a critical tool for making sure your AI is trustworthy. These commonly used metrics evaluate a model's performance.

### Prediction quality metrics

| Metric | Definition | When to prioritize |
| - | - | - |
| Accuracy | Correct predictions ÷ total predictions | Clear-cut, factual tasks (for example, classification); in contexts where incorrect outputs could lead to significant consequences (medical, legal, financial) |
| Precision | True positives ÷ predicted positives | When false positives are the most significant concern (for example, content moderation) |
| Recall | True positives ÷ actual positives | When false negatives are the most significant concern (for example, compliance monitoring) |
| F1 score | Harmonic mean of precision and recall | When balance between precision and recall is needed |

In practice, there's often a trade-off between minimizing false positives and false negatives. The relative cost of each error type helps determine whether to prioritize precision or recall when optimizing a model. The F1 score is specifically designed to balance the concerns of both, by combining precision and recall into a single metric.

Going back to the clothing returns chatbot, you might prioritize F1 score when the costs of incorrectly rejecting valid returns (customer dissatisfaction) and incorrectly accepting invalid returns (financial loss) are both significant concerns that need to be balanced.

### System performance metrics

* **Latency:** Response time per prediction
* **Throughput:** Prediction volume capacity (for example, 1,000 requests per second)

## Identify areas for improvement and iterate

### Conduct an error analysis

Categorize and investigate patterns in incorrect predictions to identify underlying causes.

### Implement targeted improvements

Apply insights from error analysis to refine the model through iterative improvements:

* [Hyperparameter tuning](/flow/sdk/fine-tuning/create-fine-tuning-job)
* Additional training data
* Model architecture modifications


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.