> ## Documentation Index
> Fetch the complete documentation index at: https://docs.seekr.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Monitor ingestion

> Track ingestion progress, interpret per-file statuses, and resolve errors through the data job detail endpoint.

Ingestion status is surfaced through the data job that triggered it. `GET /v1/flow/data-jobs/{id}` returns a canonical view of your job — including nested ingestion jobs, per-file records, timeline events, and the derived `status` that tells you whether you're ready to start alignment.

## Understand data job status

A data job runs in two phases:

1. **Ingestion** – files are converted and, for a vector database job, chunked and embedded. This begins when you attach files.
2. **Alignment** – training pairs are generated from the ingested content. This begins when you call `/start`, and only after ingestion has completed.

The phase a job is in determines which operations apply. Cancelling targets the alignment job, so a job still in ingestion has nothing to cancel. Timeline events name the phase they belong to, which is why alignment events appear on a data job.

While ingestion is in progress, the data job moves through these states:

| Status | Description |
| - | - |
| `file_processing` | At least one ingestion job is queued or running. |
| `needs_review` | Manual action required — failed ingestion records, missing files, or missing system prompt. |
| `ready_to_start` | All ingestion completed successfully and prerequisites for alignment are met. |

Once you call `/start`, the status mirrors the alignment job (`running`, `completed`, `failed`, etc.).

## Pre-generation validation

Before a data job begins generating data, SeekrFlow validates the job and stops it early when it can't produce useful results. Stopping early avoids spending tokens on a job that wouldn't succeed. A job is stopped when:

* The instructions are empty or contain only whitespace.
* The instructions can't be interpreted as a data-generation task.
* The uploaded documents don't match the instructions.
* No content in the uploaded documents is relevant enough to the instructions.

When a job stops for one of these reasons, its `status_message` names the specific cause and the adjustment to make. Review the message, revise the instructions or documents, and resubmit the job.

SeekrFlow also sends an email when a job is stopped this way. The email includes the job ID, the source file, the reason the job stopped, and the steps to resubmit.

## Check job status

List all data jobs:

**Endpoint:** [`GET /v1/flow/data-jobs`](/flow/reference/get_data_jobs_v1_flow_data_jobs_get)

<CodeGroup>
  ```python Python theme={null}
  from seekrai import SeekrFlow

  client = SeekrFlow()

  jobs = client.data_jobs.list()
  for job in jobs.data:
      print(job.id, job.status, job.created_at)
  ```
</CodeGroup>

Retrieve a specific data job:

**Endpoint:** [`GET /v1/flow/data-jobs/{id}`](/flow/reference/get_data_job_v1_flow_data_jobs__data_job_id__get)

<CodeGroup>
  ```python Python theme={null}
  detail = client.data_jobs.retrieve("dj-1234567890")
  print("Job ID:", detail.id)
  print("Status:", detail.status)
  ```
</CodeGroup>

**Sample response:**

<CodeGroup>
  ```json JSON theme={null}
  {
      "id": "dj-1b75f4d5-5c9e-4d33-b164-a2393bc5ab6d",
      "name": "Customer support refresh",
      "instructions": "Train a support assistant to answer Q4 product questions.",
      "job_type": "principle_files",
      "status": "ready_to_start",
      "created_at": "2025-11-14T07:35:21.700005Z",
      "updated_at": "2025-11-14T07:37:52.464993Z",
      "ingestion_jobs": [...],
      "files": [...],
      "timeline": [...]
  }
  ```
</CodeGroup>

## Inspect ingestion jobs and file records

The `ingestion_jobs` array contains one entry per ingestion run. Each entry includes a `records` array with independent status and timestamps for every file processed.

<CodeGroup>
  ```python Python theme={null}
  detail = client.data_jobs.retrieve("dj-1234567890")

  for ingestion_job in detail.ingestion_jobs:
      print(f"Ingestion job: {ingestion_job.id} — {ingestion_job.status}")
      for record in ingestion_job.records:
          print(f"  {record.filename}: {record.status}")
          if record.processing_at:
              print(f"    Started: {record.processing_at}")
          if record.completed_at:
              print(f"    Finished: {record.completed_at}")
          if record.status == "failed":
              print(f"    Error: {record.error_message}")
              print(f"    Fix: {record.suggested_fix}")
  ```
</CodeGroup>

### File record fields

| Field | Description |
| - | - |
| `record_id` | Unique identifier for the file record |
| `filename` | Source filename |
| `status` | Per-file processing state |
| `method` | Ingestion method used (`speed-optimized` or `accuracy-optimized`) |
| `queue_position` | Position in queue when `status` is `queued` |
| `progress` | Confirmed progress for this file, from 0 to 1 |
| `progress_at` | When `progress` was last confirmed |
| `projected_progress` | Estimated progress ahead of the last confirmation, from 0 to 1 |
| `projected_progress_at` | When `projected_progress` was calculated |
| `error_message` | Plain-language description of what went wrong |
| `suggested_fix` | Recommended action to resolve the error |
| `created_at` | When the file record was created |
| `processing_at` | When the file entered the `running` state |
| `completed_at` | When the file entered the `completed` state |
| `failed_at` | When the file entered the `failed` state |
| `progress` | Fraction of the file processed, from 0–1 |
| `progress_at` | When `progress` was last measured |
| `projected_progress` | Estimated current progress, from 0–1, for advancing a progress indicator between measurements |
| `projected_progress_at` | When `projected_progress` was calculated |

<Note>
  `progress` is not a completion signal. A finished record does not necessarily reach 1, and all four progress fields are `null` for records processed before progress reporting shipped. Use `status` to determine whether a file is complete.
</Note>

### Track per-file progress

Every file reports its own progress through both document conversion and vector database ingestion, so a job with one slow file is distinguishable from a job that has stalled.

Two values work together:

* `progress` is confirmed. It only moves when a step actually completes, so it never goes backwards, and it reaches 1 only when the file is genuinely finished.
* `projected_progress` estimates where the file has reached between confirmations. It advances during long steps, such as table extraction on a large PDF, when `progress` would otherwise appear frozen.

Confirmed progress catches up to the earlier projection as steps complete. Reading a single file across a run:

<CodeGroup>
  ```json JSON theme={null}
  { "progress": 0.0,   "projected_progress": 0.060 }
  { "progress": 0.060, "projected_progress": 0.268 }
  ```
</CodeGroup>

Use `progress` when you need a value you can trust, and `projected_progress` to drive a progress indicator that keeps moving.

### File list

The `files` array in the data job detail provides a unified view of ingestion outputs and manually uploaded Markdown files:

* Entries with a `record_id` came from ingestion and include per-file processing metadata.
* Markdown uploads have `record_id: null` because they skip ingestion and are immediately alignment-ready.

## Read the timeline

The `timeline` array contains ordered milestone events for the job lifecycle.

<CodeGroup>
  ```json JSON expandable theme={null}
  [
      {
          "timestamp": "2025-11-14T07:35:21.700005Z",
          "event_type": "Created",
          "message": "Data job created.",
          "metadata": {}
      },
      {
          "timestamp": "2025-11-14T07:35:22.483459Z",
          "event_type": "File Processing Started",
          "message": "Started processing files.",
          "metadata": {
              "ingestion_job_id": "ij-a19a1923-1d18-4fc7-8365-96d12ea734ce",
              "status": "running",
              "record_count": 5
          }
      },
      {
          "timestamp": "2025-11-14T07:37:52.435803Z",
          "event_type": "File Processing Completed",
          "message": "Finished processing files.",
          "metadata": {
              "ingestion_job_id": "ij-a19a1923-1d18-4fc7-8365-96d12ea734ce",
              "status": "completed",
              "record_count": 5
          }
      }
  ]
  ```
</CodeGroup>

Events are pre-sorted by timestamp. The timeline reports every milestone a job reached, so a job that ended early still shows the steps it completed.

### Event types

| Event type | Meaning |
| - | - |
| Created | The data job was created. |
| File Processing Started | Ingestion began processing files. |
| File Processing In Progress | Ingestion is processing files. |
| File Processing Completed | Ingestion finished processing files. |
| Alignment Job Created | An alignment job was created for this job. |
| Alignment Job Started | Alignment began. |
| Alignment Job Completed | Alignment finished successfully. |
| Alignment Job Failed | Alignment failed. |
| Alignment Job Stopped | Alignment stopped before completing. |
| Alignment Job Cancelled | Alignment was cancelled. |

Terminal events carry the reason in `metadata.status_message`. Read it to find out why a job failed, stopped, or was cancelled, rather than inferring from status alone.

## Resolve ingestion failures

When a file fails, its record includes `error_message` and `suggested_fix`. The data job remains in `needs_review` until every failed record is resolved — either fixed and retried, or removed.

<CodeGroup>
  ```python Python theme={null}
  detail = client.data_jobs.retrieve("dj-1234567890")

  for ingestion_job in detail.ingestion_jobs:
      for record in ingestion_job.records:
          if record.status == "failed":
              print(f"File: {record.filename}")
              print(f"Error: {record.error_message}")
              print(f"Fix: {record.suggested_fix}")
  ```
</CodeGroup>

To retry, re-upload the corrected file and attach it to the job again via `POST /v1/flow/data-jobs/{id}/add-files`. To skip the file, remove it via `POST /v1/flow/data-jobs/{id}/remove-files`. At least one viable file must remain before alignment can start.

## Troubleshoot common errors

| Error | Suggested fix |
| - | - |
| The file appears to be empty. | Upload a file with content. |
| The PDF may be corrupted, password-protected, or in an unsupported format. | Upload a valid, unprotected PDF. |
| The PDF contains pages that exceed the maximum supported size. | Re-export the PDF with smaller page dimensions. |
| The file was not found or is not owned by the current user. | Re-upload the file or verify the correct `file_id`. |
| Service temporarily unavailable. | Retry the job after a brief wait. |
| Internal processing failure. | If the issue persists, contact support. |

### Document processing issues

| Issue | Possible cause | Solution |
| - | - | - |
| Files fail to upload | File exceeds size limit | Split large files or compress them |
| | Invalid file format | Ensure file extension matches actual format |
| | Network timeout | Implement retry logic with exponential backoff |
| Markdown parsing errors | Improper heading hierarchy | Fix heading structure (ensure proper nesting) |
| | Unsupported Markdown syntax | Use standard Markdown formatting |
| PDF extraction issues | Protected PDF | Remove password protection before uploading |

### File ingestion issues

| Issue | Possible cause | Solution |
| - | - | - |
| Slow ingestion | Complex document structure | Adjust chunking parameters |
| | Resource constraints | Monitor system resources during ingestion |
| Failed ingestion job | Malformed content | Check files for compatibility issues |
| | Service timeout | Increase timeout settings |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.