# 11 · Jobs and Scheduling

Read this when an extension needs recurring or scheduled work.

A wasm server extension exports `jobs`: `list-jobs()` declares the jobs it runs
(each id must appear in `surfaces.server.jobs[]`) and `run-job(execution)` runs
one. The host binds each declared id into the scheduler, so any owner or agent
dispatch explicitly owned by that extension runs its executor. A global job-id
match alone is never authority. Canonical execution contract:
[`07-wasm-server-backends.md`](07-wasm-server-backends.md).

## Current model

- `surfaces.server.jobs[]` declares job ownership ids.
- Runtime schedule instances live in Cedros automation settings.
- Declaring a job id does not make extension code executable by itself — the
  wasm component must export it.
- `jobSchedules[]` with `autoCreateOnEnable: true` materializes a schedule into
  the site automation settings on enable, so the host automation runtime
  dispatches the job on its cadence without the owner adding it by hand (e.g.
  outbox dispatch, cleanup). Idempotent; never overwrites a job id the owner
  already configured. Omit it (or set false) to require manual scheduling.
  The server reconciles these schedules on startup and across browser,
  marketplace, and headless/API extension lifecycle operations. A headless
  upload or enable therefore creates the same schedules as the admin UI.
- `autoCreateOnEnable` is an **optional convenience, not a requirement**. Do
  not make your own package validator require it — omitting it is a valid
  choice (the operator schedules the job manually). An extension that
  hard-requires it fails intake for no reason.
- **Provisioning is scoped to automation settings.** The enable flow reads and
  writes only the site automation object via the focused
  `GET`/`PUT /admin/runtime/site/automation` endpoint — it does **not**
  re-serialize the whole runtime site settings. The request body is just the
  job list, so provisioning stays well under the host's request-body cap and
  never clobbers unrelated settings, even on large sites with many schedules.
  (Earlier host versions provisioned by re-saving the entire site settings
  through `PUT /admin/runtime/site`, which could exceed the 2 MiB body cap and
  return `413 Payload Too Large` on large sites — bouncing the enable so the
  settings/sidebar contribution never settled. Keep schedule counts reasonable
  regardless, and prefer one schedule per genuinely-recurring job.)
- Job outcomes that matter to owners should emit site-brain records and
  structured logs ([`16-analytics-and-site-brain.md`](16-analytics-and-site-brain.md));
  use wasm `telemetry.emit-record` when the extension has the required
  capability scope.

## Runtime schedule schema

Manifest `jobSchedules[]` field reference:
[`04-manifest-reference.md`](04-manifest-reference.md). The materialized
automation schedule looks like:

```json
{
  "id": "schedule_123",
  "extensionId": "acme-demo",
  "jobId": "acme-demo:nightly-sync",
  "enabled": true,
  "schedule": {
    "cron": "0 2 * * *",
    "timezone": "America/Los_Angeles",
    "catchUp": false
  },
  "timeoutSeconds": 300,
  "maxRetries": 0,
  "concurrency": {
    "mode": "single-flight",
    "scope": "site"
  },
  "metadata": {
    "extensionId": "acme-demo"
  },
  "channelIds": []
}
```

Runtime schedule rules:

- prefer 5-field cron unless Cedros automation settings explicitly require a
  seconds field; document when 6-field cron is used
- timezone must be an IANA timezone; document DST behavior
- missed-run, manual-run, disable/re-enable, and extension-disable behavior
  must be explicit (`catchUp: true` runs one missed occurrence on next start)
- a `retryable` outcome is re-dispatched with the same dispatch id when
  `maxRetries` is greater than 0; retries use durable exponential backoff
  starting at 30 seconds and capped at 24 hours
- retries are bounded to 10 and increment `attempt` from 1; jobs must still be
  idempotent because a process or network failure can make an external side
  effect ambiguous
- concurrency modes are `single-flight`, `per-site`, `per-account`, or
  `concurrent` unless Cedros publishes others
- metadata must be small, safe JSON; never put secrets in job metadata
- manually configured extension jobs must preserve the host-selected
  `metadata.extensionId`; dispatch fails closed when it is absent or differs
  from the extension that owns the bound job
- notification channels need type, destination reference, and redaction rules

Default extension-disable behavior: pause the schedule and fail closed until
the extension is re-enabled, unless the release handoff documents a different
host-approved policy.

## Job execution contract

The guest receives one `job-execution` per run:

```
job-execution {
    job-id: string,
    metadata-json: string,   // caller-supplied JSON (empty for none)
    attempt: u32,            // 1-based; increments for durable retries
}
```

and returns a `job-outcome` with `status` (`succeeded`, `failed`, or
`retryable`) plus `details-json` for logs/notifications. Execution runs under
the same wasm limits as routes (fresh instance per run, 5 s CPU epoch, 60 s
whole-invocation timeout, 55 s per host call). A job that needs "now" can call
`chrono::Utc::now()` in-guest — the WASI clocks are linked.

A `retryable` outcome becomes `retry_scheduled` while retry capacity remains,
then is re-dispatched after backoff. Once retries are exhausted, it becomes a
terminal `failed` outcome. Host failures with ambiguous ownership or timeout
remain terminal rather than risking an unsafe duplicate side effect. Treat
`skipped`, `timed_out`, and `cancelled` as SDK/platform extension requests
until Cedros publishes them.

## Job design guidance

For each job, document:

- job id and purpose
- schedule expression and timezone
- `timeoutSeconds` (1–86400)
- `maxRetries` (optional, 0–10; defaults to 0)
- safe metadata JSON
- idempotency key strategy
- downstream API limits
- failure notification behavior

The default `database.read` page size is 100, so the canonical sweep-job shape
is: paginate with `limit`/`offset`, make each page's work idempotent, and let a
retry or the next scheduled run pick up anything missed. Long multi-page
sweeps must also fit the 30 s invocation timeout — bound each run's work and
carry a cursor in settings or a control record when a full sweep cannot fit in
one run.

## Rules

- Do not invent a second cron manifest.
- Do not ship a private scheduler that bypasses Cedros automation.
- Cron strings must be valid 5-field or 6-field expressions.
- Keep `timeoutSeconds` between 1 and 86400.
- Keep job metadata safe JSON; never secrets.
- Jobs must be idempotent and self-recovering on the next scheduled run.

## Acceptance checks

- Job ids appear in `surfaces.server.jobs[]` and are exported by the component
  (`list-jobs`); a declared-but-not-exported id fails install atomically.
- Runtime schedules can be represented as Cedros automation settings.
- Jobs are idempotent and self-recovering on the next scheduled run (the host
  does not retry).
- Failures produce owner-visible recovery context (site-brain record and/or
  notification channel).
- If the extension is enabled headlessly (API), the release notes say how the
  schedules get created.

Next: [`12-admin-pages-and-dashboard-cards.md`](12-admin-pages-and-dashboard-cards.md).
