Task-manager usage¶
For operators: initializing workspaces, submitting jobs, running managers, and inspecting or repairing what they leave behind.
Every command below is spelled the canonical way, httk workflow …. The
httk-taskmanager executable installed beside it is an alias of the same tree
(init, submit, run, status, and request); see
the project and workflow command line for the mapping.
Initialize a workspace¶
Every httk workflow command names a registered workspace — a name bound once
to a place — never a bare path. workspace init both creates the workspace and
registers the name, and the place is always explicit: --remote local for this
machine (being local is never implied), plus the --path the workspace lives at.
WORKSPACE below is that registered name.
httk workflow workspace init WORKSPACE --remote local --path runs/WORKSPACE
This creates a core-v2 workspace and registers the name for it. Optional
supported extensions are enabled at initialization:
httk workflow workspace init WORKSPACE --remote local --path runs/WORKSPACE \
--extension transactional-data-v1 \
--extension detached-transfer-v1
A workspace on a cluster names the remote it lives on instead of local; it is
created there over the adapter and its name is registered here. workspace list
shows every registered name and where it resolves, workspace forget deregisters
a name, and workspace delete --force destroys the workspace and deregisters it.
A library caller still constructs Workspace(path) directly; the registry is the
command-line contract. See Project and workflow command line for the whole workspace group and
Campaigns for spreading a very large run across many workspaces.
Protocol publications are synchronized to storage by default. --no-durable
turns that off for throwaway workspaces and makes submission and transitions
faster at the price of correctness after a node crash: an unsynchronized
journal frame can be lost while the marker naming it survives, which leaves a
job whose state cannot be read until workspace fsck --repair restores it.
--durable is still accepted and does nothing, since it is now the default.
Workspace policy¶
Four tunables belong to the workspace rather than to any one process, so that
every manager, CLI, and independent implementation attaching it agrees on them.
They live in .httk-workflow/format.json and are read and written with:
httk workflow workspace policy show WORKSPACE
httk workflow workspace policy set WORKSPACE visibility_deadline_seconds 60
httk workflow workspace policy set WORKSPACE retention.journal_days 90
Key |
Default |
Meaning |
|---|---|---|
|
|
How long a marker rename or a referenced journal frame may take to become visible before it is called damage. |
|
|
The claim lease of a manager started without |
|
|
The size at which a journal writer rotates to its next segment. |
|
|
Optional |
Values are given as JSON and validated on write; an unknown key is refused rather than stored. A change reaches a manager when it attaches, so restart long-running managers after changing policy. Concurrent policy writers are not serialized: the write itself is atomic, but the last writer wins.
Application settings¶
Separate from that engine policy, a workspace also holds application settings: a flat, dotted-name map of small values a runner resolves at run time — the VASP command, a pseudopotential library — rather than tunables the scheduler reads.
httk workflow workspace settings set WORKSPACE vasp.command '"srun -n 32 vasp_std"'
httk workflow workspace settings show WORKSPACE
A runner reads one through a.setting("vasp.command"), resolved in layers — the
job’s inputs, then the environment (HTTK_VASP_COMMAND), then the workspace
setting, then the runner’s default — so an operator configures the command once
per workspace instead of exporting it for every job, while a machine’s own
environment still wins. The manager snapshots the settings into each attempt’s
context.json at claim time, so a runner sees the values the workspace held when
its job was claimed. See Packaged VASP runners and Python and Bash authoring parity.
Freeing disk on a quota’d filesystem¶
A manager frees nothing while it runs, by design: it is never required to execute cleanup code, so it can disappear between any two instructions. Over a long campaign the workspace therefore accumulates one control directory per attempt, one journal writer and manager directory per process start, a full copy of every tree a transaction replaced, and an intact bundle for every transfer already acknowledged. On a quota’d HPC filesystem that is what fails first.
Collection is a separate, explicit operation. Configure the retention limits once, then run it from a maintenance job or by hand:
httk workflow workspace policy set WORKSPACE retention.attempt_control_days 14
httk workflow workspace policy set WORKSPACE retention.trash_days 14
httk workflow workspace policy set WORKSPACE retention.journal_days 90
httk workflow workspace gc WORKSPACE --dry-run
httk workflow workspace gc WORKSPACE
It is safe to run against a live workspace: a manager that is still heartbeating keeps its own directory and every journal segment it wrote, no marker or payload is touched beyond the aged attempt-control directories of terminal jobs, and pruning an empty placement mirror that a transition is recreating underneath is an ordinary outcome rather than an error. A limit left unset means keep. See the command guide for the full category table and for what collecting journal history costs.
A long-lived manager can also do this itself, which is convenient where no maintenance job exists:
httk workflow manager run WORKSPACE --gc-interval 3600
The manager then collects at most once per interval, at the end of a tick and
never between observing a marker and acting on it, obeying exactly the same
policy.retention limits. It is off by default, and a failed collection is
logged rather than allowed to disturb scheduling. Keep the interval long: a
collection walks the state tree and the journal directory, which is work the
scheduling passes do not need done often.
Submit a job¶
A prepared payload is a directory containing an immutable job.json and its
runner. Submit it at any arbitrary placement:
httk workflow job submit WORKSPACE PAYLOAD --placement project-a/00/17
Submission copies by default. --move performs a same-filesystem rename and
consumes the source directory.
Run¶
httk workflow manager run WORKSPACE --workers 8
Without pool configuration, a manager advertises the reserved default pool.
Additional routing and capability labels are explicit:
httk workflow manager run WORKSPACE \
--pool vasp \
--capability gpu \
--workers 4
A manager claims work under the workspace’s lease_seconds unless
--lease-seconds overrides it for that manager alone.
--until-idle is useful for batch invocations and tests.
Taking over another manager’s attempt. An expired lease says that a manager stopped heartbeating, which is not the same as its attempt having stopped, so neither workdir mode relaunches on lease expiry alone:
Workdir mode |
What admits a takeover |
Relaxed by |
|---|---|---|
|
The recorded process is provably gone on this host. A second writer would corrupt the shared directory. |
|
|
The recorded process is provably gone, or the heartbeat has been silent for |
|
Both unsafe options and the evidence of every takeover — which rule admitted
it and how old the heartbeat was — are recorded in the new attempt’s state
frame, so job log shows exactly why a job was relaunched.
Long scans. A manager heartbeats between its scheduling passes and inside
long ones, and bounds how many markers of one kind it processes per pass,
resuming the rest on the next pass in a stable order. A workspace too large to
scan inside one lease is therefore served round-robin instead of making the
manager look abandoned to its peers. A pass that still consumes half of the
lease is logged as a warning, and nine tenths of it as an error: raise
lease_seconds, split the workspace, or reduce what the manager scans.
Every claim, launch, transition, recovery decision, and refused request is
logged. The console reports warnings and errors, while the complete info-level
record is rotated into .httk-workflow/managers/MANAGER_ID/log. --log-level
raises or lowers both, --log-file moves the file, and --json-logs emits one
JSON object per line for ingestion.
A manager drains on SIGTERM or SIGINT, which is what a batch system sends
at walltime. The first signal stops claiming, terminates the running attempts,
and keeps committing their outcomes for --drain-timeout seconds before
exiting successfully; a second signal exits immediately. Anything left behind
is recovered from its expired lease by the next manager.
Scheduling¶
A manager never reads the whole workspace on a tick. Every scheduling pass
discovers its work by streaming the state tree of one active kind — one of
submitted, ready, claimed, running, committing, waiting, and
cancelling — and never opens the terminal succeeded, failed, or
cancelled trees at all. The in-memory marker index and every scheduling scan
therefore grow with the active work in flight rather than with the accumulated
history of a workspace that has run for years.
Bounded streaming discovery. A pass walks directory entries with
os.scandir instead of materializing an rglob of the tree, and it stops early
on two independent budgets: it visits at most discovery_budget directory
entries — 4096 by default — and it collects at most maximum_pass_markers
markers — 256 by default — before it yields the tick. It also takes a
heartbeat opportunity every 512 entries inside the walk, so even one enormous
flat placement directory keeps a manager’s lease alive from within the scan
exactly as crossing many placements does, rather than only between passes. The
walk keeps a resume cursor per top-level placement root, held in the manager’s
memory alone — nothing is written to disk, so two managers of one workspace
never contend on a shared position and a restarted manager simply begins a fresh
cycle. The roots are served in a round-robin rotation with per-root resume, so a
one large placement subtree can never starve a smaller sibling, and
the next tick continues precisely where this one stopped. A concurrent
transition that renames or removes a marker underneath the walk is tolerated
silently, consistent with how a vanished marker becomes a miss rather than a
fault.
The exhaustive workspace operations — fsck, gc, harvest, status, and
job list — use the same scandir walker in an exhaustive mode with no cursor
and no budget, so their semantics are unchanged; only the bounded scheduling
passes carry the budgets.
Best-within-window priority. Claiming ready work scans a single bounded window and then claims the best-priority candidates found within that window, in a stable order among equal priorities, up to the number of free worker slots. Priority is therefore best-within-window rather than exact-global: that is the deliberate price of bounded discovery, and the round-robin rotation is what eventually reaches a starved subtree on a later tick. Recovering exact global order would require a derived priority index, which this implementation does not build; it remains a possible future addition only where a deployment measures that it needs one.
Restricting a manager to placement prefixes. A manager may be told to scan only part of the tree, exactly the way pools and capabilities restrict what it claims:
httk workflow manager run WORKSPACE \
--placement-prefix project-a \
--placement-prefix project-b/2026
The flag is repeatable, and every scheduling scan — bounded window and
exhaustive walk alike — is then confined to those subtrees. With no
--placement-prefix a manager scans the whole workspace, which is the default.
Overlapping assignments stay safe because the marker rename still arbitrates a
claim, so two managers assigned the same subtree never both run one job;
disjoint assignments simply divide the scanning, so neither manager pays to walk
the other’s trees. The assignment is deployment policy and not a protocol
change — placement values remain project-owned semantics that the engine only
validates and filters on — and it is recorded in the manager’s manifest, so job why reports a prefix mismatch when a live manager’s placement prefixes exclude
the placement of the job being diagnosed.
Laying out placements across a large campaign and assigning their subtrees to managers by a written recipe rather than by hand is out of scope here; the Phase 14 campaign recipes add it.
Inspect and control¶
httk workflow workspace status WORKSPACE
httk workflow workspace status WORKSPACE --json
httk workflow job request WORKSPACE JOB_UUID pause \
--operator "$USER" --reason "inspection"
httk workflow job request WORKSPACE JOB_UUID continue \
--operator "$USER" --reason "inputs repaired"
Requests capture the exact current marker generation and record reference.
A delayed request therefore cannot mutate a newer job state. One that can never
apply again — because the job has moved on — is moved to
.httk-workflow/requests/retired/ with the reason recorded beside it instead
of being reread on every pass; a request for a runner backend this manager does
not serve is left alone for a manager that does.
When the publishing installation has an operator identity key — created by
httk workflow config init — the request also carries a detached Ed25519
signature over its canonical JSON, and the manager records the verified
operator_key in the journalled state frame beside operator and reason. The
signature is optional in both directions: a request without one is applied
exactly as before, so a mixed deployment needs no flag day, while a request
whose signature does not verify is quarantined with that reason rather than
applied. It is attribution and not authorization; see
the project CLI guide.
Cancelling a running job is fenced and verified, not a single signal:
httk workflow job request WORKSPACE JOB_UUID cancel \
--operator "$USER" --reason "wrong inputs"
The manager first renames the marker running → cancelling, which fences the
attempt so it can no longer commit an outcome. Only then does it SIGTERM the
process group, SIGKILL it if it has not exited within the grace period, and
verify that it is actually gone; only a verified exit moves the job to
cancelled, and how it was verified is recorded in the terminal frame. A
manager that dies mid-cancellation leaves a cancelling marker, and the next
manager finishes exactly the same procedure. A process recorded on another host
cannot be proven stopped here: the job stays cancelling, the reason is
journaled, and a warning is logged on every retry — which is the safe answer,
because cancelled asserts that nothing is still writing the workdir.
Checking and repairing a workspace¶
workspace fsck verifies the one thing a manager cannot route around: that
every state marker still resolves to its journal frame.
httk workflow workspace fsck WORKSPACE
httk workflow workspace fsck WORKSPACE --json
httk workflow workspace fsck WORKSPACE --repair
httk workflow workspace fsck WORKSPACE --repair --quarantine-unrepairable
It reads every marker of every state kind and checks that its record reference
resolves — within the configured visibility deadline, so a merely slow network
filesystem is never mistaken for damage — to a readable frame whose checksum
verifies and whose job, kind, and generation agree with the marker name. Each
problem is reported with a stable code: missing_segment, short_read,
checksum_mismatch, reference_mismatch, identity_mismatch,
unparseable_name, and their siblings. Without --repair nothing is written.
The command exits 0 when the workspace is clean or everything found was
repaired, and 1 when something is left for an operator.
--repair re-points a damaged marker at the last good frame of its job. Since
the frame holding the backward link is the unreadable one, the repair scans the
journal for readable frames naming that job and adopts the newest one older
than the marker’s own generation — never a newer one, which would be either the
damaged frame itself or a transition no marker ever committed. It then writes
one fsck_repair state frame, chained to the recovered frame and carrying its
step, activation, and attempt counters forward, and renames the marker onto it
at the next generation. History is added, never rewritten, and the job is
schedulable again.
Two things are deliberately never repaired:
a
claimed,running, orcommittingmarker whose manager is still heartbeating within its lease is reported and left exactly as it is, because that manager owns the transition that comes next. Stop the manager, or wait for its lease to expire, and run the repair again;a marker with no readable older frame — typically a job damaged before its second transition — cannot be restored at all. It is reported, and moved into
.httk-workflow/quarantine/with an audit record only if--quarantine-unrepairableis also given.
Run it when a node crashed while writing, when a filesystem was restored from a
snapshot, or whenever job show reports that a state frame is not readable.
Inspecting jobs¶
Five commands read one job the way a manager reads it — the authoritative marker,
the journal frame that marker names, and the immutable job.json — and none of
them writes protocol state:
httk workflow job list WORKSPACE --kind ready --placement project-a
httk workflow job show WORKSPACE JOB
httk workflow job log WORKSPACE JOB --limit 20
httk workflow job why WORKSPACE JOB
JOB is a job UUID, a complete tag--uuid job key, or any unique prefix of
either; an ambiguous prefix is refused with the jobs it matched. Every command
also accepts --json and prints one object: a report, a frame array, a
diagnosis, or a job array.
job show reports the state kind, placement, priority, generation, job digest,
runner identity, the budgets of the retry policy against what has been consumed,
the current and initial step, any step set the runner declared, the last
failure, the join and per-child state of a waiting job, and the payload,
workdir, and data paths.
job log walks the journal backward from the marker through
previous_record_ref and prints one line per state frame, oldest first, with the
timestamp, the transition, the step, the attempt ordinal, the reason, and any
failure code. A frame that cannot be read is reported in place; whatever history
remains readable is still shown.
job why answers “why is this job not running?” for every state:
submitted: whether any manager has registered it, and which live managers serve its runner backend;ready: every claim precondition, one line each — runner backend, claim pool, required capabilities, the maintenance lock, the workspace core profile, the attempt budgets, and which live manager would accept the job;claimedandrunning: the owning manager, its heartbeat age against the recorded lease, and whether an expired lease means recovery rather than a stuck job;committing: that a published outcome is being committed and any manager serving the backend resumes it;waiting: the join condition, every child with its label and state, which children block, and which cannot be resolved in this workspace;failed: the failure, whether an operatorcontinuestill fits inside the retry budget, and theerror.jsonbreadcrumb of the last attempt;paused,succeeded, andcancelled: the state and how to proceed.
The job side of every precondition comes from job.json and cannot drift. The
other side — pools, capabilities, and served backends — is deployment policy of
whichever manager is running and is read from the manifest each manager
publishes, so a manager that is not running is reported as absent rather than
assumed.
Reading results rather than status is a harvest: httk workflow harvest WORKSPACE, or httk.workflow.harvest, streams one record per finished job for a
data layer to store; see Harvesting results.
The foreground debug runner¶
httk workflow job debug WORKSPACE PAYLOAD --step relax
httk workflow job debug WORKSPACE JOB --follow-children
job debug drives exactly one job to a terminal state in the foreground and
streams the attempt’s stdout.log and stderr.log to the console as they grow.
Every transition is performed by a private task manager whose scans are
restricted to that one job, so the debugged job runs through exactly the code
paths a production manager uses and no unrelated work is claimed. Lines are
prefixed with the step that produced them, and [debug] marks each transition
the polling loop observed; job log always holds the complete record afterwards.
--log-level raises the private manager’s own console log, which is quiet by
default.
The first argument is either a payload directory, which is submitted fresh at
--placement (debug by default), or a selector of a job that already exists.
--step overrides the initial step of a fresh payload; overriding the step of a
job that already has a history is refused, because rewriting history is what the
recorded override_step request is for. --follow-children drives the children a
waiting job spawned, depth first, and then resumes the parent.
The exit status is 0 when the job succeeded, 3 when it failed, and 4 when
it stopped without finishing — paused, cancelled, or waiting for children
without --follow-children. A live maintenance lock is refused up front, since
it would stop every launch anyway.
Runner contract¶
The runner executes in the selected persistent or isolated workdir. It reads
the context named by HTTK_WORKFLOW_CONTEXT and publishes
outcome.tmp.<nonce>/ as outcome.ready/ beneath
HTTK_WORKFLOW_CONTROL_DIR. See the
Workflow filesystem API for the complete protocol, and
Native runner helpers, Native Bash runner API, or the Python and Bash authoring parity table
for the two authoring SDKs that implement it.
The local executor starts runners behind a one-byte launch gate. It records
the process identity and commits the running marker before releasing that
gate. If the manager disappears during this narrow launch interval, the gated
process observes end-of-file and exits without executing the runner.
httk workflow manager run executes only the normal path runner backend. Legacy
ht_steps jobs use a distinct backend and are intentionally left untouched;
run those with httk v1 task compatibility.