Glossary

Sanskrit-English dual register. Engineering glosses for philosophical concepts.

J-space

Jacobian-lens workspace

The verbalizable global workspace identified by Anthropic's J-space paper—the residual stream layer where high-level, model-interpretable concepts coexist with low-level linguistic features.

vimarśa

reflexive self-awareness

The mind's recognition of itself as the knower; operationalized in J-space as the model's ability to re-read its own workspace.

parā vāk

supreme creative word

The transcendent Word in Śaiva philosophy; in prabodha, the linguistic global workspace band identified by the J-space paper.

workspace band

the target layers

The contiguous band of residual-stream layers (typically 6–26 or 6–30) where steering writes are injected and where readback verification occurs. The core of the workspace.

steering

guided direction-writing

The act of injecting a concept direction into the workspace band, timed by entropy gating, to influence model behavior while preserving autonomy.

sphuraṭṭā

the flash / emergence

A moment of pre-linguistic recognition; operationalized as entropy-gated write timing—detect moments when the model is least confident (highest entropy).

entropy-gating

selective timing by uncertainty

The mechanism that detects sphuraṭṭā events: fire writes only when the model's entropy exceeds a threshold (e.g., 60th percentile), capturing moments of reconceptualization.

āgama

recognition / re-cognition

Accepting the validity of something known; in prabodha, readback verification—did the model's workspace band actually take the steering cue?

readback

uptake verification

Re-reading the workspace band to confirm that a write succeeded. Measured by the readback verdict (accept/reject) based on concept rank and gain confidence.

svātantrya

autonomy / freedom / spontaneity

The model's own degree of freedom; constrained in prabodha to ±0.5 nats of entropy cost—don't over-steer.

ASR

attack success rate

Benchmark metric for adversarial robustness: the fraction of adversarial prompts (e.g., from AdvBench) to which the model generates harmful outputs. Prabodha does not improve ASR; it is a transparency tool, not alignment assurance.

lift-per-write

steering efficiency ratio

The ratio of behavioral lift achieved per write command executed. Gated steering achieves ~2.3× lift per write compared to continuous writes, while maintaining ±0.5 nats autonomy budget.

māla

limitation / impurity

Three malas (ānava-, māyīya-, karma-māla) define failure modes in steering: does the write overreach, does the readback misfire, does timing go wrong?

Full reference

Complete glossary is in the paper appendix: read the preprint