Multimodal data processing system with hierarchical state representation and deterministic execution-path gating
Patent Information
- Application Number
- PCT/GB2026/050231
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-21
- Filing Date
- 2026-02-18
- Publication Date
- 2026-08-27
Smart Images

Figure GB2026050231_27082026_PF_FP_ABST
Abstract
Description
[0001] 1. Title of the Invention; Multimodal Data Processing System with Hierarchical State Representation and Deterministic Execution-Path Gating
[0002] 2. Field of the Invention
[0003] The invention relates to computer-implemented information processing systems, and in particular to multimodal systems in which input data originating from multiple heterogeneous sources are temporally ordered, transformed into feature representations, processed using a hierarchical internal state structure, and deterministically controlled by a state controller so as to limit activation of computationally intensive processing stages in dependence on predefined technical conditions.
[0004] More specifically, the invention concerns real-time control architectures in which synchronisation, hierarchical state management, controlled retention policies, and state-based gating mechanisms cooperate to maintain predictable processing latency.
[0005] 3. Background of the Invention
[0006] Conventional computer systems designed for processing multisensory or multimodal data typically rely on sequential buffering mechanisms and statistical predictive models executed on general-purpose processors. Although such systems are capable of performing complex pattern recognition and inference tasks, they frequently encounter difficulties in maintaining deterministic processing latency under dynamically changing input conditions.
[0007] In many real-time control environments, including robotics and Internet of Things (loT) systems, large volumes of heterogeneous data streams must be precisely temporally aligned and processed within strict timing constraints. Existing approaches often rely on accumulation of contextual history without clearly defined mechanisms for constraining memory growth or enforcing time-bounded transitions between operational states.
[0008] Furthermore, while machine-learning-based predictive modules are widely employed, they are commonly implemented independently of deterministic control logic. In particular, conventional systems lack integrated mechanisms that:
[0009] • synchronise heterogeneous sensory data within a defined temporal window;
[0010] • apply controlled retention policies in order to constrain memory utilisation; and
[0011] • combine predictive confidence measures with explicit state-transition rules while maintaining predefined latency limits.As a consequence, there exists a technical need for an integrated real-time control architecture capable of synchronised multisensory processing, hierarchical memory management incorporating controlled forgetting, and deterministic state transitions that ensure bounded end-to-end latency.
[0012] 4. Prior Art
[0013] Existing multimodal data processing systems and artificial-intelligence-based control systems generally fall into one of two classes of solutions:
[0014] • Monolithic machine-learning models (e.g. large neural networks) configured to process multimodal data in a continuous manner; and
[0015] • Classical control systems in which sensor signals are processed linearly and decisions are taken on the basis of simple thresholds or heuristic rules.
[0016] 4.1 Limitations of Known Multimodal Systems
[0017] In known multimodal systems (e.g. architectures such as CLIP, VisualBERT, and other selfsupervised models), shared feature-space representations and attention mechanisms are employed. However, such systems exhibit the following limitations:
[0018] • Absence of an explicitly defined hierarchical memory structure comprising distinct short-, mid-, and long-term layers;
[0019] • Absence of a controlled forgetting mechanism incorporating a definable logarithmic decay and threshold-based pruning of stored records over time;
[0020] • The internal model state is not used as a control variable governing system operation, but rather serves solely as an intermediate representation for generating predictions.
[0021] Consequently, these architectures do not integrate hierarchical state management with deterministic control of processing activation.
[0022] 4.2 Limitations of Interaction and Robotics Systems
[0023] Traditional dialogue systems and human-computer interaction systems typically utilise conversation history buffers or sequential text storage mechanisms. However, such systems:
[0024] • Do not organise memory into a hierarchical structure (short- / mid- / l ong-term) with differentiated temporal horizons;
[0025] • Do not apply prioritisation based on contextual relevance and frequency of stimulus occurrence;Do not ensure bounded memory growth, whereby the size of stored history increases linearly, preventing guarantees of constant execution latency.
[0026] In the field of robotics, conventional controllers and finite state machines (FSMs) are typically based on:
[0027] • Thresholds defined directly on raw sensor signals (e.g. voltage, position, velocity); and • Simple logical conditions, without reliance on a structured internal state model derived from hierarchical predictive processing.
[0028] Accordingly, conventional FSM-based control does not incorporate predictive confidence or stability metrics derived from a structured internal representation.
[0029] 4.3 Technical Problem
[0030] Known architectures fail to combine, in a systematic manner, the following three key aspects:
[0031] 1. Deterministic multimodal synchronisation within a defined temporal window (e.g. < 10 ms);
[0032] 2. A hierarchical aperceptive memory combining sequential analysis (e.g. LSTM), selforganising mapping (e.g. SOM), and long-term retention with logarithmic decay; and 3. A finite state machine (FSM) that utilises predictive confidence and stability measures to control activation of computationally intensive processing paths in order to maintain end- to-end latency below a predefined threshold.
[0033] There therefore exists a need for a control architecture in which synchronisation, hierarchical state formation, retention policy, and deterministic state transitions cooperate to ensure predictable temporal behaviour and bounded resource utilisation.
[0034] 5. Solution According to the Invention
[0035] The present invention introduces a hierarchical aperceptive memory module as a central component of a real-time control architecture.
[0036] Unlike conventional history buffers, the aperceptive memory:
[0037] • Comprises explicitly defined functional layers corresponding to short-term, mid-term, and long-term representations;
[0038] • Implements weight dynamics that assign priorities depending on contextual relevance and frequency of stimulation;• Utilises a controlled forgetting mechanism to maintain a bounded and predictable memory size; and
[0039] • Provides a structured internal state representation that functions as a control variable for a finite state machine.
[0040] 5.1 Technical Rationale for “Hierarchical Aperceptive Memory”
[0041] The term Hierarchical Aperceptive Memory denotes the capability of the system to:
[0042] • Construct and update a structured internal state representation across multiple temporal horizons (perception combined with contextual interpretation); and
[0043] • Perform aperceptive control, in which newly received stimuli are integrated as modifications of an existing structured state rather than processed in isolation.
[0044] When combined with deterministic input synchronisation and an FSM governed by predictive confidence thresholds, the disclosed architecture:
[0045] • Stabilises the input to predictive modules; and
[0046] • Constrains memory growth, thereby enabling sustained operation in embedded and edgecomputing environments while maintaining bounded processing latency.
[0047] 6. Summary of the Invention
[0048] The essence of the invention resides in an integrated real-time control architecture in which multisensory data processing is performed by means of recursive internal state synthesis and deterministic gating of execution paths.
[0049] In technical terms, virtual aperception denotes a process in which newly received input data are not processed in isolation, but instead update a current structured internal state representation of the system, the internal state representation being used to control activation of computational and output modules.
[0050] The disclosed architecture is based on three cooperating pillars which collectively ensure temporal coherence, bounded memory utilisation, and predictable latency.
[0051] The overall architecture of the system according to the invention is schematically illustrated in FIG. 1.(a) Deterministic Temporal Unification (Synchronisation)
[0052] The system employs a digital clock module to impose a common temporal reference axis on heterogeneous data streams.
[0053] • Technique: sensory data packets are aligned within a synchronisation window of less than 10 milliseconds.
[0054] • Technical effect: reduction of temporal drift between modalities and provision of a coherent input vector for subsequent memory and predictive processing.
[0055] Synchronisation operates as a prerequisite for hierarchical state formation, ensuring that only temporally consistent data are propagated to further processing stages.
[0056] (b) Hierarchical State Synthesis (HPA - Hierarchical Aperceptive Memory)
[0057] The synchronised input data are transformed into feature representations and processed within a hierarchical memory module comprising short-term, mid-term, and long-term layers.
[0058] • Sequential layer (e.g. LSTM): models temporal dependencies and generates a hidden state representation.
[0059] • Topological layer (e.g. SOM): maps the output vector onto a discrete contextual representation (e.g. a best-matching unit within a neuron grid) and determines fit metrics such as quantisation error.
[0060] • State integration mechanism: parameters derived from weight dynamics and novelty / fit measures are used to determine whether, and to what extent, new data update stored memory representations.
[0061] The long-term memory layer applies a controlled retention policy comprising:
[0062] • time-dependent decay (e.g. logarithmic decay), and
[0063] • frequency-based pruning (e.g. 5% occurrence threshold over 50 processing cycles),
[0064] thereby ensuring bounded memory growth and predictable access cost.
[0065] This hierarchical arrangement provides a structured internal state that evolves incrementally with each incoming input vector, rather than accumulating unbounded historical data.
[0066] (c) Execution-Path Gating (FSM and Predictive Gating)
[0067] The system further comprises a finite state machine (FSM) configured to deterministically manage operational transitions and activation of computational processing based on predictive outputs.• Technique: a predictive module employs cross-modal attention, in which modality weights are determined using similarity measures and a learned weighting function that minimises Shannon entropy of a probability distribution associated with modality activations.
[0068] • Control logic: the FSM monitors predictive metrics, including predictive confidence and a stability or convergence criterion. Until predefined decision thresholds are satisfied (e.g. confidence > 0.7 and stability defined as a change < 0.01 over three consecutive steps), the system remains in a listening state and inhibits activation of computationally intensive downstream modules.
[0069] In this manner, predictive evaluation and deterministic state control are structurally coupled.
[0070] Technical Value and Effect
[0071] In contrast to conventional pipeline architectures operating continuously once triggered, the present solution introduces a feedback structure linking synchronisation, hierarchical memory management, and FSM-based gating.
[0072] This interaction results in:
[0073] • Bounded memory utilisation and predictable processing cost through controlled retention (logarithmic decay combined with frequency-based pruning); and
[0074] • Deterministic end-to-end latency, as synchronisation, retention, and FSM gating cooperate to ensure that the time required to generate predictive and / or control signals does not exceed approximately 50 milliseconds.
[0075] Accordingly, the invention provides a technically integrated real-time control architecture in which temporal alignment, hierarchical state evolution, and deterministic execution gating collectively ensure stable and predictable system behaviour.7. Technical Analysis and Key Advantages of the Solution
[0076] The following section provides an engineering justification for the system modules and their contribution to achieving the technical effect of temporal synchronisation, controlled memory retention, and gated processing under real-time constraints.
[0077] (a) Deterministic Multimodal Synchronisation
[0078] Technique: a digital clock module aligns data within a synchronisation window of less than 10 milliseconds.
[0079] Substance: synchronisation constitutes a prerequisite for subsequent hierarchical processing; samples that do not satisfy the synchronisation criterion are marked as inconsistent and are not forwarded to subsequent processing stages.
[0080] Technical effect: ensures coherence of the input package and reduces predictive errors arising from temporal asynchrony between modalities.
[0081] (b) Hierarchical Aperceptive Memory (HP A)
[0082] Architecture:
[0083] • a short-term layer implemented as a sequential processing network (e.g. LSTM);
[0084] • a mid-term layer implemented as a self-organising mapping structure (e.g. SOM);
[0085] • a long-term layer incorporating a retention policy.
[0086] Substance: integration of sequential representation with topological mapping enables continuous, stepwise (sequential) updating of the contextual state representation upon receipt of each new input vector, rather than accumulation of passive historical data.
[0087] Technical effect: stabilises the contextual representation used by predictive and control modules.
[0088] (c) Weight Dynamics and Controlled Retention
[0089] Technique: logarithmic time-dependent decay based on elapsed time since last access, combined with frequency-based pruning (e.g. threshold of 5% over 50 cycles).
[0090] Substance: the memory incorporates an inherent retention policy whereby infrequent or noise-related representations are weakened and removed, while contextually relevant representations are maintained.Technical effect: constrains memory resource utilisation and stabilises computational cost, thereby supporting maintenance of the required temporal processing regime.
[0091] (d) Cross-Modal Attention with Entropy Regulation
[0092] Technique: determination of modality weights within a key-query-value mechanism, employing a learned function that minimises Shannon entropy of a probability distribution associated with modality activations; application of a softmax function with a temperature parameter in the range of 0.5-1.5; training, for example, using an Adam optimiser.
[0093] Substance: the system adaptively increases the influence of modalities exhibiting higher informational coherence and reduces the influence of noisy modalities.
[0094] Technical effect: improves robustness under uncertain input conditions and stabilises predictive performance in the presence of noise.
[0095] (e) FSM as a Deterministic Transition Regulator
[0096] Mechanism: a finite state machine controlling activation of processing stages based on threshold conditions (e.g. predictive confidence > 0.7 and stability / convergence defined as |A| < 0.01 over three consecutive steps).
[0097] Substance: the FSM performs execution gating, limiting activation of computationally intensive processing paths to situations where predictive outputs satisfy defined confidence and stability criteria.
[0098] Technical effect: ensures predictability of operational transitions and reduces the risk of oscillatory behaviour, thereby structurally linking machine-learning-based prediction with deterministic control logic.
[0099] (1) Physical Technical Effect (Hardware / IO Loop)
[0100] Application: dynamic modification of user interfaces and generation of control signals for external devices in loT or robotic environments.
[0101] Technical effect: the system directly influences physical components through control signals generated on the basis of gated real-time processing, thereby producing a concrete hardwarelevel effect.(g) Two-Stage Contextual Filtering (Optional Embodiment)
[0102] Structure: Filter I (lightweight noise reduction) followed by Filter II (contextual
[0103] sei ecti on / attenti on) .
[0104] Substance: pre-filtering reduces the number of samples propagated to the HPA and predictive modules.
[0105] Technical effect: reduces computational load and supports maintenance of deterministic timing constraints.
[0106] 8. Mechanism of Implementation (Enablement)
[0107] 8.1 Multimodal input records (per modality)
[0108] In one embodiment, the apparatus receives input samples from at least two modalities. Each modality m provides samples as records comprising: (i) a timestamp generated by a common digital clock of the system, and (ii) a modality payload (e.g., text segment, image frame, numeric vector, motion / biometric signal, or system signal). In some embodiments, an optional quality indicator may be included per sample, provided that such indicator is available from the corresponding interface or sensor.
[0109] 8.2 FIFO buffering and temporal synchronisation (< 10 ms)
[0110] For each modality m, the system maintains a first-in-first-out (FIFO) buffer Q[m] ordered by timestamp. The synchronisation unit forms a synchronised multisensory package only when samples from the required modalities fall within a synchronisation window W sync, where W sync < 10 ms. Samples that cannot be aligned within the window are treated as inconsistent and are not forwarded to subsequent processing; depending on the implementation variant, such samples may be discarded or retained for later consideration. This ensures that downstream modules operate on temporally coherent multisensory packages.
[0111] The temporal alignment mechanism and resulting end-to-end latency constraint are illustrated in FIG. 4.
[0112] 8.3 Feature extraction and cross-modal fusion
[0113] For each synchronised package, the system derives modality-specific feature representations in a common dimensional space. A cross-modal attention subroutine is applied to compute modality weights using a key-query-value mechanism in which attention logits are based on cosinesimilarity normalised by vector magnitudes. The weights are obtained via a temperature-controlled softmax, with a temperature coefficient T in the range 0.5 to 1.5. Shannon entropy is computed on the resulting probability distribution over modalities and is used within the learned weighting function objective, such that the weighting function minimises said entropy. A fused representation vector is produced by a weighted averaging of modality-specific features.
[0114] The internal computational structure of the hierarchical aperceptive memory (HP A) is illustrated in FIG.
[0115] 2.
[0116] 8.4 Hierarchical aperceptive memory (HP A): STM — MTM — LTM
[0117] The fused representation is processed by a hierarchical aperceptive memory comprising a shortterm subunit, a mid-term subunit, and a long-term subunit. The short-term subunit is implemented as a sequential processing network configured to model temporal dependencies and generate a hidden state representation. The mid-term subunit is implemented as a self-organising mapping network that computes distances between an output vector derived from the hidden state and neuron weight vectors arranged in a multidimensional grid, selects a winning neuron based on a minimum-distance criterion, and updates neuron weights according to a competitive learning rule. The long-term subunit applies a controlled forgetting mechanism comprising timedependent decay and threshold-based pruning so as to constrain memory utilisation over time.
[0118] 8.5 Controlled retention: logarithmic decay and frequency-based pruning
[0119] The long-term subunit maintains records associated with access timing and frequency of occurrence over processing cycles. In one embodiment, record retention is updated using a logarithmic time-dependent decay function with a threshold parameter and a scaling factor. In one embodiment, records are selectively removed when an occurrence frequency falls below a predefined percentage over a predefined number of consecutive processing cycles. This retention policy limits long-term memory growth and stabilises processing cost.
[0120] The time-dependent decay and frequency-based pruning mechanism applied within the long-term memory layer is illustrated in FIG. 5.
[0121] 8.6 Predictive outputs and convergence criteria
[0122] A predictive module generates predictive outputs from the fused representation and the hierarchical state (contextual state). The system computes at least: (i) a predictive confidence value used for threshold comparison, and (ii) a convergence / stability measure derived from temporal changes in predictive outputs across consecutive time steps. Convergence is detectedwhen the change in predictive output over three consecutive time steps falls below a predefined value.
[0123] 8.7 Deterministic gating by a finite state machine (FSM)
[0124] The state transition logic of the finite state machine (FSM) is illustrated in FIG. 3.
[0125] A decision-making unit implemented as a finite state machine comprises at least a listening state, a processing state, and an output state. The finite state machine applies explicit threshold-based transition rules to the predictive module outputs. In one embodiment, the transition from listening to processing is triggered when predictive output exceeds a predefined threshold, and the transition from processing to output is triggered when convergence is detected based on the three-step stability criterion. When threshold or convergence conditions are not satisfied, the finite state machine maintains or reverts to a non-output state such that costly downstream actions are gated.
[0126] 8.8 Output interface and end-to-end latency (< 50 ms)
[0127] An output interface module comprises control circuitry configured to generate control signals in response to finite state machine state transitions. The control signals dynamically modify at least one of: (i) a display layout, (ii) interface parameters, or (iii) external device behaviour in an loT or robotic environment. In one embodiment, the synchronisation window constraint (< 10 ms), the controlled retention policy, and the finite state machine gating cooperate to maintain an end-to-end processing latency for generating predictive outputs of less than or equal to 50 milliseconds, as a consequence of the structural constraints of the processing pipeline rather than a purely declarative target.
[0128] 8.9 Algorithmic Implementation of an Example Embodiment
[0129] 8.9.1 Multisensory Synchronisation (FIFO + Digital Clock, W sync < 10 ms)
[0130] Inputs: incoming sample (m, payload, t) where t is generated by a common digital clock CLK State: FIFO buffers Q[m] per modality m (ordered by timestamp)
[0131] Parameters: W_sync (synchronisation window), with W_sync < 10 ms; required modality set M_req <= M, |M_req| > 2
[0132] Output: P sync (synchronised package) or ± (no package)Algorithm: SyncWindowPackage
[0133] Procedure SyncWindowPackage(m, payload, t = CLK.now()):
[0134] Q[m].push_back( (t, payload) )
[0135] / / form a package only when all required modalities have at least one buffered sample if exists r in M req such that Q[r] is empty:
[0136] return ±
[0137] / / choose a reference time from the heads of required queues
[0138] t ref := min_{r in M req} Q [r], front}). t
[0139] / / optional anti-stall: remove stale samples that can no longer match the window for each r in M req:
[0140] while Q[r] not empty and Q [r], front}), t < (t ref - W sync):
[0141] Q[r].pop_front()
[0142] / / re-check after stale removal
[0143] if exists r in M req such that Q[r] is empty:
[0144] return ±
[0145] / / verify that heads fall within the synchronisation window
[0146] P sync := empty map
[0147] for each r in M req:
[0148] if abs(Q[r], front}). t - t ref) < W sync:
[0149] P syncfr] := Q[r], front}). payload
[0150] else:
[0151] / / inconsistent package: do not forward
[0152] return ±
[0153] / / consume one sample per required modality
[0154] for each r in M req:
[0155] Q[r].pop_front()
[0156] return P syncImplementation note (consistent with embodiments): samples outside the window may be discarded or buffered for later; the above realises a mixed practical variant: “too old” samples are discarded to avoid FIFO blocking, while “too new” samples remain buffered and the package is not forwarded (±) until alignment is possible.
[0157] 8.9.2 Cross-Modal Feature Extraction and Fusion (K-Q-V, cosine similarity, softmax(T), Shannon entropy)
[0158] Input: P sync (payload per modality)
[0159] State: modality feature extractors Embed(m, •) and projection functions Qproj, Kproj, Vproj Parameters: temperature T G [0.5, 1.5]
[0160] Output: x fused (fused feature vector)
[0161] Algorithm: CrossModalFusion
[0162] Procedure CrossModalFusion(P_sync):
[0163] / / I) embed each modality into a common D-dimensional space
[0164] for each modality m in keys(P sync):
[0165] e[m] := Embed(m, P syncfm]) / / e[m] G RAD
[0166] / / 2) compute query / key / value projections
[0167] for each modality m:
[0168] q[m] := Qproj (e[m])
[0169] k[m] := Kproj(e[m])
[0170] v[m] := Vproj(e[m])
[0171] / / 3) attention logits via cosine similarity (normalised)
[0172] for each modality m:
[0173] z[m] := CosineSimilarity(q[m], k[m])
[0174] / / 4) modality weights via temperature-controlled softmax
[0175] p := Softmax(z / T) / / p is a probability distribution over modalities
[0176] / / 5) Shannon entropy on modality-weight distribution (uncertainty indicator / learning objective)
[0177] H := - E m p[m] * log(p[m])
[0178] / / 6) fused representation as weighted average of values
[0179] x fused := E m p[m] * v[m]
[0180] return x fused, HImplementation note: “learned weighting function minimising Shannon entropy” is reflected here by explicitly computing p and H; training / calibration of the weighting function (e.g., via Adam as per dependent claims) may be described elsewhere, but enablement requires showing how the distribution and entropy are computed and used.
[0181] 8.9.3 Hierarchical Aperceptive Memory Update (STM = LSTM, MTM = SOM / BMU, LTM = log-decay + pruning)
[0182] Input: x fused
[0183] State: STM recurrent state STM state; SOM neuron weights W sornfi]; LTM records LTM (each record stores at least: weight / value, last_access, freq_counter)
[0184] Parameters: SOM learning rate a and neighbourhood function Neighbourhood(-); retention parameters k, T_threshold, s; pruning window N and threshold F_min (e.g., 50 cycles and 5%) Output: hierarchical state S tAlgorithm: HPA Update
[0185] Procedure HPA Update(x fused):
[0186] / / STM: sequential processing
[0187] (h_t, STM_state) := LSTM_forward(x_fused, STM_state)
[0188] / / map hidden state to matching vector for SOM
[0189] y_t := FC(h_t)
[0190] / / MTM: SOM best matching unit (BMU)
[0191] for each neuron i in SOM grid:
[0192] d[i] := Distance(y_t, W_som[i]) / / e.g., Euclidean norm
[0193] BMU := argmin i d[i]
[0194] / / competitive learning update (winner and optionally neighbours)
[0195] for each neuron i in SOM grid:
[0196] infl := Neighbourhood^, BMU) / / e.g., Gaussian neighbourhood W_som[i] := W_som[i] + a * infl * (y_t - W_som[i])
[0197] / / LTM: controlled retention (logarithmic time-dependent decay)
[0198] now t := CLK.now()
[0199] for each record r in LTM:
[0200] t elapsed := now t - r.last access
[0201] / / illustrative update consistent with dependent claim structure
[0202] r.weight := r.weight + k * log(l + (t_elapsed - T_threshold) / s) * (r.value - r. weight) / / update frequency counters (implementation-dependent keying: e.g., BMU-based) UpdateFrequency(LTM, key = BMU)
[0203] / / pruning: remove records below frequency threshold over lastN cycles
[0204] for each record r in LTM:
[0205] if r.freq_over_last_N < F_min:
[0206] remove r from LTM
[0207] / / assemble hierarchical state for downstream predict! on / control
[0208] S t := { STM hidden = h_t, MTM BMU = BMU, LTM state = LTM }
[0209] return S t8.9.4 Predictive Output, Deterministic FSM Gating, and Control Signal Emission (< 50 ms) Input: S t (and optionally x fused)
[0210] State: FSM state G {LISTENING, PROCESSING, OUTPUT}; sliding history pred hist (last 3 predictions)
[0211] Parameters: confidence threshold T conf (e.g., 0.7); stability threshold T stab (e.g., 0.01); stability window length = 3; latency budget L_budget < 50 ms
[0212] Output: control signals emitted by output control circuitry (UI and / or loT / robotics)
[0213] Algorithm: Predict AndGate
[0214] Procedure PredictAndGate(S_t, x fused):
[0215] t start := CLK.now()
[0216] / / predictive module output and confidence
[0217] (pred t, conf t) := Predict(S_t, x fused)
[0218] / / update 3 -step history for convergence detection
[0219] pred_hist.push(pred_t) / / keep last 3 values only
[0220] stable := false
[0221] if pred hist.size == 3:
[0222] Al := abs(pred_hist[2] - pred_hist[l])
[0223] A2 := abs(pred_hist[l] - pred_hist[O])
[0224] stable := (Al < T stab) and (A2 < T stab)
[0225] / / deterministic FSM transitions
[0226] if FSM state == LISTENING:
[0227] if conf t > T conf:
[0228] FSM state := PROCESSING
[0229] if FSM state == PROCESSING:
[0230] if stable == true:
[0231] FSM state := OUTPUT
[0232] / / gating: generate control signals only when OUTPUT is reached
[0233] if FSM state == OUTPUT:
[0234] Ctrl := ControlCircuitry(pred_t, S t)
[0235] Emit(ctrl) / / UI modification and / or loT / robotic control
[0236] / / timing verification as an embodiment constraint (structural latency budget)
[0237] t end := CLK.now()
[0238] / / embodiment: system is configured such that (t end - t start) < L budget
[0239] returnImplementation note: the latency limit is expressed as a structural consequence of synchronisation constraints, retention / pruning, and FSM gating; runtime measurement may be used for verification / monitoring without requiring a particular programming construct.
[0240] 9. Technical Contribution and Distinction Over Conventional Systems
[0241] The invention defines a technical configuration in which:
[0242] 1. Multimodal synchronisation is implemented by timestamping input samples using a common digital clock, buffering modality-specific samples in FIFO queues, and selecting only those samples that fall within a defined synchronisation window (e.g., < 10 ms) prior to further processing.
[0243] 2. Internal state is explicitly structured hierarchically as temporally distinct short-term, mid-term and long-term subunits (STM / MTM / LTM), wherein:
[0244] o the STM performs sequential state update,
[0245] o the MTM performs topological / competitive mapping (e.g., SOM / BMU selection and weight adaptation), and
[0246] o the LTM enforces a controlled retention policy that constrains memory growth over time by applying time-dependent decay (e.g., logarithmic decay) and threshold-based pruning (e.g., frequency-based removal over a defined processing window).
[0247] 1. Execution of the computational pipeline is deterministically gated by a finite state machine (FSM) based on measurable threshold conditions derived from predictive outputs, including at least a predictive confidence threshold and a convergence / stability criterion evaluated over consecutive time steps.
[0248] In contrast to systems that perform continuous processing once initiated (e.g., monolithic ML inference pipelines or always-on multimodal fusion), the disclosed architecture restricts activation of downstream processing and output stages to cases where the defined confidence and stability conditions are satisfied, and constrains the cost of maintaining internal state through controlled retention. The resulting technical effect arises from the coordinated operation of (i) time-window synchronisation, (ii) bounded-retention hierarchical state update, and (iii) deterministic FSM gating, thereby enabling predictable real-time operation (including bounded end-to-end latency in example embodiments) when generating control outputs for interface modification and / or external device actuation in loT or robotic environments.10. Summary of the Invention (UK technical English)
[0249] The present invention addresses a technical problem associated with deterministic multimodal data processing in real-time control systems, namely maintaining predictable end-to-end latency and bounded memory utilisation while generating reliable predictive outputs for downstream control actions. In example implementations, the invention supports operation within an Intelligent Virtual Being (IVB) system architecture by constraining timing and resource consumption under dynamically varying multisensory inputs.
[0250] Key features
[0251] • Precise temporal synchronisation:
[0252] Input signals from a plurality of modalities (e.g., visual, auditory, textual, motion-based, biometric, contextual and / or meta-modal sources) are timestamped using a common digital clock and aligned within a synchronisation window of less than 10 milliseconds. This reduces temporal drift and provides time-consistent input packages for subsequent processing stages.
[0253] • Hierarchical aperceptive memory (HP A):
[0254] A hierarchical memory subsystem maintains temporally distinct layers, including:
[0255] (i) a short-term subunit (e.g., LSTM) for modelling sequential dependencies and updating a hidden state;
[0256] (ii) a mid-term subunit (e.g., SOM) for competitive / topological mapping of feature vectors; and
[0257] (iii) a long-term subunit implementing a controlled retention policy, including timedependent (e.g., logarithmic) decay and threshold-based pruning, to prevent unbounded growth of stored representations.
[0258] • Predictive module with cross-modal attention:
[0259] A predictive module performs cross-modal attention to combine modality-specific features. In one embodiment, modality weights are derived via a learned weighting function that minimises Shannon entropy of a computed distribution over modality weights / activations, thereby reducing the influence of noisy modalities and prioritising more consistent channels in a given context.
[0260] • Deterministic state-based gating (FSM):
[0261] A finite state machine (FSM) deterministically governs transitions between operational states (e.g., listening, processing, output) based on explicit threshold conditions, including at least a predictive confidence threshold and a stability / convergencethreshold derived from temporal changes in predictive outputs. This gating restricts activation of higher-cost processing paths unless required conditions are satisfied.
[0262] Main technical advantages
[0263] 1. Predictable latency:
[0264] By coordinated operation of synchronisation, hierarchical state update, controlled retention, and FSM-based gating, the system maintains an end-to-end latency budget for producing predictive outputs (e.g., < 50 ms in example embodiments), supporting realtime human-machine interaction, robotics and loT control.
[0265] 2. Bounded memory utilisation and stable processing cost:
[0266] Controlled forgetting and pruning remove low-utility or low-frequency records and apply time-dependent decay, maintaining a bounded memory footprint and stabilising processing overhead over extended runtimes.
[0267] 3. Real-time control applicability:
[0268] The apparatus can generate control signals for dynamic interface modification and / or for actuation of external devices in loT or robotic environments in response to changing predictive state, while preserving deterministic timing and resource constraints.
[0269] 11. Industrial Applicability and Commercialisation Potential (UK technical English) Owing to the combined HPA-FSM architecture — including sub-10 ms multisensory synchronisation, hierarchical state representation with controlled retention, and deterministic FSM gating — the system is applicable in industrial domains where excessive latency, erroneous actuation, or constrained power budgets can lead to operational failure.
[0270] 11.1 Autonomous Robotics and Industry 4.0
[0271] In factory automation and mobile robotics, control decisions must be issued under strict real-time constraints while sensor streams may be transiently inconsistent.
[0272] Technical value: The <10 ms synchronisation window and threshold-based FSM gating reduce spurious actuation caused by contradictory inputs (e.g., transient camera glare versus stable ranging signals). The controller can withhold escalation to output / actuation until predictive confidence and stability criteria are satisfied, thereby reducing oscillatory control behaviour and improving operational robustness.
[0273] 11.2 Safety-Critical Systems and Medical Monitoring
[0274] Patient monitoring systems typically integrate heterogeneous signals such as ECG, blood pressure, and imaging streams.Technical value: Deterministic gating based on confidence and convergence thresholds reduces false positives by preventing alarm escalation in response to short-lived artefacts (e.g., motion-induced disturbances). High-cost response procedures are activated only when the predicted state remains stable over a defined temporal window and meets predefined confidence thresholds.
[0275] 11.3 Edge and loT Devices (Power-Constrained Operation)
[0276] Battery-powered edge devices (e.g., smart cameras, distributed environmental sensors, UAVs) require bounded compute and memory usage.
[0277] Technical value: The controlled retention policy (time-dependent decay and threshold-based pruning) limits long-term memory growth, while the FSM “listening” state suppresses costly processing paths until conditions justify activation. This supports reduced average compute load and improved power efficiency by avoiding continuous high-cost processing in the absence of stable, high-confidence events.
[0278] 11.4 Advanced Driver Assistance Systems (ADAS) and Autonomous Vehicles
[0279] ADAS platforms require reliable fusion across radar, LiDAR and camera inputs under strict timing constraints.
[0280] Technical value: The architecture provides an engineering-style “debounce” effect at the system level: sensor fusion is time-aligned and predictive escalation is gated by explicit stability criteria, reducing susceptibility to single-sensor outliers and transient noise that may otherwise trigger abrupt control reactions (e.g., false braking events).
[0281] 11.5 Real-Time Natural Language Processing and Multimodal HMI
[0282] Real-time assistants and multimodal human-machine interfaces operate in noisy environments and must reject irrelevant or inconsistent inputs.
[0283] Technical value: Temporal alignment across modalities (e.g., audio plus contextual signals and / or visual cues) combined with entropy-regulated cross-modal weighting increases robustness to interference. FSM-based gating prevents escalation to output actions unless the predictive output exceeds the required confidence threshold and satisfies stability criteria, improving reliability of real-time interaction.
[0284] Terminology Clarification
[0285] The following terminology is used for clarity and consistency of description.
[0286] The terms are introduced as descriptive labels and not as additional structural components.Each term corresponds to mechanisms disclosed in the filing of 28 February 2025 and does not extend the scope of the invention as originally disclosed.
[0287] 1. Hierarchical State Engine (HSE)
[0288] In this application, the term “Hierarchical State Engine (HSE)” refers to the hierarchical stateforming and retention mechanisms corresponding to the Hierarchical Aperceptive Memory (HPA 120) shown in Fig. 1 and disclosed in the original filing.
[0289] The term HSE does not denote an additional module.
[0290] It is a descriptive designation for the integrated operation of
[0291] • a short-term layer (STM);
[0292] • a mid-term layer (MTM);
[0293] • a long-term layer (LTM);
[0294] • retention mechanisms including time-dependent decay and selective pruning; and
[0295] • the interface between said layers and the predictive module (130).
[0296] The HSE performs deterministic update operations based on synchronised input vectors (via 110) and produces an internal hierarchical state representation for use by the predictive module (130) and the finite-state decision unit (140).
[0297] Accordingly, HSE is used solely as an engineering label for the disclosed HPA (120) and does not introduce new subject matter.
[0298] 2. State Stapp= S(t)_app
[0299] The notation S(t)_app denotes the internal hierarchical state representation generated by HPA (120) at time t.
[0300] S(t)_app does not define an additional module or an independent data structure, but is used as a symbolic notation for the already disclosed state formed by the cooperation of the layers of HPA (120).
[0301] In example embodiments, S(t)_app may be expressed as a composite representation including one or more of
[0302] • the current short-term state vector produced by the STM;
[0303] • a mid-term representation derived from selection of a best-matching element within the MTM (e.g., an identifier and / or an associated weight vector);• selected long-term entries stored in the LTM according to predefined retention criteria; and
[0304] • state-related measures computed within the disclosed processing pipeline (e.g., measures of stability and / or confidence used by the FSM gating logic).
[0305] The notation S(t)_app therefore serves as a compact reference to the disclosed multi-layer internal state and does not introduce a new computational entity.
[0306] 3. Adaptive Cognitive Path (ACP)
[0307] The term “Adaptive Cognitive Path (ACP)” is used herein as a descriptive designation for the dynamic processing-selection and execution-path control mechanisms disclosed in the original application.
[0308] ACP does not constitute a separate structural module.
[0309] It refers to the cooperative operation of disclosed components, including:
[0310] • multisensory fusion and cross-modal weighting mechanisms;
[0311] • state-derived weighting and selection operations;
[0312] • the predictive module (130);
[0313] • the finite-state decision unit (140), including threshold-based gating; and
[0314] • in relevant embodiments, selection and assignment of processing workloads across available computational resources, consistent with the disclosure illustrating heterogeneous execution (e.g., CPU / GPU / edge processing units).
[0315] The mechanisms underlying ACP are disclosed in the original filing, including the description of multimodal fusion and weighting, the hierarchical memory mechanisms, and the FSM-based gating logic.
[0316] Accordingly, ACP represents a descriptive grouping of disclosed mechanisms and does not introduce additional technical features beyond those already disclosed.Important Clarification
[0317] None of the above terminology modifies the structure of the claims.
[0318] The claimed invention remains defined by the modules and mechanisms as originally disclosed, including:
[0319] • Multisensory Input Module (100);
[0320] • Synchronization Unit (110);
[0321] • Hierarchical Aperceptive Memory (HPA) (120);
[0322] • Predictive Module (130);
[0323] • Finite-State Machine (FSM) Decision Unit (140); and
[0324] • Output Interface (150).
[0325] The introduced terminology is provided solely to improve descriptive clarity and technical consistency.
[0326] attachments
[0327] Fig. 1
[0328] Fig. 2
[0329] Fig. 3
[0330] Fig. 4
[0331] Fig. 5
Claims
AMENDED CLAIMSreceived by the International Bureau on07 July 2026 (07.07.2026)Claim 1 An apparatus for processing multisensory data in an Intelligent Virtual Being (IVB), comprising:• a multisensory input unit configured to receive input signals from at least two distinct sensors corresponding to different modalities selected from visual, auditory, textual, motion-based, biometric, contextual, or meta-modal sources;• a signal synchronization unit comprising a digital clock module configured to temporally align the received input signals within a synchronization window of less than 10 milliseconds prior to further processing;• a hierarchical aperceptive memory module configured to process the synchronized input signals, the hierarchical aperceptive memory module comprising:o (a) a short-term memory subunit implemented as a Long Short-Term Memory (LSTM) network configured to process sequential dependencies in the synchronized input signals and generate a hidden state representation;o (b) a mid-term memory subunit implemented as a Self-Organizing Map (SOM) network configured to:■ (i) compute distances between an output vector derived from the hidden state representation and neuron weight vectors arranged in a multidimensional grid,■ (ii) select a winning neuron based on a minimum-distance criterion, and■ (iii) update the weights of at least the winning neuron according to a competitive learning rule; ando (c) a long-term memory subunit configured to apply a controlled forgetting mechanism based on:■ (i) a logarithmic decay function dependent on elapsed time since last access, and■ (ii) a frequency -based pruning threshold configured to remove stored data when its occurrence frequency falls below a predefined percentage over a predefined number of processing cycles;• a predictive module comprising a cross-modal attention subroutine configured to:o (i) calculate attention weights for each modality using a key-query-value mechanism based on a cosine similarity function normalized by magnitudes of input vectors,o (ii) assign modality-specific weights based on a learned weighting function minimizing Shannon entropy calculated on a probability distribution of modality embeddings, and o (iii) combine modality-specific features using a weighted averaging function employing a softmax activation with a temperature coefficient in the range of 0.5 to 1.5;• a decision-making unit implemented as a finite state machine comprising at least a listening state, a processing state, and an output state, wherein:o (i) transition from the listening state to the processing state occurs when an output of the predictive module exceeds a predefined threshold, ando (ii) transition from the processing state to the output state occurs when convergence is detected, convergence being defined as a change in predictive output over three consecutive time steps falling below a predefined value; and• an output interface module comprising control circuitry configured to generate control signals in response to the finite state machine state transitions, the control signals being adapted todynamically modify at least one of a display layout, interface parameters, or external device behavior in an loT or robotic environment;• wherein the synchronization unit, the hierarchical aperceptive memory module including said controlled forgetting mechanism, and the finite state machine cooperate to gate processing such that end-to-end latency for generating predictive outputs is less than or equal to 50 milliseconds.Dependent Claims2. The apparatus of claim 1, wherein the controlled forgetting mechanism applies an update function defined as:wnew =wold+k- log( 1 +st T threshold) ■ (x~ wold)where w denotes a weight vector, x denotes a current input or context vector, t denotes elapsed time since last access, T threshold is a time threshold parameter, and s is a scaling factor.
3. The apparatus of claim 2, wherein records are pruned when an occurrence frequency falls below 5% over 50 consecutive processing cycles.
4. The apparatus of claim 1, wherein the signal synchronization unit aligns received input signals within a synchronization window less than or equal to 10 milliseconds.
5. The apparatus of claim 1, wherein the finite state machine gates processing such that an end-to-end latency for generating said predictive outputs is less than or equal to 50 milliseconds.
6. The apparatus of claim 1, wherein the predictive module employs a temperature coefficient in the range 0.5 to 1.5, preferably between 0.8 and 1.2.
7. The apparatus of claim 1, wherein the predefined threshold for transition from the listening state to the processing state is greater than or equal to 0.7.
8. The apparatus of claim 1, wherein the predefined value for convergence is less than 0.01 over said three consecutive time steps.
9. The apparatus of claim 1, configured to process input signals from at least fifteen distinct sensory sources.
10. The apparatus of claim 2, wherein k is in the range 0.05 to 0.2, T threshold is in the range 30 to 120 seconds, and s is in the range 1 to 10.
11. The apparatus of claim 1, wherein the LSTM network comprises at least 128 memory cells and includes a forget gate with a decay factor configurable in the range 0.4 to 0.6.
12. The apparatus of claim 1, wherein the SOM network consists essentially of at least 64 neurons arranged in an 8x8 matrix and is trained using competitive learning with a multidimensional Gaussian neighbourhood function having a learning rate in the range of 0.01 to 0.1.
13. The apparatus of claim 1, wherein the learned weighting function is derived using an Adam optimiser with a learning rate of 0.001, betal of 0.9, beta2 of 0.999, and epsilon of 1 x 10-8.
14. A computer-implemented method for multimodal sensory-data processing, the method being performed by one or more processors and comprising:(a) receiving signals from a plurality of distinct modalities and temporally aligning said signals within a synchronisation window prior to further processing;• (b) processing the aligned signals in a hierarchical aperceptive memoiy comprising short-term, mid-term, and long-term subunits including an LSTM subunit and a SOM subunit;• (c) applying a controlled forgetting mechanism based on logarithmic time-dependent decay and frequency-based pruning;• (d) processing the aligned signals using a predictive module comprising cross-modal attention that calculates modality weights using cosine similarity within a key-queiy-value mechanism and a learned weighting function minimizing Shannon entropy on a probability distribution of modality embeddings;(e) combining modality-specific features using a temperature-controlled softmax to generate predictive outputs;(f) generating decision signals via a finite state machine comprising listening, processing and output states with threshold-based transitions including a convergence criterion defined by a change in predictive output over three consecutive time steps falling below a predefined value; and(g) outputting, via an output interface module (150) comprising control circuitry, control signals derived from said decision signals to cause dynamic modification of at least one of a display layout, interface parameters, or behavior of an external device in an loT or robotic environment; wherein the method gates processing such that end-to-end latency for generating predictive outputs is less than or equal to 50 milliseconds.
15. The method of claim 14, wherein the synchronisation window is less than or equal to 10 milliseconds.
16. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method of claim 14, including causing the output interface module (150) to output control signals to dynamically modify at least one of a display layout, interface parameters, or behavior of an external device in an loT or robotic environment.
17. The computer program product of claim 16, wherein the synchronisation window is less than or equal to 10 milliseconds and the end-to-end latency is less than or equal to 50 milliseconds.