Array perception and confidence isolated interlocked human-machine collaborative control system and method

CN122808796APending Publication Date: 2026-09-25BEIJING MAIJIN TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610969620.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0011]本发明提供一种阵列感知与置信隔离的联锁人机协同控制系统及方法,旨在解决现有多模态交互在轨道交通联锁安全苛求系统中硬件感知能力浪费、AI裁决不可信的技术问题

Benefits of technology

(一)物理层零感知防攻击

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122808796A_ABST
    Figure CN122808796A_ABST
Patent Text Reader

Abstract

The application discloses an array perception and confidence isolation interlocking man-machine cooperative control system and method, relates to the field of rail transit signal control technology, and is arranged on an interlocking host computer and operates as a module independent of original software of the interlocking host computer. A microphone array extracts spatial phase features in parallel while performing beamforming voice enhancement, and a software comparison and decision sub-module realizes passive living body detection based on a near-field acoustic physical model; after an interactive auxiliary domain generates a candidate instruction with a confidence label, a deterministic mapping engine forcibly converts and peels off the confidence label; a safe execution domain only performs space-time conflict pre-checking and resource arbitration according to a digital twin and an interval algebra model; and a ruling instruction is converted into a simulated mouse click operation by a mouse click simulator and is sent to an interlocking system for execution through an existing communication protocol of the host computer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rail transit signal control technology, and in particular to an interlocking human-machine collaborative control system and method based on array sensing and confidence isolation. Background Technology

[0002] Computer-based interlocking systems are the core safety system for rail transit signal control, used to control trackside equipment such as switches and signals and manage routes. Their safety integrity level must meet SIL4 requirements. Currently, most in-service computer-based interlocking systems adopt a single-modal human-machine interaction mode. Operators click on command buttons and signal equipment buttons on the host computer's station map using a mouse, and the host computer sends the operation commands to the interlocking logic unit for execution via existing communication protocols. This mode has the following inherent drawbacks: Low operational efficiency: The serial mode of "observe the station - locate the cursor - click the target - confirm feedback" forces the operator's visual attention to frequently switch between target location and status monitoring, which can easily lead to misoperation due to distraction, especially in emergency response scenarios.

[0003] Passive security verification: Existing systems mainly rely on post-execution device interlock checks, lacking continuous verification of operator identity and proactive intent verification before command issuance, as well as prediction of spatiotemporal conflicts. Misoperations can only be passively discovered after execution, which is "post-event remediation" rather than "pre-event prevention." Operator authentication is usually only performed during login, and subsequent operations lack a continuous biometric authentication mechanism.

[0004] Poor multitasking capabilities: Mouse operation requires the operator to keep their hands focused on the host computer. In scenarios such as emergency response where operators need to simultaneously record dispatch commands and review ledgers, it is difficult for operators to handle multiple hand tasks, posing a risk of operation delays or interruptions. Voice interaction can provide a supplementary operation channel when both hands are occupied, achieving decoupling of hand operations.

[0005] Insufficient traceability: Operation logs only record click results, not operation intentions, making it difficult to distinguish between intentional and accidental operations. Furthermore, existing operation log protection methods generally involve setting access permissions or passwords, but access permissions are easily modified by humans, and access passwords are easily cracked, leading to a risk of log tampering.

[0006] In recent years, voice control technology has been widely used in the consumer electronics field. However, the extreme requirements for safety, real-time performance, and determinism in rail transit interlocking systems pose significant security risks if general-purpose voice recognition technology is directly integrated. Existing voice control solutions suffer from the following architectural flaws: (1) Limited hardware functionality: The microphone is only used for sound pickup and cannot defend against recording replay attacks at the physical layer. Recording replay attacks are a recognized security threat to voiceprint recognition systems. To defend against such attacks, existing solutions generally require users to read aloud randomly generated dynamic number strings or specific text (active liveness detection). Each authentication requires additional steps, which is unacceptable in interlocking control scenarios with extremely high time requirements, such as emergency response.

[0007] (2) AI overstepping its bounds: The probabilistic results (confidence levels) output by deep learning models are directly or indirectly used to influence safety operation decisions. According to functional safety standards, it is strongly recommended that formal proof methods be used for modeling and verification of safety-critical systems. However, the black-box nature of deep learning models is fundamentally contradictory to the requirements of formal verification, and their probabilistic outputs are difficult to meet the rigid requirements of safety-critical systems for deterministic logic.

[0008] It is worth noting that passive liveness detection technology using microphone arrays has been explored in the consumer electronics field. For example, the ArrayID system (Your Microphone Array Retains Your Identity, USENIX Security 2022) uses the microphone array of a smart speaker to extract an "array fingerprint" for passive liveness detection, and announcement number CN114155850B discloses a voice spoofing attack detection system based on a microphone array. However, the above solutions all rely on deep learning models (such as convolutional neural networks, support vector machines, etc.) to classify and decide on the extracted spatial features. Their outputs are probabilistic, and the black-box nature of the models makes it difficult to formally verify the decision results. When such solutions are directly applied to the stringent safety requirements of rail transit interlocking, they face the following fundamental obstacles: (1) The misjudgment rate of deep learning models (even the best in laboratory environments still has an error of 0.16%) is unacceptable in interlocking scenarios, and a single misjudgment may lead to a serious safety accident; (2) Probabilistic outputs cannot meet the rigid requirements of SIL2 functional safety for deterministic logic; (3) The change management introduced by model updates and retraining is complex and difficult to meet the configuration management requirements of safety certification. This invention proposes a novel solution based on near-field acoustic physical models and deterministic numerical comparisons to address the special requirements of passive liveness detection in the aforementioned stringent safety scenarios, which is fundamentally different from existing deep learning-based solutions.

[0009] For a long time, the following technical biases have existed in this field: (1) It is generally believed that voice interaction is inherently unsuitable for systems with stringent safety requirements because it is based on a probabilistic AI model, and therefore the voice channel has always been regarded as a "forbidden zone" in the field of interlocking control; (2) The deterministic requirements of existing safety standards (such as EN50129) for software verification are regarded as an insurmountable obstacle for voice interaction; (3) The closed nature of existing interlocking systems and the SIL4 safety level requirements mean that the introduction of any new interaction method is considered to require modification of existing systems or development of dedicated safety interfaces, which is costly and risky. This invention fundamentally breaks the above biases through the "confidence isolation" architecture - the AI ​​probabilistic output is strictly limited to the auxiliary layer, and only deterministic logic exists in the safety decision chain, thus proving that voice interaction can be safely introduced into interlocking systems without reducing the SIL level; at the same time, zero-intrusion docking with existing systems is achieved through a mouse click simulator, avoiding expensive and high-risk system modification.

[0010] There is currently no multimodal human-machine collaborative control scheme that can achieve low-intrusion deployment at the interlocking host computer level while meeting safety requirements. This invention fundamentally breaks through the above-mentioned technical biases by combining a trust isolation architecture with a mouse click simulator. At the same time, it interfaces with the existing SLL4 level interlocking system in a zero-intrusion manner, proving the feasibility of voice interaction in scenarios with stringent safety requirements. Summary of the Invention

[0011] This invention provides a human-machine collaborative control system and method for interlocking systems based on array sensing and confidence isolation. It aims to address the technical problems of wasted hardware sensing capabilities and unreliable AI decisions in existing multimodal interaction systems with stringent safety requirements in rail transit interlocking systems. This invention is entirely deployed on the interlocking host computer. It achieves non-intrusive liveness detection through the multiplexing of spatial features of the microphone array, completely excludes AI probability outputs from the safety decision chain through a confidence isolation architecture, and converts voice commands into simulated mouse clicks. These are then sent to the interlocking logic unit for execution via the host computer's existing communication protocol. This achieves safe and reliable human-machine collaborative operation without modifying the existing interlocking system hardware and software. The synergistic effect of the overall architecture of this invention lies in the seamless integration of acoustic and physical liveness detection, functional safety isolation architecture, and mouse simulation. For the first time, it introduces a voice interaction channel to SIL2-level host computer security without lowering the existing SIL4 safety level or modifying the existing system. The modules collectively solve the systemic challenges of non-intrusive attack prevention, confidence domain isolation, and zero-intrusion compatibility, rather than simply being superimposed.

[0012] This invention constructs an interactive auxiliary domain and a secure execution domain that are physically and logically isolated. All modules are deployed inside the interlocking host computer. The specific technical solution is as follows: On the one hand, an interlocking human-machine collaborative control system with array sensing and confidence isolation is provided, deployed on an interlocking host computer, characterized by comprising: The interactive assistance domain, including a microphone array interface, a spatial awareness module, and a semantic parsing module, is used to convert operator speech into candidate instruction data packets with confidence labels. The spatial awareness module integrates a passive liveness detection unit and a beamforming module. The semantic parsing module is used to extract structured operation instructions from the text. The secure execution domain, deployed inside the host computer, is connected to the interactive assistance domain via a one-way data channel and includes a deterministic mapping engine, a pre-verification unit, and a resource arbitration unit. After receiving the candidate instruction data packets, the secure execution domain discards the confidence labels and makes security decisions based solely on the deterministic physical model and interlocking rules. A hierarchical confirmation unit is used to confirm the secure execution of candidate instruction data packets. Before entering the secure execution domain, a differentiated confirmation strategy is implemented based on the confidence level; confidence labels are physically stripped before being passed to the secure execution domain; a mouse click simulator is used to generate corresponding mouse click operation sequences based on the arbitrated instructions, simulating the operator clicking the instruction buttons and signal device buttons on the host computer's site diagram; an execution feedback unit is used to send the simulated click operations to the interlocking logic operation unit for execution via the host computer's existing communication protocol, and synchronously provide feedback on the execution results through visual and auditory channels; a log audit unit, with a built-in hardware security module and read-only solid-state memory, is used to record the entire process operation information using hash chain and hardware security module digital signature methods, forming an immutable operation traceability chain.

[0013] On the other hand, a method for interlocking human-machine collaborative control based on array sensing and confidence isolation is provided, which is applied to an interlocking human-machine collaborative control system based on array sensing and confidence isolation, including the following steps: Multi-channel audio signals are acquired via a microphone array; target speech is extracted using a beamforming module based on the audio signals, and phase group delay features are extracted simultaneously. A threshold envelope is dynamically calculated based on a near-field physical model and array parameters for comparison. If the threshold exceeds the envelope, it is determined to be a replay attack and blocked; if it falls within the envelope, it is determined to be genuine human voice and subsequent processing is allowed. The target speech is converted into text, and structured operation instructions are extracted to generate candidate instruction data packets. A hierarchical confirmation unit executes a differentiated confirmation strategy based on confidence labels and physically removes the labels before the data packets are transmitted to the secure execution domain. A deterministic mapping engine forcibly converts the confirmed instructions into standard equipment operation codes within the interlocking system. The secure execution domain receives the standard equipment operation codes and discards them. The credibility label calls the digital twin model of the station to calculate the inviolable spatiotemporal envelope. If the target device is within the envelope, a security veto is triggered and the instruction is intercepted; otherwise, it is latched as an instruction to be executed and enters resource arbitration. The resource arbitration unit maps the instruction to be executed as a request to occupy the resources of a specific device and arbitrates based on the principle of physical mutual exclusion. The instruction that passes the arbitration is sent to the mouse click simulator to generate a simulated click sequence to click the button on the station map, and then sent to the interlocking logic operation unit for execution through the existing communication protocol. After the interlocking logic operation unit executes, it provides synchronous feedback on the results through the visual and auditory channels. The entire operation information is recorded using a hash chain digital signature method based on the hardware security module, forming an immutable operation traceability chain.

[0014] On the other hand, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the method described in any of the preceding methods.

[0015] Beneficial effects The present invention has the following beneficial effects: (a) Physical layer zero-awareness attack prevention While performing beamforming voice enhancement, the microphone array simultaneously extracts spatial phase features for passive liveness detection. This process is completed automatically when the operator issues normal voice commands, without requiring the reading of random numbers or any additional steps, thus completely eliminating the threat of recording and playback attacks and achieving "unobtrusive" authentication.

[0016] (ii) Confidence isolation ensures security and compliance At the architectural level, this invention forcibly strips any confidence values ​​output by the AI ​​model (noise reduction, speech recognition, semantic parsing) when it enters the safe execution domain. Safety decisions are executed entirely by a deterministic physical model based on digital twins and interval algebra, fundamentally avoiding the uncertainty issues of purely data-driven black-box models in safety-critical scenarios. This meets the SIL2 functional safety requirements of the host computer for deterministic logic without lowering the original SIL4 safety level of the interlocking logic operation unit.

[0017] (iii) Seamless compatibility with existing systems and minimal invasive deployment This invention is deployed within the interlocking host computer, operating as a module independent of the original host computer software. It does not alter any code, configuration, or communication protocol of the original host computer software, nor does it change the hardware architecture and safety verification mechanism of the existing interlocking logic operation unit. Voice commands are converted into simulated mouse clicks via a mouse click simulator—calling the operating system API to simulate the operator clicking command buttons and signal device buttons on the host computer's station map sequentially, completely equivalent to the operator manually clicking on the host computer interface. The source of the commands received by the interlocking logic operation unit is indistinguishable from manual mouse operation, requiring no hardware modification or interface adaptation. Furthermore, by embedding a safety execution domain within the host computer as a daemon process / background verification module, this invention allows the operator to immediately switch back to traditional mouse operation in the event of system degradation or failure. The original functions of the host computer and the operation of the interlocking logic operation unit remain completely unaffected, ensuring uninterrupted train operation safety.

[0018] (iv) Dynamic physical arbitration eliminates logical blind spots The static rules that determine security priority based on the semantic meaning of operation command text (such as "open / close" and "lock / unlock") are abandoned. Instead, a dynamic arbitration mechanism based on the real-time occupancy topology of the device and the physical mutual exclusion state is adopted to ensure that the system's behavior has physical determinism under any edge conditions.

[0019] (v) Tamper-proof traceability throughout the entire chain The hash pointer chain and the hardware security module digital signature bind the operation intent, voiceprint credentials, instruction parameters, and execution results throughout the entire process, meeting the highest traceability requirements of SIL4 level. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the dual-domain isolation system architecture deployed on the interlocking host computer provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the principle of parallel processing of microphone array spatial features and passive liveness detection provided in an embodiment of the present invention; Figure 3 This is a flowchart of the interlocking human-machine collaborative control method with array sensing and confidence isolation provided in an embodiment of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be considered as limitations on the invention, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this invention.

[0022] like Figure 1 The diagram shown is a schematic of a dual-domain isolation system architecture deployed on an interlocking host computer according to an embodiment of the present invention. It describes an interlocking human-machine collaborative control system with array sensing and confidence isolation, deployed on an interlocking host computer, characterized by comprising: The interactive assistance domain, including a microphone array interface, a spatial awareness module, and a semantic parsing module, is used to convert operator speech into candidate instruction data packets with confidence labels. The microphone array is a uniform linear array with 4 to 8 channels and a spacing of 3 to 5 cm between adjacent microphones. The spatial awareness module integrates a passive liveness detection unit and a beamforming module. The semantic parsing module is used to extract structured operation instructions from the text. The passive liveness detection unit and the beamforming module share the same microphone array data stream, which is copied, split, and processed in parallel through the data link layer. Specifically, the passive liveness detection unit extracts the phase group delay features between each channel of the array, dynamically calculates the threshold envelope based on the near-field spherical wave propagation physical model and array geometric parameters, and determines whether it is a real person speaking or a replay attack through deterministic numerical comparison. The beamforming module is used to extract the speech signal in the target direction through spatial filtering.

[0023] The secure execution domain, deployed within the host computer, connects to the interactive auxiliary domain via a unidirectional data channel. It includes a deterministic mapping engine, a pre-verification unit, and a resource arbitration unit. The input of the deterministic mapping engine is connected to the interactive auxiliary domain via the unidirectional data channel, and its output is connected to the input of the pre-verification unit, which in turn is connected to the input of the resource arbitration unit. Upon receiving candidate instruction data packets, the secure execution domain discards the confidence labels and makes security decisions solely based on the deterministic physical model and interlocking rules. The deterministic mapping engine is used to forcibly convert natural language entities into standard equipment operation codes within the interlocking system. The pre-verification unit invokes the station's digital twin model and uses deterministic interval algebra to calculate the inviolable spatiotemporal envelope, performing security conflict pre-verification on the instructions. The resource arbitration unit arbitrates instructions based on the physical mutual exclusion principle of the intersection of resource occupancy sets. The deterministic mapping engine uses an Aho-Corasick automaton as its index structure and employs the longest match priority principle to perform forced conversion of natural language entities into internal standard device operation codes. By receiving candidate instructions from the semantic parsing module, it performs forced matching in the mapping table in read-only memory, strictly converting the operator's natural spoken entity words into internal standard device operation codes that are uniquely recognizable by the interlocking system. If the match is successful, a unique machine code is output; if the match fails, the process terminates. The device identification code mapping table in read-only memory cannot be modified during operation. After receiving the standard equipment operation code, the pre-verification unit discards the confidence label, calls the station's digital twin model stored in a graph data structure, and obtains the current status of the equipment, including occupied and idle states. The digital twin model asynchronously updates node attributes through interrupt feedback from the interlocking logic operation unit. Using the train's current emergency braking distance and the latest turnout response time as boundaries, it calculates the inviolable spatiotemporal envelope within the future time window. If the target equipment is within the spatiotemporal envelope, a safety veto is triggered, the current instruction is intercepted, the process is terminated, and the operator is prompted to manually handle the situation using a mouse through both visual and auditory channels. If the target equipment is not within the spatiotemporal envelope... Outside the spatiotemporal envelope, the instruction is latched as a pending execution instruction. If the digital twin model of the station does not receive a status update from the lower-level machine within the preset synchronization interval, the model is deemed out of sync, and the safe execution domain automatically degrades to a mouse-only operation mode, prohibiting all voice commands. The preset synchronization interval is determined based on the device status refresh cycle of the interlocking logic operation unit, and is no more than 100ms by default. The resource arbitration unit maps each operation instruction to a request for the use of specific device resources. If the resource use sets of multiple requests overlap, they are handled according to the physical mutual exclusion principle of first-come, first-served, and subsequent requests wait and are canceled upon timeout. Arbitration does not use the semantic meaning of the instruction text as a judgment weight.

[0024] The interlocking logic operation unit receives equipment operation commands generated by simulated clicks through the existing communication protocol of the host computer; it performs interlocking logic checks on the operation commands, including route condition verification, turnout position consistency verification, section occupancy status verification, and adversarial route mutual exclusion verification; if the interlocking logic check passes, it outputs control commands to the corresponding trackside equipment drive unit to drive turnout switching or signal lighting; after the trackside equipment drive unit executes the control command, it collects the equipment status and feeds it back to the interlocking logic operation unit; the interlocking logic operation unit transmits the execution result back to the host computer through the existing communication protocol for the execution feedback unit to perform synchronous feedback through the visual and auditory channels; the interlocking logic operation unit operates independently of the interactive auxiliary domain and safety execution domain in the host computer, and its internal interlocking logic does not contain any neural network or machine learning model.

[0025] The tiered confirmation unit, deployed between the interactive assistance domain and the host computer user interface layer, is used to execute differentiated confirmation strategies based on confidence levels before candidate instruction data packets are passed into the secure execution domain. Confidence labels are physically removed before being passed into the secure execution domain. The differentiated strategies executed by the tiered confirmation unit are as follows: when the confidence level reaches a preset first confidence threshold (e.g., 0.85), a confirmation prompt is displayed on the interface, and if no cancellation operation is performed within a set time, the process automatically proceeds to the next step; when the confidence level reaches a preset second confidence threshold (e.g., 0.6) but does not reach the preset first confidence threshold, the instruction content is read back via speech synthesis, and the process proceeds to the next step after the operator confirms via voice; when the confidence level does not reach the preset second confidence threshold, the process refuses to automatically proceed to the next step, and the operator is forced to manually select the target device and operation by clicking on the host computer site map with the mouse; confirmation is only used to verify the operator's intent and does not change the decision conclusion of the secure execution domain.

[0026] The mouse click simulator connects to the output of the resource arbitration unit and is used to generate corresponding mouse click operation sequences based on the arbitration instructions, simulating the operator clicking instruction buttons and signal device buttons on the host computer's site map. The mouse click simulator converts the arbitration instructions into a sequence of screen coordinates on the host computer's site map according to a predefined operation instruction and screen coordinate mapping table, storing the screen coordinates as a percentage of the screen area. The mouse click simulator calls the host computer's operating system's mouse / touch application programming interface to generate simulated mouse movement and click operations according to the coordinate sequence. The host computer supports existing communication protocols including CAN bus, Ethernet UDP / TCP, or serial communication.

[0027] The execution feedback unit is used to send the simulated click operation to the interlocking logic operation unit for execution through the existing communication protocol of the host computer, and to provide synchronous feedback of the execution result through the visual and auditory channels. The log auditing unit, with its built-in hardware security module and read-only solid-state storage, records the entire operation process using a hash chain and digital signatures from the hardware security module, forming an immutable operation traceability chain. Each record kept by the log auditing unit includes a timestamp, operator identifier, liveness detection result, standard device operation code, simulated click coordinate sequence, execution result, the complete hash value of the previous record, and the digital signature of the hash value of that record from the hardware security module.

[0028] The system achieves complete isolation from the original software of the interlocking host computer through the operating system process boundary; the system does not modify the code or configuration files of the original software of the interlocking host computer, and only interacts with the software of the interlocking host computer by simulating mouse clicks; when the system is downgraded or fails, the operator can directly switch back to the traditional mouse operation mode.

[0029] Figure 2This is a schematic diagram of the parallel processing of microphone array spatial features and passive liveness detection principle provided in the embodiment of the present invention. The schematic diagram shows the functional architecture of the microphone array spatial perception module. An 8-channel uniform linear (element spacing 4cm) microphone array is used as the signal acquisition front end. The multi-channel audio data is acquired using a link layer splitting mechanism to realize parallel processing of speech enhancement and passive liveness detection. The speech enhancement pathway first uses the GCC-PATH algorithm to locate the sound source, and then combines delay summation beamforming to achieve spatial filtering and extraction of speech in the target direction. The signal enhancement is then completed by the CRNN deep learning noise reduction module, and finally high-quality speech data is output to the speech recognition module. The passive liveness detection pathway first extracts the phase group delay rate difference feature between multi-channel signals. Based on the near-field spherical wave propagation physical model and array geometric parameters, a decision envelope with tolerance threshold is dynamically generated. The software decision submodule performs a deterministic comparison between the measured phase features and the envelope range. If the phase features are continuous, smooth and stably fall within the envelope range, it is determined to be a real person's near-field voice, and the signal is allowed to enter the subsequent processing flow. If there is nonlinear distortion in the phase and the jitter exceeds the envelope threshold, it is determined to be a speaker playback attack, triggering a logic blocking and voice alarm mechanism.

[0030] like Figure 3 The diagram shown is a flowchart of an interlocking human-machine collaborative control method based on array sensing and confidence isolation provided in an embodiment of the present invention. The method includes: Multi-channel audio signals are acquired via a microphone array; target speech is extracted using a beamforming module based on the audio signals, and phase group delay features are extracted simultaneously. A threshold envelope is dynamically calculated based on a near-field physical model and array parameters for comparison. If the threshold exceeds the envelope, it is determined to be a replay attack and blocked; if it falls within the envelope, it is determined to be genuine human voice and subsequent processing is allowed. The target speech is converted into text, and structured operation instructions are extracted to generate candidate instruction data packets. A hierarchical confirmation unit executes a differentiated confirmation strategy based on confidence labels and physically removes the labels before the data packets are transmitted to the secure execution domain. A deterministic mapping engine forcibly converts the confirmed instructions into standard equipment operation codes within the interlocking system. The secure execution domain receives the standard equipment operation codes and discards them. The credibility label calls the digital twin model of the station to calculate the inviolable spatiotemporal envelope. If the target device is within the envelope, a security veto is triggered and the instruction is intercepted; otherwise, it is latched as an instruction to be executed and enters resource arbitration. The resource arbitration unit maps the instruction to be executed as a request to occupy the resources of a specific device and arbitrates based on the principle of physical mutual exclusion. The instruction that passes the arbitration is sent to the mouse click simulator to generate a simulated click sequence to click the button on the station map, and then sent to the interlocking logic operation unit for execution through the existing communication protocol. After the interlocking logic operation unit executes, it provides synchronous feedback on the results through the visual and auditory channels. The entire operation information is recorded using a hash chain digital signature method based on the hardware security module, forming an immutable operation traceability chain.

[0031] An embodiment of the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to execute an interlocking human-machine collaborative control system of array sensing and confidence isolation.

[0032] Example 1: Array Spatial Feature Extraction and Software Comparison for Liveness Detection In this embodiment, the control panel is equipped with an 8-channel uniform linear microphone array (e.g., Figure 2 As shown in the attached figure (labeled 100), the adjacent spacing is 4 cm, and the sampling rate is 16 kHz.

[0033] When the operator issues the voice command "arrange downlink routes", the system performs the following parallel processing on the multi-channel audio data: Path A (Speech Enhancement): The target direction (0°±15°) signal is extracted using delayed summation beamforming, fed into a convolutional recurrent neural network noise reduction model with 1.8M parameters, and the enhanced audio is output to the speech recognition engine. The specific structure of the CRNN model is as follows: The input layer receives a single-channel 16kHz time-domain speech signal, which is converted into frequency-domain features through short-time Fourier transform (STFT, window length 25ms, frame shift 10ms); the encoder consists of three two-dimensional convolutional layers (layer 1: kernel size (5,5), number of channels 32, stride (2,2); layer 2: kernel size (3,3), number of channels 64, stride (2,1); layer 3: kernel size (3,3), number of channels 128, stride (2,1)), each layer is followed by batch normalization and ReLU activation function; the intermediate layer is a two-layer gated recurrent unit (GRU, 256 hidden units per layer); the decoder consists of three transposed convolutional layers, symmetrical to the encoder, to restore the original time resolution; the output layer outputs an enhanced time-domain signal through inverse short-time Fourier transform (iSTFT). The model was pre-trained using publicly available speech enhancement datasets (such as the DNSChallenge dataset), and then fine-tuned using real-world speech data from rail transit dispatching (no less than 200 hours, covering different speakers such as station staff and dispatchers, as well as various background noise environments). The loss function was a weighted combination of scale-invariant signal-to-noise ratio (SI-SNR) and mean squared error. The optimizer was Adam, with an initial learning rate of 1e-3 and a cosine annealing decay strategy.

[0034] It should be noted that, to prevent the AI ​​model from freezing due to abnormal input or resource contention, causing the interactive auxiliary domain to remain unresponsive for extended periods, this invention establishes a unified timeout monitoring mechanism for each AI processing step (including the convolutional recurrent neural network denoising inference in this step and the semantic parsing BERT inference in subsequent steps). Specifically, each AI inference call runs in an independent operating system thread and starts a real-time timer with a preset maximum execution time threshold of 100ms (this threshold can be adjusted based on actual testing of the host computer's computing resources, but should not exceed the watchdog synchronization interval). The timer starts when the inference call begins. If a result is returned normally within 100ms, the timer is reset, and the process continues; if no result is returned after the timeout, the operating system immediately terminates the inference thread, discards the currently processed audio frame data, records a log entry containing a timestamp, model name, and timeout duration, and releases GPU / CPU computing resources and memory buffers. This mechanism ensures that even if the AI ​​module freezes, the interactive auxiliary domain can recover to an idle state within milliseconds, awaiting the operator's next wake-up command, thus eliminating the risk of prolonged 'silence' of the voice channel due to AI module failure.

[0035] Path B (Passive Liveness Detection): Real-time calculation of the group delay rate of change of the phase difference between channel 1 and channel 8. The system has a built-in dynamic threshold envelope function pre-calculated based on the near-field spherical wave propagation equation. Specifically, assuming the distance between adjacent array elements is d (4cm), the incident angle of the sound source is θ (relative to the array normal), and the sound wave when a real person speaks is a near-field spherical wave, the time delay difference τ reaching the adjacent array elements satisfies: τ = (d·sinθ) / c + Δτ_diffraction, where c is the speed of sound, and Δτ_diffraction is the additional time delay caused by head diffraction (which can be pre-measured using the head correlation transfer function HRTF or approximated using a spherical diffraction model). In a preferred embodiment of the present invention, Δτ_diffraction is approximated using a spherical diffraction model, specifically: Δτ_diffraction ≈ (r_head / c)·(α+sinα), where r_head is the equivalent radius of the human head (empirical value is 0.085m~0.095m), and α is the azimuth angle of the sound source relative to the midpoint of the line connecting the two ears (which has a geometric conversion relationship with the incident angle θ of the array normal direction and is determined by the array geometry). The physical basis of this model is that the diffraction effect of the head on near-field spherical waves causes the wavefront to bend and delay, while loudspeaker reproduction (far-field approximates a plane wave) does not have this characteristic. Therefore, by comparing the measured group delay rate of change with the theoretical value (including Δτ_diffraction), it is possible to distinguish between human voices and loudspeaker reproduction without relying on any statistical model or machine learning classifier. For loudspeaker reproduction, the sound source is located in the far field and is approximately a plane wave. Its time delay difference is determined only by (d·sinθ) / c, lacking the group delay change caused by the wavefront curvature unique to near-field spherical waves. This invention pre-calculates the theoretical group delay rate of change based on the near-field spherical wave physical model for different incident angles θ, and sets a tolerance δ (δ is obtained through offline calibration. During calibration, multiple sets of human voice samples are collected in a real operating environment, and the standard deviation of their group delay rate of change is calculated. Three times the standard deviation is taken as the δ value), to obtain the threshold envelope range [f(d,θ)-δ,f(d,θ)+δ].

[0036] Human voices are near-field spherical waves, and their group delay rate of change has a continuous characteristic value that conforms to the physical laws of head diffraction, continuously falling within this envelope range; while loudspeaker reproduction is a far-field approximate plane wave or has nonlinear phase distortion due to electroacoustic conversion, and its group delay rate of change exceeds this envelope range.

[0037] The liveness detection decision is made by the software comparison decision submodule (see attached figure 101), which is deployed in the host computer and executes the following deterministic comparison logic: The group delay rate of change value extracted from receiver path B; Read the threshold envelope function f(d,θ)±δ pre-calculated based on the array geometry parameters; Perform numerical comparison: If the group delay rate of change ∈ [f(d,θ)-δ,f(d,θ)+δ], it is determined to be a real human voice, and the "pass" flag is output, allowing the audio data of path A to continue to be transmitted to the speech recognition engine; If the group delay change rate is ∉[f(d,θ)-δ,f(d,θ)+δ], it is determined to be a speaker replay attack. The software logic blocks the subsequent voice signal path (i.e. discards the current audio frame data and does not forward it to the speech recognition engine), triggers a "suspected replay attack" voice alarm, and terminates the current instruction process.

[0038] It should be noted that the system performs copying and splitting of the original multi-channel audio data at the data link layer (e.g., through memory copying or data mirroring). One path is supplied to the beamforming module (path A), and the other path is directly supplied to the liveness detection feature extraction submodule (path B). This ensures that the phase compensation operation in beamforming does not negatively affect the original phase features on which path B depends. The two processes are processed in parallel without interference.

[0039] The software's comparison and decision submodule follows deterministic logic (performing only numerical comparisons and threshold checks), does not contain any neural networks or machine learning models, and has good interpretability. The entire decision process takes less than 1ms, meeting the real-time requirements of interlocking systems.

[0040] Example 2: Confidence Segregation and Deterministic Decision In this embodiment, the interactive auxiliary domain semantic parsing module (based on the BERT-base-Chinese pre-trained model, fine-tuned using the rail transit dispatch command dataset) outputs the following candidate instruction JSON format data packet: { "intent":"open_signal", "object":"S1", "confidence":0.92, "timestamp":1702800000000 } The specific parameters for model fine-tuning are as follows: The training data comes from the dispatching command logs and simulator training data of a railway bureau, totaling approximately 500,000 labeled samples after anonymization, covering 12 intent categories such as "arranging routes," "open signals," "turnout locking / unlocking," "canceling routes," and "confirming departure," as well as approximately 800 entity words; a supervised fine-tuning strategy is adopted, with a batch size of 32, a learning rate of 2e-5 (using a linear warm-up + decay strategy, with warm-up steps accounting for 10% of the total), 5 training epochs, and a cross-entropy loss function. On the validation set, the intent recognition accuracy is 96.7%, and the entity recognition F1 score is 92.3%. It should be noted that the above accuracy and F1 score are only reference indicators of model performance and do not affect the deterministic decision-making in the safe execution domain.

[0041] The data packet is sent to the secure execution domain (Figure 201, deployed in the host computer, running as a daemon / background service) via a one-way data channel (Figure 200).

[0042] The secure execution domain software module performs the following deterministic steps, ignoring the 0.92 confidence level throughout: Step 1 – Deterministic Mapping: Query the device identifier mapping table loaded in memory (indexed by an Aho-Corasick automaton, containing approximately 5000 rules) to convert the natural language entity “S1” into the internal standard code 0x1A3F. In this mapping table, “S1” corresponds to only one unique standard code, eliminating one-to-many ambiguity.

[0043] Step 2 - Resource Request Parsing: Query the device-resource association table based on the standard code 0x1A3F to obtain the resource occupancy list of all switches and track sections within the protected route of this signal.

[0044] Step 3 – Spatiotemporal Envelope Conflict Detection: Read the current emergency braking distance envelope calculated by the digital twin model (e.g., the envelope already covers the section inside signal S1). The digital twin model stores the station topology in a graph data structure G=(V,E). Node V contains equipment identification code, equipment type, dynamic response delay parameters, and current status attributes, while edge E contains track distance and maximum permissible speed attributes. This model asynchronously updates node attributes through interrupt feedback from the interlocking logic unit. It should be noted that, as an alternative or supplementary method to interrupt feedback, the interlocking logic unit can also provide status updates to the digital twin model in the following way: the safety execution domain actively polls the equipment status register of the interlocking logic unit at a period not greater than a preset synchronization interval, and updates the node attributes of the digital twin model accordingly. When using the polling method, the polling period is no greater than 1 / 2 of the preset synchronization interval to ensure the timeliness and reliability of out-of-synchronization detection.

[0045] Step 4 - Deterministic Decision: Since the requested resource is located within the inviolable envelope, the secure execution domain directly outputs a "deny" signal and sends an independent prompt to the host computer interface: "Security domain interception: Segment occupied, instruction denied by the security domain".

[0046] Step 5 – Synchronization Degradation Protection: The system is equipped with a watchdog timer (see attached diagram, label 202) to monitor the state synchronization cycle between the digital twin model and the interlocking logic unit. If no device state update interruption is received within the preset synchronization interval (default 100ms), the model is considered out of sync, and the secure execution domain automatically degrades to "mouse-only" mode until the model is resynchronized. During the degradation period, all voice channel commands are unconditionally intercepted by the secure execution domain, and a "System Degradation, please use mouse operation" prompt is issued through both the interface and voice channels.

[0047] This embodiment clearly demonstrates the absolute veto power of the secure execution domain over the interactive assistance domain at the logic layer, and the complete removal of the AI ​​confidence label at the security adjudication entry point.

[0048] Example 2A: Processing Boundaries of Hierarchical Confirmation Strategy and Confidence Labels In this embodiment, the candidate instruction data packet output by the interactive auxiliary domain includes a confidence level label (value range 0~1). The hierarchical confirmation strategy is executed entirely within the interactive auxiliary domain and the host computer user interface layer. It belongs to the human-computer interaction confirmation link of the auxiliary layer, and its purpose is to verify the clarity of the operator's intention, rather than being part of the safety decision. The specific differentiated strategy is as follows: When the confidence level is ≥0.85, the system only displays a lightweight confirmation prompt on the interface (such as flashing a highlighted candidate instruction, waiting for the operator to look at it for confirmation or lightly touch the confirmation button). If there is no cancellation operation within 1 second, it will automatically proceed to the next step; when the confidence level is 0.60≤confidence level<0.85, the system reads back the instruction content through voice synthesis (such as "Do you want to open the S1 signal machine?"), and the operator replies "confirm" or "yes" before proceeding to the next step; when the confidence level is <0.60, the system refuses to automatically proceed to the subsequent process and forces the operator to manually click on the host computer site map to select the target device and operation (i.e., return to the traditional mouse operation process). The voice content of this instruction is only used as an auxiliary prompt and is not automatically converted into a candidate instruction. It is important to reiterate that any output generated during the aforementioned hierarchical verification process (including successful or failed verification) does not alter the subsequent deterministic security decision made by the secure execution domain. The secure execution domain receives only the standard device opcodes converted by the deterministic mapping engine. The confidence label is physically stripped when the data packet enters the unidirectional data channel (i.e., the confidence field is filtered out by software logic in the data path between the output of the interactive auxiliary domain and the input of the secure execution domain). No submodule within the secure execution domain can obtain the confidence value. Therefore, a clear boundary is formed between hierarchical verification and security decision-making: "intent verification at the auxiliary layer, security decision-making at the deterministic layer." The two do not interfere with each other, fully conforming to the design principles of a confidence isolation architecture.

[0049] Example 3: Simulating Mouse Clicks and Command Sending In this embodiment, after the secure execution domain adjudication is passed, the instruction to be executed is sent to the mouse click simulator (reference numeral 203).

[0050] Click the emulator with the mouse to perform the following operations: Step 1 – Command to Coordinate Mapping: Based on the predefined “Operation Command-Screen Coordinate Mapping Table”, convert the standard equipment operation code into a sequence of screen coordinates on the host computer's site map. For example, the “Open S1 Signal” command is mapped as follows: Coordinate 1 (x1, y1): Command button area (such as the "Signal Control" panel) Coordinate 2 (x2, y2): Device button area (S1 signal light icon) The coordinate mapping table is automatically generated during the initial system installation via "teach mode." The operator clicks the command and device buttons on the station map sequentially as prompted on the screen. The system records the screen coordinates corresponding to each operation and stores them as a mapping table. When the station map version is updated or the screen resolution changes, the system automatically detects the change (by calculating the current station map interface window handle and control tree hash value and comparing it with the baseline hash value recorded during installation). If a change is detected, an automatic recalibration process for the mapping table is triggered. The operator only needs to re-execute the complete teach mode according to the same process to update the mapping table, without requiring manual modification of configuration files or writing of code. All coordinates are stored as screen percentage coordinates (i.e., normalized coordinates relative to the upper left corner of the host computer window) to adapt to different screen resolutions.

[0051] Step 2 – Simulated Click Execution: Call the mouse / touch application programming interface (such as SendInput for Windows, uinput for Linux) of the host computer operating system to generate simulated mouse click operations according to the coordinate sequence. Move the mouse to coordinate 1 → Left-click Move the mouse to coordinate 2 → Left-click Click the confirmation dialog box if necessary. Step 3 - Operation Sending: Simulate a click to trigger the host computer to send the operation command to the interlocking logic unit for execution via existing communication protocols (such as CAN bus, Ethernet UDP / TCP, serial communication, etc.).

[0052] It should be noted that the modules of this invention are isolated from the original software of the interlocking host computer through the operating system process boundary. The system interacts with the host computer by simulating mouse clicks, without modifying any code or configuration files of the original host computer software, ensuring that the system operation does not affect the existing functions of the host computer.

[0053] The core of this embodiment is that the source of the instructions received by the interlocking logic operation unit is completely equivalent to that of manual mouse operation, and it can be compatible with the voice interaction capability of the present invention without any modification.

[0054] Example 3A: Processing of concurrent voice input by multiple operators In this embodiment, the system supports concurrent scenarios where multiple operators simultaneously issue voice commands on the same control panel. The microphone array uses beamforming technology to spatially separate sound sources by azimuth angle, forming independent voice channels for each detected independent sound source direction. Each voice channel corresponds to an independent interactive auxiliary domain processing instance (including liveness detection, speech recognition, and semantic parsing), and each instance is bound to an operator ID (associated with a pre-configured operator seat orientation mapping table via the sound source azimuth angle). When multiple voice channels simultaneously generate candidate command data packets, the secure execution domain sequentially sends each candidate command to the deterministic adjudication pipeline according to its arrival timestamp. The adjudication process follows a single-threaded serial processing principle, ensuring that only one command enters the resource arbitration stage at any given time. If the preceding command is undergoing resource arbitration (occupying an average execution time of approximately 15ms, with a worst-case execution time of 50ms), subsequent commands are queued in a waiting queue. The waiting timeout threshold is set to 500ms. After a timeout, the system issues a voice prompt: "Command congestion, please try again later." The resource arbitration phase employs a physical resource occupancy set check mechanism. If the resource occupancy sets of two commands do not overlap, they are allowed to pass in parallel (e.g., operator A issues "turnout P1 locked," operator B issues "open signal S2," and if the protection routes of P1 and S2 do not overlap, both commands can simultaneously enter the mouse click simulator queue, and the simulator executes the mouse clicks serially according to the coordinate sequence). If the resource occupancy sets overlap, they are handled according to the first-come, first-served principle; the later command is intercepted and triggers a "resource conflict, intercepted" voice alarm. This concurrency mechanism achieves sound source diversion through spatial separation and ensures safe mutual exclusion through deterministic resource arbitration, effectively improving operational efficiency without compromising safety in multi-operator collaborative operation scenarios.

[0055] Example 4: Device Resource Topology Arbitration In this embodiment, the system receives two operation requests for the same device within an 800ms time window: Voice channel request: "Switch P1 locked" Mouse channel request: "Switch P1 unlock" The security execution domain arbitration unit does not determine the semantic security of "locking" and "unlocking," but instead performs the following physical state determination: Query the current real-time physical status bit of turnout P1 (status data transmitted back from the interlocking logic unit via the host computer communication protocol). If the current status is "locked", the "lock" request is determined to be an invalid duplicate operation (and is discarded directly). The "unlock" request has the ability to physically change things, enters the confirmation queue, and forces the mouse to confirm twice. If the current status is "not locked", the "unlock" request is determined to be an invalid operation (and is discarded directly), and the "lock" request enters the execution queue.

[0056] This arbitration mechanism is based entirely on the physical facts of the equipment, eliminating misjudgments of edge conditions caused by static semantic classification.

[0057] Example 5: Unalterable Audit Traceability In this embodiment, each operation record in step S8 includes the following fields:

[0058] The Hardware Security Module (HSM) is deployed as an onboard component of the host computer motherboard. It communicates with the host computer CPU via PCIe or SPI bus and has a built-in independent security processor and non-volatile memory area for storing private keys and generating digital signatures. For host computer hardware configurations without an onboard HSM, an external USB HSM device conforming to the FIPS140-2 Level 3 standard can be used as an alternative. It connects via a USB 3.0 interface and requires a physical key card to be inserted and a PIN code to be entered for activation during operation.

[0059] Calculate the hash of this record: SHA-256 (previous hash + sequence number + timestamp + operator ID + liveness detection result + standard device operation code + simulated click coordinates + execution result). Use a hardware security module (see attached diagram 300) to perform asymmetric encrypted digital signature on the hash value of each record. All records are sequentially appended and stored in read-only solid-state storage (see attached diagram 301).

[0060] During auditing, the hash chain and digital signature are verified one by one, starting from the first record. If the verification of any record fails, the log is determined to have been tampered with. This mechanism can verify 3,000 logs per second, meeting the highest traceability requirements of SIL4.

[0061] Explanation of Functional Safety Compatibility and Isolated Deployment It should be noted that: First, regarding the allocation of safety levels: The rail transit signaling system adopts a hierarchical allocation principle. The interlocking logic unit directly controls trackside equipment such as turnouts and signals, and its safety integrity level is SIL4. The interlocking host computer, as the human-machine interface between the operator and the interlocking logic unit, has a safety integrity level of SIL2. This invention is entirely deployed on the interlocking host computer, where the safety execution domain (deterministic mapping, digital twin pre-verification, and resource arbitration) meets the SIL2 requirements; the interactive auxiliary domain (including the AI ​​module) does not participate in safety adjudication, and its output is only used as auxiliary suggestions, without claiming a SIL level.

[0062] Second, regarding the liveness detection software comparison module: The liveness detection decision is executed by the software comparison decision submodule deployed on the interlocking host computer. This module only performs deterministic numerical comparison logic—comparing the measured phase group delay characteristics with the threshold envelope pre-calculated based on the physical model, and outputting a "pass / reject" binary decision. This process does not involve any neural networks, machine learning, or statistical probability models; it is entirely based on a deterministic physical model and numerical comparison, possessing good interpretability, verifiability, and testability, and meeting the SIL2 functional safety requirements for deterministic logic.

[0063] Third, regarding the secure execution domain software module: all functions of the secure execution domain—deterministic mapping, digital twin spatiotemporal conflict pre-verification, and resource occupancy arbitration—are executed by a software module deployed in the interlocking host computer. This software module runs as a daemon process / background service, and its core algorithms are all deterministic logic (string matching, graph search, numerical comparison, interval algebra calculation), without involving any machine learning or probabilistic inference models, and thus also meet the requirements of SIL2 functional safety for deterministic logic.

[0064] Fourth, regarding the deployment boundaries of the AI ​​modules: All deep learning modules in this invention (such as speech enhancement CRNN and semantic parsing BERT) are deployed in the interaction-assisted domain (speech signal preprocessing and perception layer). Their outputs (confidence labels, enhanced speech signals, and intent classification probability distributions) are physically prohibited from flowing into the safe execution domain. The safe execution domain and the interaction-assisted domain are isolated through a unidirectional data channel; the outputs of the safe execution domain are not returned to the interaction-assisted domain.

[0065] Fifth, regarding isolation from existing systems: Each module of this invention is completely isolated from the original interlocking host computer software through operating system process boundaries. This invention does not modify any code, configuration, or communication protocol of the original interlocking host computer software; it only interacts with the original interlocking host computer software by simulating mouse clicks through the operating system API. When the system is downgraded or malfunctions, the operator can immediately switch back to traditional mouse operation, and the original functions of the interlocking host computer software remain completely unaffected.

[0066] The system architecture design follows the fault-safe principle: When AI inference times out: The interactive auxiliary domain actively terminates the timed-out thread, discards the current frame and releases resources, and the process terminates naturally, waiting for the next wake-up; When a liveness detection result indicates an "attack": the software logic blocks the voice signal path and terminates the process; When the confidence level of the deep learning module is lower than a preset threshold: the interaction auxiliary domain refuses to generate candidate instructions, and the process terminates; When deterministic mapping fails: the safe execution domain terminates the process directly and reports an error; When a potential conflict is detected during digital twin pre-verification: the secure execution domain triggers a deterministic security veto; When the digital twin model and physical device lose synchronization: the watchdog timeout is triggered, and the system degrades to mouse-only operation mode.

[0067] The aforementioned six-level degradation mechanism ensures that while introducing intelligent voice interaction capabilities, the original SIL4 security integrity level of the interlocking logic operation unit is not reduced, and the present invention itself meets the SIL2 security requirements of the host computer, achieving complete isolation from existing systems and minimal intrusive deployment.

Claims

1. An interlocking human-machine collaborative control system based on array sensing and confidence isolation, deployed on an interlocking host computer, characterized in that, include: The interactive assistance domain, including a microphone array interface, a spatial awareness module, and a semantic parsing module, is used to convert operator speech into candidate instruction data packets with confidence labels. The spatial perception module integrates a passive liveness detection unit and a beamforming module; the semantic parsing module is used to extract structured operation instructions from the text. The secure execution domain, deployed inside the host computer, is connected to the interactive auxiliary domain via a one-way data channel. It includes a deterministic mapping engine, a pre-verification unit, and a resource arbitration unit. After receiving the candidate instruction data packet, the secure execution domain discards the confidence labels and makes security decisions based solely on the deterministic physical model and interlocking rules. A graded confirmation unit is used to perform a differentiated confirmation strategy based on the confidence level before the candidate instruction data packet is passed into the secure execution domain; the confidence level label is physically stripped before being passed into the secure execution domain. A mouse click simulator is used to generate corresponding mouse click operation sequences based on arbitrated instructions, simulating the operator clicking instruction buttons and signal device buttons on the host computer's site map; The execution feedback unit is used to send the simulated click operation to the interlocking logic operation unit for execution through the existing communication protocol of the host computer, and to provide synchronous feedback of the execution result through the visual and auditory channels. The log auditing unit, with its built-in hardware security module and read-only solid-state storage, is used to record the entire process operation information using hash chain and hardware security module digital signature, forming an immutable operation traceability chain.

2. The system according to claim 1, characterized in that, The passive liveness detection unit and the beamforming module share the same microphone array data stream, which is copied, split, and processed in parallel through the data link layer. The passive liveness detection unit extracts the phase group delay features between each channel of the array, dynamically calculates the threshold envelope based on the near-field spherical wave propagation physical model and array geometric parameters, and determines whether it is a real person speaking or a replay attack through deterministic numerical comparison. The beamforming module is used to extract the speech signal in the target direction through spatial filtering.

3. The system according to claim 1, characterized in that, The deterministic mapping engine is used to forcibly convert natural language entities into standard equipment operation codes within the interlocking system; the pre-verification unit is used to call the station's digital twin model, use deterministic interval algebra to calculate the inviolable spatiotemporal envelope, and perform security conflict pre-verification on the instructions; the resource arbitration unit is used to arbitrate the instructions based on the physical mutual exclusion principle of the intersection of resource occupancy sets. The deterministic mapping engine receives candidate instructions from the semantic parsing module and performs forced matching in the mapping table in the read-only memory, strictly converting the operator's natural spoken entity words into internal standard device operation codes that are uniquely identifiable by the interlocking system. If the match is successful, a unique machine code is output; if the match fails, the process is terminated. The device identification code mapping table in the read-only memory cannot be rewritten during operation.

4. The system according to claim 3, characterized in that, After receiving the standard equipment operation code, the pre-verification unit discards the confidence label, calls the station digital twin model stored in a graph data structure, and obtains the current status of the equipment, including occupied and idle status. The digital twin model asynchronously updates node attributes through interrupt feedback from the interlocking logic operation unit. Using the current emergency braking distance of the train and the latest response time of the turnout as boundaries, it calculates the inviolable spatiotemporal envelope within the future time window. If the target equipment is located within the spatiotemporal envelope, a safety veto is triggered, the current instruction is intercepted, the process is terminated, and the operator is prompted to manually handle the matter using a mouse through both visual and auditory channels. If the target equipment is located outside the spatiotemporal envelope, it is latched as an instruction to be executed. If the digital twin model of the station does not receive a status update from the lower-level machine for more than a preset synchronization interval, it will be downgraded to a mouse-only operation mode and all voice commands will be prohibited. The resource arbitration unit maps each operation instruction to a request to occupy specific device resources. If the resource occupation sets of multiple requests have an intersection, they are handled according to the physical mutual exclusion principle of first-come, first-served, and subsequent requests wait and are canceled upon timeout. The arbitration does not use the semantic meaning of the instruction text as a judgment weight.

5. The system according to claim 4, characterized in that, The interlocking logic operation unit receives device operation instructions generated by simulated clicks through the existing communication protocol of the host computer. The operation command is subjected to an interlocking logic check, which includes route condition verification, turnout position consistency verification, section occupancy status verification, and opposing route mutual exclusion verification. If the interlocking logic check passes, a control command is output to the corresponding trackside equipment drive unit to drive the turnout switching or the signal lights to turn on. After the trackside equipment drive unit executes the control command, it collects the equipment status and feeds it back to the interlocking logic operation unit; The interlocking logic operation unit will transmit the execution result back to the host computer through the existing communication protocol, so that the execution feedback unit can provide synchronous feedback between the visual and auditory channels. The interlocking logic operation unit operates independently of the interactive auxiliary domain and the security execution domain in the host computer, and its internal interlocking logic does not contain any neural network or machine learning model.

6. The system according to claim 1, characterized in that, The differentiated strategy executed by the hierarchical confirmation unit is as follows: when the confidence level reaches the preset first confidence threshold, a confirmation prompt is displayed on the interface, and if no cancellation operation is performed within a set time, the subsequent process is automatically entered; when the confidence level reaches the preset second confidence threshold but does not reach the preset first confidence threshold, the instruction content is read back through speech synthesis, and the subsequent process is entered after the operator confirms it by voice; when the confidence level does not reach the preset second confidence threshold, the automatic entry into the subsequent process is refused, and the operator is forced to manually click to select the target device and operation on the host computer site map using the mouse; the confirmation is only used to verify the operator's intention and does not change the decision conclusion of the safe execution domain.

7. The system according to claim 1, characterized in that, The mouse click simulator converts the arbitrated instructions into a sequence of screen coordinates on the host computer's site map according to a predefined operation instruction and screen coordinate mapping table. The screen coordinates are stored as screen percentage coordinates. The mouse click simulator calls the mouse / touch application programming interface of the host computer operating system to generate simulated mouse movement and click operations according to the coordinate sequence; the host computer has existing communication protocols including CAN bus, Ethernet UDP / TCP or serial communication.

8. The system according to claim 1, characterized in that, The system is completely isolated from the original software of the interlocking host computer through the operating system process boundary; the system does not modify the code or configuration files of the original software of the interlocking host computer, and only interacts with the software of the interlocking host computer by simulating mouse clicks; when the system is downgraded or fails, the operator can directly switch back to the traditional mouse operation mode.

9. A method for interlocking human-machine collaborative control based on array sensing and confidence isolation, applied to the system as described in any one of claims 1 to 8, characterized in that, include: Multi-channel audio signals are acquired using a microphone array; Based on the audio signal, the target speech is extracted by the beamforming module, and the phase group delay feature is extracted simultaneously. The threshold envelope is dynamically calculated based on the near-field physical model and array parameters for comparison. If it exceeds the threshold, it is determined to be a replay attack and blocked. If it falls within the threshold, it is determined to be a real person's voice and subsequent processing is allowed. The target speech is converted into text and structured operation instructions are extracted to generate candidate instruction data packets; The hierarchical confirmation unit executes a differentiated confirmation strategy based on the confidence level label, and physically strips the label before the data packet is passed into the security execution domain; The deterministic mapping engine forcibly converts the confirmed instructions into standard equipment operation codes within the interlocking system; After receiving the standard device operation code, the security execution domain discards the confidence label, calls the station digital twin model to calculate the inviolable spatiotemporal envelope, and if the target device is located within the envelope, a security veto is triggered and the instruction is intercepted; otherwise, it is latched as an instruction to be executed and enters resource arbitration. The resource arbitration unit maps the instructions to be executed into requests for the use of specific device resources and arbitrates them based on the principle of physical mutual exclusion. The arbitration command is sent to the mouse click simulator to generate a simulated click sequence to click the station map button, and then sent to the interlocking logic operation unit for execution via the existing communication protocol; After the interlocking logic operation unit executes, the result is fed back synchronously through the visual and auditory channels; The hash chain digital signature method based on hardware security modules records the entire process operation information, forming an immutable operation traceability chain.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Voice deception attack detection system and method based on microphone array

    CN114155850B