Method and system for verifying intention with body based on cross-layer feature coupling

By employing a cross-layer feature coupling embodied intent verification method, and utilizing a causal calculus model to identify user autonomous behavior in an encrypted communication environment, this method solves the problems of privacy leakage and real-time intervention in existing technologies. It achieves high-precision intent recognition and low-power security protection, and is suitable for diverse terminal environments.

CN122001669APending Publication Date: 2026-05-08HANZHONG BIG FRUIT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANZHONG BIG FRUIT TECHNOLOGY CO LTD
Filing Date
2026-03-17
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies cannot accurately identify user-initiated behavior from automated attacks, remote control, or malicious internal calls in encrypted communication environments, and pose privacy and legal risks, thus failing to meet the need for real-time intervention.

Method used

The method of embodied intent verification adopts cross-layer feature coupling. By monitoring digital behavior flow and physical interaction data, it uses a causal calculation model to quantify the causal correlation strength between digital behavior and physical entities, and achieves accurate judgment of user autonomous behavior. The steps include S1 digital situation awareness, S2 physical state mapping, S3 causal consistency calculation and S4 dynamic decision intervention.

Benefits of technology

It achieves high-precision intent recognition in encrypted environments with zero privacy leakage, meets real-time intervention requirements, is suitable for diverse terminal environments, and features low power consumption and high flexibility. It is applicable to scenarios such as protection of minors, fraud prevention for the elderly, prevention of misuse of smart devices, and financial fraud prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122001669A_ABST
    Figure CN122001669A_ABST
Patent Text Reader

Abstract

The invention discloses an end side security dynamic management and control method and device based on cross-layer causal verification and terminal equipment, and relates to the technical field of computer security and privacy computing. The method comprises the following steps: asynchronously and parallelly collecting a first feature sequence representing a digital behavior intention of a user and a second feature sequence of physical interaction and environment states coupled with the user in a space-time manner at a local terminal; the causal association strength between the two sequences is calculated in real time through a causal calculation model, whether a digital instruction is driven by real physical interaction or not is quantified, and a dynamic credibility score is generated; and dynamically arbitrating at the terminal locally based on the score, switching to a matched security role operation environment, and executing a fine-grained access control strategy. According to the method, a traditional scheme depending on plaintext content or static rules is abandoned, and a new intention verification normal form of digital-physical causal closed-loop verification in an encryption environment is constructed. All sensitive data are processed on the end side and do not need to be uploaded; the asynchronous trigger mechanism ensures low power consumption and millisecond response. The scheme can be flexibly deployed in heterogeneous terminals such as a mobile phone, the Internet of Things, an intelligent connected automobile and an industrial system, non-personal attacks such as remote control and automatic scripts are effectively detected, and the initiative, accuracy and reliability of end side safety protection are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer security, human-computer interaction and privacy computing technology. Specifically, it relates to a method and system for implementing high-precision intent verification and proactive security protection locally on a smart terminal, which is particularly suitable for distinguishing between user-initiated behavior and automated attacks, remote control or internal malicious calls in an encrypted communication environment. Background Technology

[0002] With the acceleration of digitalization, smart terminals (such as smartphones, smart cars, and industrial control systems) have become core entry points for critical business and sensitive operations. At the same time, security threats targeting these terminals are becoming increasingly complex and covert. The industry urgently needs an innovative technological solution that can accurately and in real-time verify and intervene in high-risk digital behaviors without decrypting communication content or compromising user privacy, all locally on the terminal.

[0003] However, existing technical solutions face two major challenges: First, the widespread adoption of encrypted communication renders traditional detection technologies ineffective. End-to-end encryption protocols such as TLS 1.3 and QUIC have become standards, rendering content analysis and behavior recognition technologies based on deep packet inspection (DPI) completely ineffective. Security systems face the dilemma of "not being able to see the content." Second, privacy regulations and user experience requirements necessitate on-device processing. Globally, regulations such as GDPR and China's Personal Information Protection Law strictly limit the uploading of raw user data (such as touch data, biometric signals, and network traffic) to the cloud for analysis. Furthermore, the latency and power consumption incurred from sending all data to the cloud for analysis severely degrade the user experience.

[0004] Existing technical solutions attempt to address the above problems, but all have fundamental limitations: Cloud-based analytics solutions require uploading raw or lightly anonymized data, posing legal and ethical risks of privacy breaches, and suffer from high response delays, making them unsuitable for real-time intervention. Pure client-side rule engine: It can only make simple allow or block based on coarse-grained information such as application identifier and IP address, and cannot identify specific risky operations within the application (such as distinguishing between normal browsing and fraudulent transfers). It is easy to be bypassed and has a high false alarm rate. Single-dimensional behavioral analysis: This involves analyzing only network traffic statistics (such as time series and packet length) or only user interaction patterns (such as click frequency). The former cannot distinguish between active user behavior and silent background traffic; the latter cannot correlate physical interactions with specific digital business risks. This separation leads to a sharp decline in recognition accuracy in complex scenarios (such as multi-task parallelism and background updates). Content-based guessing "DPI-like" schemes attempt to guess the content of encrypted traffic through machine learning. This method is not only unreliable, but may also cross the legal line of illegally parsing communication content.

[0005] To address the complex challenges of edge security protection, related technological fields have undertaken multi-dimensional exploration and development. For example, existing technologies include content classification methods based on encrypted traffic metadata (such as time-series and SNI), providing network-side feature awareness capabilities for edge behavior recognition; other solutions have further designed a dynamic permission management system that integrates behavior recognition and incentive strategies, demonstrating the application value of recognition results in control; in addition, some studies have proposed a local security role environment switching mechanism for terminals, providing underlying support for fine-grained control; and still others have proposed a flexible analysis framework and cross-device state synchronization protocol to adapt to the capabilities of heterogeneous devices, thereby improving system practicality and coverage.

[0006] However, when applying the aforementioned technologies to high-value, high-security scenarios such as protecting the mental health of minors and preventing cyberbullying, preventing mis-payment by the elderly, preventing mis-control of smart devices, and combating financial fraud, the inventors have discovered a common and urgent new problem: existing solutions can effectively identify 'what the behavior is' and implement control, but none can penetrate the encryption barrier at the moment the behavior occurs to verify whether the behavior is truly driven by the user's real physical interaction. Faced with 'non-embodied' attacks that do not generate or forge physical interaction signals, such as remote control, automated scripts, malicious calls within the system, and even social engineering-induced hasty user operations, existing solutions have detection blind spots. This lack of 'intent authenticity verification' capability has become a key bottleneck in elevating edge security to a new level. This invention aims to overcome this bottleneck. Summary of the Invention

[0007] Technical problems to be solved This invention aims to overcome the shortcomings of the prior art and solve the following core problems: The challenge of verifying intent: How can we accurately determine the true intent behind digital behavior (i.e., whether it is driven by the user) without obtaining plaintext communications or infringing on user privacy? The challenge of real-time system design: How to design a low-power, high-real-time edge system that can accurately extract abnormal behaviors driven by non-user intent (such as Trojans, automated scripts, remote control, and induced operations) from complex mixed traffic and interactive noise? The challenge of a universal framework: How to build a universal framework that can be flexibly deployed in diverse terminal environments, from consumer-grade apps to automotive-grade and industrial control-grade hardware, and cover a wide range of scenarios, from minor protection to financial anti-fraud? Technical solution

[0008] To address the aforementioned technical problems, this invention proposes the core concept of "embodied intent verification through cross-layer feature coupling." The principle is that genuine user-initiated behavior leaves traces with spatiotemporal coupling in both the digital space (generating specific network traffic or system commands) and the physical space (causing changes in the user's biomechanics, physiological rhythms, or environmental state). In contrast, non-initiated behaviors such as automated attacks and remote control create a "logical break" between these two.

[0009] Based on this, the present invention provides the following solution: A method for verifying embodied intent based on cross-layer feature coupling, characterized by the following steps: Step S1 (Digital Situation Awareness): Monitor the digital behavior flow of the target object in real time and extract the first feature sequence representing the behavioral intent and risk level; the digital behavior flow includes network communication data, system call instructions or application layer operation events; Step S2 (Physical State Mapping): In response to the risk triggering condition of the first feature sequence, the physical perception module is asynchronously activated to collect physical environment data or biological interaction data that are associated with the target object in time and space, forming a second feature sequence; Step S3 (Causal Consistency Calculus): The first feature sequence and the second feature sequence are spatiotemporally aligned, and the causal correlation strength between them is calculated using a causal calculus model; the causal correlation strength is used to quantify the degree of consistency between the initiation motivation of digital behavior and the interaction state of physical entities. Step S4 (Dynamic Decision Intervention): Generate verification conclusions based on the strength of the causal relationship; when the strength of the causal relationship indicates a logical break or inconsistency between the digital behavior and the physical entity state, it is determined to be an abnormal intent, and corresponding security intervention strategies are executed.

[0010] Preferably, the first feature sequence in step S1 may include transport layer features such as network traffic timing, packet length, and cryptographic fingerprints, as well as business semantic features characterized by instruction type and API call sequence. Preferably, the second feature sequence in step S2 may include biomechanical features derived from sensors (such as micro-jitter and pressure intensity), physiological rhythm features (such as heartbeat and eye movement), and environmental state features. Preferably, the causal calculus model in step S3 quantifies the correlation strength by calculating the "forced response coefficient" and the "causal break index," thereby distinguishing between user-driven and non-driven behaviors. Preferably, the security intervention strategy in step S4 adopts a tiered execution mechanism, including a progressive response from user prompts to process blocking and then to linked alarms.

[0011] An embodied intent verification system based on cross-layer feature coupling, used to implement the above method, is characterized by comprising: The digital sensing module is used to execute step S1; The physical acquisition and scheduling module is used to execute step S2; The causal calculus engine is used to execute step S3; The safety decision execution module is used to execute step S4.

[0012] The system can be configured to take various forms, such as standalone applications (APP), software development kits (SDK), operating system services, kernel modules, and even standalone security chips, and deployed in various smart terminals.

[0013] To further enhance security, the system's core processing logic and sensitive raw data can be further configured to run in the trusted execution environment (TEE) or hardware security isolation domain of the terminal device, achieving logical isolation from the main operating system and ensuring tamper-proof and privacy isolation. The hardware security isolation domain includes, but is not limited to, Apple's Secure Enclave, Qualcomm's Security Processing Unit (SPU), or a separate security chip (SE). Beneficial effects

[0014] Compared with the prior art, the present invention has the following significant advantages: Zero privacy breaches and legal compliance: The entire process is handled locally on the device. Original sensitive data (such as touch trajectories, original audio, and network packets) is immediately destroyed after feature extraction, outputting only a de-identified "verification result" label. This fundamentally meets the most stringent privacy protection regulations and avoids legal risks. High-precision identification: Through cross-layer (digital-physical) causal correlation analysis, it can effectively distinguish between "traffic generated when the user is operating" and "silent background traffic" or "malicious script traffic". It achieves identification accuracy close to plaintext analysis in an encrypted environment, and the false alarm rate is significantly lower than that of traditional statistical methods. Real-time performance and low power consumption: The system adopts an "asynchronous triggering" mechanism, which monitors digital features with low power consumption under normal conditions and activates high-precision physical sensing only when there is a suspected risk. This results in extremely low resident power consumption of the system while ensuring millisecond-level real-time intervention capability. Deployment flexibility and versatility: The technical solution is not tied to specific hardware or operating system permissions. It can be used as a regular third-party app, integrated as a system-level SDK or kernel module, or embedded in a TEE or security chip to meet the high-security needs of financial, automotive, and industrial control scenarios. It can be widely applied in various fields such as online protection for minors, fraud prevention for the elderly, enterprise data leakage prevention, IoT device security, and intelligent vehicle control verification. Proactive defense: Upgraded from passive feature matching to proactive "intent verification", it can determine and block malicious operations before they cause actual losses (such as fund transfers or data leaks), achieving true proactive security protection. Terminology Definition

[0015] To facilitate understanding of the technical solution of this invention, the key terms used in this specification and claims are defined as follows: 1. Regarding "Embodied" vs. "Non-embodied": "Embodied": In this invention, it refers to a state in which digital behavior or instructions are initiated by a real physical entity (including natural persons, robots, or controlled intelligent devices) through direct physical interaction (such as touch, voice, body movements, changes in physiological signals, etc.). Its core characteristic is that the generation of digital logic is accompanied by observable and quantifiable physical energy consumption or changes in physical state, that is, there is an inevitable causal relationship between "digital and physical". "Non-embodied" refers to digital behaviors or instructions that do not originate from the aforementioned direct physical interactions, but are generated by pure software scripts, automated programs, remote network injection, internal system simulation calls, or virtualization environments. Its core characteristic is that the generation of digital logic lacks corresponding, real-time evidence of physical entity interaction, manifesting as a disconnect between "digital activity" and "physical silence" or "conflict between physical state and digital instruction logic."

[0016] 2. Regarding "Causal Correlation Strength": This refers to a numerical index quantified through a causal calculus model (see Module 130), used to characterize whether there exists a driving and driven relationship between the "first characteristic sequence" (digital behavior flow) and the "second characteristic sequence" (physical perception data) that conforms to physical or physiological laws. When a digital instruction occurs before or simultaneously with a specific physical response in time, and the trends, energy levels, or logical semantics of both conform to the well-known physical / physiological models in the application scenario, it is determined to be of high causal correlation strength and indicates the true intention of "embodiment". When a digital instruction occurs, if the expected physical response is not detected, or if the detected physical response contradicts the instruction logic (e.g., a high-precision mouse click instruction is executed, but no corresponding finger movement or pressure change is detected), it is judged as having low causal correlation strength or causal break, indicating an abnormal intent that is "non-embodied".

[0017] 3. Regarding "Spatio-Temporal Alignment": This refers to the process of mapping "digital behavioral streams" and "physical sensing data" from different sampling frequencies, clock sources, and coordinate systems onto a unified time axis and spatial reference system. Time alignment: Eliminate transmission delays and clock drift between digital logs and sensor data through timestamp interpolation, sliding window matching, or dynamic time warping (DTW) algorithms to ensure that causal analysis is based on the same time slice; Spatial alignment: Mapping the coordinates of physical sensors (such as camera pixel coordinates, accelerometer vector direction) to the target area of ​​digital operations (such as screen click coordinates, virtual spatial position) to verify the spatial logical consistency between the location of the physical action and the object of the digital instruction.

[0018] Those skilled in the art will understand that the above definitions are intended to define the logical boundaries of the technical solutions of the present invention. The specific alignment algorithms, causal model construction methods, and embodied / non-embodied determination thresholds can be conventionally adapted to known technologies in specific application scenarios (such as financial transactions, industrial control, medical surgery, etc.) without departing from the core concept of the present invention. Attached Figure Description

[0019] Figure 1 System overall architecture diagram Figure 2 Overall Flowchart of the Method Figure 3 Schematic diagram of the working principle of the causal calculus engine Figure 4 Dynamic asynchronous acquisition timing diagram Figure 5 Layered Deployment Architecture Diagram Figure 6 Example 1 - A Psychological Crisis Protection Method Based on Embodied Intent Verification System 100 Figure 7 Example 2 - Anti-fraud intervention method based on embodied intent verification system 100 Figure 8 Example 3 - Black Screen Remote Control Defense Method Based on Embodied Intent Verification System 100 Figure 9 Example 4 - Vehicle Network Security Defense Method Based on Embodied Intent Verification System 100 Figure 10 Example 5 - Collaborative Safety Method for Industrial Robots Based on Embodied Intent Verification System 100 Figure 11Example 6 - XR / Metaverse Virtual Asset Security Method Based on Embodied Intent Verification System 100 Figure 12 Example 7 - IIoT Remote Operation Security Method Based on Embodied Intent Verification System 100 Figure 13 Example 8 - AI Countermeasures and Black Market Cleanup Method Based on Embodied Intent Verification System 100 Figure 14 Example 9 - A Dynamic Boundary Control Method for AI Agent Authority Based on Embodied Intent Verification System 100 Figure 15 Example 10 - A Ubiquitous IoT Dual-Mode Trusted Interconnection Method Based on Embodied Intent Verification System 100 Detailed Implementation

[0020] The core principles of the embodied intent verification method and system based on cross-layer feature coupling provided by the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the embodiments described herein are intended to aid in understanding the invention and not to limit its scope. 4.1 System Overall Architecture

[0021] refer to Figure 1 The embodied intent verification system (100) provided by the present invention is integrated into a smart terminal (200). The smart terminal (200) refers to any device with computing, sensing and communication capabilities, such as smartphones, tablets, smart cars, extended reality (XR) devices, service robots or industrial Internet of Things nodes.

[0022] The system (100) includes four core functional modules that work together to form a complete "perception-decision-execution" closed loop: The digital perception module (110) is configured to monitor the digital behavior flow of the target object in real time through the software or hardware interface of the intelligent terminal (200), and extract a first feature sequence representing the behavioral intent and risk level. The digital behavior flow includes, but is not limited to, network communication data, system call instructions, or application layer operation events. It should be noted that the digital behavior flow has specific manifestations in different types of terminals. For example, in intelligent robots, it can be manifested as path planning instructions or joint execution sequences; in XR devices, it can be manifested as the displacement or operation events of the user in the virtual interactive space. The digital perception module (110) can also be configured to perform lightweight local semantic feature extraction on the collected digital behavior flow. For example, it can perform localized keyword hash matching or sentiment analysis on the text data carried in the network traffic load; or perform non-semantic acoustic event detection on the audio stream (such as spectrum recognition of specific emotional tone); or perform anonymized scene classification on the video stream (such as determining whether it is a transaction interface). All semantic analysis is completed locally on the terminal. The original data is not stored or uploaded. Only desensitized feature vectors or risk labels are output as supplementary or enhancing information to the first feature sequence.

[0023] Physical acquisition scheduling module (120): It is configured to asynchronously activate the perception of physical interaction behavior or device environment status in response to the risk triggering condition of the first feature sequence, and collect physical environment data or biological interaction data to form a second feature sequence.

[0024] Causal calculus engine (130): It is configured to spatiotemporally align the first feature sequence with the second feature sequence and calculate the causal correlation strength between them through a causal calculus model. This correlation strength is used to quantify the degree of consistency between the initiation motivation of digital behavior and the interaction state of physical entities.

[0025] Security decision execution module (140): It is configured to generate a verification conclusion based on the strength of the causal relationship, and execute the corresponding security intervention strategy when it is determined to be an abnormal intent.

[0026] In a preferred workflow, the digital sensing module (110) acts as a resident "sentinel," operating continuously with low power consumption. Once it detects suspicious digital behavioral characteristics (the first feature sequence), it sends a trigger signal to the physical acquisition scheduling module (120). The physical acquisition scheduling module (120) then wakes up high-precision sensors to perform short-term intensive acquisition, obtaining the second feature sequence. Both sequences are sent to the causal calculation engine (130) for core analysis, and the resulting verification conclusion ultimately drives the security decision execution module (140) to take corresponding actions. 4.2 Overall Method Flow

[0027] refer to Figure 2The embodied intent verification method based on cross-layer feature coupling provided by this invention includes the following steps: Step S210: Digital Situation Awareness. Real-time monitoring of the digital behavior flow of the target object, and extraction of the first feature sequence. This step corresponds to the function of the digital awareness module (110). The risk triggering condition can be determined based on the first feature sequence itself, such as detecting highly sensitive business semantic features, atypical operation modes, or matching with known risk modes; Step S220: Physical State Mapping. In response to the risk triggering condition hit in step S210, the physical sensing module is asynchronously activated to collect physical environment data or biological interaction data that are spatiotemporally associated with the target object, forming a second feature sequence. This step corresponds to the function of the physical acquisition scheduling module (120). Its "asynchronous activation" mechanism ensures that the system maintains extremely low power consumption during risk-free periods. This mechanism allows high-power physical sensors to be briefly woken up only when necessary, which can significantly reduce the average power consumption of the system and extend the battery life of terminal devices compared with the continuous full-time monitoring scheme. Step S230: Causal Consistency Calculus. The first feature sequence and the second feature sequence are spatiotemporally aligned, and the causal correlation strength between them is calculated using a causal calculus model. This step corresponds to the function of the causal calculus engine (130). The calculus includes analyzing the temporal correlation and consistency between the two sequences, and may quantify the degree of "disruption" in the physical interaction when digital behavior occurs; Step S240: Dynamic Decision Intervention. A verification conclusion is generated based on the causal correlation strength calculated in step S230. When this strength indicates a logical break or inconsistency between digital behavior and the physical entity state, it is determined to be an abnormal intent, and a corresponding security intervention strategy is executed. This step corresponds to the function of the security decision execution module (140). The intervention strategy can be graded according to the risk level, ranging from alerts to forced blocking. 4.3 Working principle of the causal calculus engine

[0028] refer to Figure 3 The working principle of the causal calculation engine (130) is the core of this invention for achieving high-precision intent verification. Its processing can be abstracted into the following stages: Spatiotemporal alignment and feature fusion: Since digital behavioral flows (such as network packets) and physical interaction flows (such as sensor sampling) typically have different timestamps and sampling frequencies, the engine first performs spatiotemporal alignment. This alignment can be achieved through a unified timestamp synchronization mechanism (such as a network time protocol or hardware clock interrupt), or by using a sliding time window to map digital events with different sampling rates and physical signals into the same analysis frame, ensuring that subsequent analysis is performed within the correct causal time window.

[0029] Association Strength Calculation: On the aligned feature sequences, the engine performs causal association analysis. This is not a simple correlation calculation, but rather aims to discover evidence of whether digital events are driven by physical events. The causal calculus model is the algorithmic vehicle for implementing this calculation function, and can be based on predefined rules, statistical models, or lightweight machine learning models (such as pruned neural networks). For example: Calculate the forced response coefficient: Analyze the instantaneous correlation between the rate of change of energy in digital commands (e.g., bursts of download requests) and the rate of change of response in physical interaction signals (e.g., increased touch pressure, torque fluctuations in robot end effectors). A high coefficient may indicate that the interaction is forced to be driven by external rhythms (e.g., induced voice or remote command streams); The causal break index is calculated as follows: When a high-weighted digital behavioral command (such as a large transfer or a critical equipment control command) is detected, the corresponding physical interaction channel is checked. If the second feature sequence exhibits "physical silence" (such as no effective touch, screen off, or no pilot in the control room) or "absence of biological rhythms" (such as no corresponding heartbeat or eye movement feedback), the causal break index is determined to be elevated. It should be noted that when processing biological signals such as physiological rhythms, this invention only extracts anonymized dynamic features such as their temporal fluctuation frequency, without collecting, reconstructing, or comparing any static biometric templates (such as fingerprints, irises, or facial images) that can uniquely identify a specific natural person. This ensures that while verifying intent, the processing activity does not fall within the strictly regulated scope of biometrics, thus guaranteeing the compliance of the technology.

[0030] Optionally, when calculating the correlation strength, the causal calculation engine (130) can receive and fuse the output of the semantic analysis submodule from the digital perception module (110) as an auxiliary input. For example, when a payment behavior is detected accompanied by external inducement audio, if the semantic analysis module simultaneously identifies that the audio has inducement acoustic features of 'high frequency and rapidity', it can provide additional evidence for the determination of the intention of 'induced payment', thereby potentially increasing the confidence of the final verification conclusion. It should be emphasized that this semantic analysis is an optional enhancement function, and the system can independently complete the intention verification through the core cross-layer causal analysis.

[0031] Intent label generation: Based on the calculated causal association strength value, the engine compares it with a preset threshold (such as...). Figure 3 The values ​​Th_high and Th_low shown are compared and mapped to readable intent verification conclusions. For example, when the intensity value is higher than the threshold Th_high, it is determined to be "user-initiated intent"; when it is lower than the threshold Th_low, it is determined to be "non-embodied anomalous intent"; and when it is in between, it is determined to be "suspicious intent, requiring further observation". 4.4 Dynamic Asynchronous Acquisition Mechanism

[0032] refer to Figure 4 This invention achieves extremely low resident power consumption while ensuring real-time response through a dynamic asynchronous acquisition mechanism. The core of this mechanism is "event-driven" rather than "continuous polling": Normal low-power monitoring (Phase T1): For the vast majority of the time, the system only activates the digital sensing module (110) for low-power flow or command monitoring. The physical acquisition scheduling module (120) and its controlled high-power sensors (such as cameras and microphones) are in sleep or low-power state. The system power consumption is extremely low during this phase; Risk Triggering and Asynchronous Wake-up (Stage T2): When the first feature sequence extracted by the digital sensing module (110) hits the preset risk triggering condition (such as detecting a large payment flow of a specific application), it immediately sends a wake-up signal to the physical acquisition scheduling module (120). Figure 4 As shown, time T2 is the critical turning point when physical acquisition is triggered; High-precision short-term acquisition (stage T3): The physical acquisition scheduling module (120) is awakened and precisely schedules the required sensors according to the risk type. High-frequency, high-precision data acquisition is performed within a short window (e.g., 500 milliseconds) to form a high-quality second feature sequence. Subsequently, the sensors enter sleep mode again.

[0033] Closed-loop feedback and adaptation: The results of causal calculation and security intervention can form a feedback loop to optimize the system's adaptive capabilities. For example, if the system confirms a false alarm after multiple interventions, it can automatically increase the risk trigger threshold for that scenario; conversely, if it is confirmed as a real attack but data collection was not initially triggered, it can automatically decrease the threshold or increase the types of data collection sensors, thereby achieving dynamic learning and optimization. 4.5 Layered Deployment Architecture

[0034] refer to Figure 5 The system architecture of this invention possesses high flexibility and scalability, capable of adapting to terminal devices with different security levels and resource constraints. Its core modules can be deployed at different levels, from the application layer to the hardware layer: Application layer deployment (510): All or part of the modules of the system (100) can be encapsulated as an independent third-party application (APP) or software development kit (SDK). In this mode, the system obtains information such as network status and sensor data through the standard application programming interface (API) exposed by the operating system. This is the most accessible and universal method for deployment. System-level / kernel-level deployment (520): To achieve higher data acquisition efficiency and execution privileges, modules of the system (100) (especially the digital perception module and the security decision execution module) can be integrated as operating system services, background processes, or kernel modules. This approach can intercept network data and system calls at a lower level, achieving more powerful intervention capabilities; Hardware security layer deployment (530): For scenarios with extreme security and privacy requirements, such as finance, automotive, and industrial control, the causal computation engine (130) and the raw data processing logic of the second feature sequence can be configured to run in a trusted execution environment or an independent security chip of the smart terminal (200). This environment is isolated from the main operating system hardware, ensuring that the core verification logic and sensitive biometric data are protected from tampering or theft by malicious software, achieving the highest level of "privacy computing"; Cloud-based collaborative expansion (540): In addition, to address scenarios requiring complex model reasoning or group intelligence sharing, some logic of the causal calculation engine (130) (such as large-scale historical baseline comparison and complex threat pattern recognition) can be offloaded to cloud-based collaborative services for execution. The terminal only needs to upload encrypted feature vectors or desensitized intermediate results, and the cloud will return enhanced verification conclusions after completing the calculations, thereby achieving continuous evolution of computing power expansion and defense capabilities. Psychological Crisis Protection Method Based on Embodied Intention Verification System 100

[0035] refer to Figure 6 This embodiment is applied to the field of protecting the digital health of minors, and is particularly suitable for early warning scenarios of psychological crises in adolescents when they encounter emotional setbacks, academic pressure, or cyberbullying. With the widespread use of social media, existing technical solutions mainly rely on delayed manual observation or privacy-infringing cloud-based content review, failing to achieve real-time, accurate, and compliant intervention. The industry urgently needs a technical solution that can accurately identify early signs of psychological crises locally on the device without infringing on privacy.

[0036] This invention provides an innovative approach to addressing the aforementioned pain points through a framework of "embodied intent verification through cross-layer feature coupling." Its core lies in analyzing the causal consistency between "specific digital behavior patterns" and "real-time physiological behavior states" rather than analyzing the content of communication, thereby inferring psychological risk by detecting "disruptions" between the two.

[0037] In a preferred embodiment of the present invention, the digital sensing module (110) adopts a lightweight approach, focusing on monitoring network behavior patterns to construct a first feature sequence: Anomaly Social Pattern Identification: By establishing a system-level Virtual Private Network (VPN) tunnel on the device, network traffic is monitored. Based on the Server Name Indication (SNI) during the TLS handshake phase, continuous encrypted traffic or sudden surges in traffic patterns from social applications (such as WeChat, QQ, and short video platforms) during abnormal periods (such as from 11 PM to 5 AM the next day). Behavioral mutation detection: Detects for operation sequences that are significantly inconsistent with the user's daily habits through the operating system's Accessibility Service or application usage interface, such as frequently performing operations like "deleting friends," "leaving groups," or "blocking contacts" within a short period of time. Risk-related metadata: Combining a locally encrypted security list, the system detects whether network connections point to known domains associated with high-risk online communities or negative content. When this combination of characteristics indicates a potential risk, the system triggers the risk assessment. This approach is entirely based on traffic metadata and system APIs, without touching any plaintext communication content.

[0038] The physical acquisition scheduling module (120) is asynchronously triggered to collect user status through the terminal's standard sensors in a privacy-preserving manner and construct a second feature sequence. All raw sensor data is processed in real time in memory and destroyed immediately, outputting only anonymized feature vectors.

[0039] An adaptive approach is adopted for both static and physiological rhythm perception: In priority mode, if the device supports it, the built-in millimeter-wave radar can be used to monitor whether the user is abnormally still for a long time and extract whether the breathing rate is rapid or irregular (typical characteristics of crying or tension). In general mode, accelerometer and gyroscope data are obtained via the SensorManager API to analyze whether the device is in a stationary state without being held by anyone, or whether there is uncontrollable low-frequency subtle vibration when the user holds it (distinguishing it from normal use). For abnormal interaction behavior detection, touch sensor data is analyzed to identify whether there is "reduced pressure," "rhythm disorder," or prolonged "interaction interruption" (no screen operation). For non-semantic acoustic event detection, the energy spectrum characteristics of ambient sound are analyzed via microphone, and a lightweight on-device model is used to determine whether there are spectral characteristics of specific sound events such as "crying" or "intense arguing," without performing speech recognition or content understanding.

[0040] The causal calculus engine (130) performs spatiotemporal alignment and correlation analysis on the two feature sequences mentioned above, such as... Figure 6 As shown: First, feature alignment and baseline comparison are performed to align the first features, such as "high-frequency social traffic at night" and "sudden deletion of social relationships," with the second features, such as "rapid breathing," "body stillness and trembling," and "interaction interruption," on the timeline. Next, the causal break index is calculated. The core analysis focuses on whether the accompanying physiological signals of high-risk digital behavior deviate significantly from the user's established health baseline model. For example, monitoring teenagers continuously accessing social media apps late at night, accompanied by rapid breathing and prolonged stillness with trembling, reveals a coupling between "online activity" and "offline physiological stress," which is severely inconsistent with the baseline model of a healthy, relaxed state. This results in a low causal association strength (i.e., a high "causal break index"). This indicates that digital behavior may be a passive immersion or avoidance response under emotional distress, which the system identifies as a "potential psychological risk" intention.

[0041] For scenarios requiring ultimate data security, the causal calculation engine (130) and user baseline model can be deployed in a trusted execution environment (TEE) on the terminal. All comparison calculations of sensitive data are completed within the TEE, ensuring that "data does not leave the domain" and providing resistance to side-channel attacks.

[0042] The safety decision execution module (140) implements tiered and flexible intervention strategies based on the verification results, forming a closed loop of "monitoring-early warning-guidance": Primary intervention (low risk): Support resources such as psychological assistance hotlines and contact information for the school's psychological counseling center are displayed at the top of the device interface; Intermediate Intervention (Medium Risk): Temporarily throttle the encrypted conversation streams suspected of causing stress and trigger a local device alert to guide the user to rest. Simultaneously, send a de-identified alert (e.g., "Child may be experiencing low mood; please pay attention") to the parental control app, without including any specific chat content or screenshots. Advanced Intervention (High Risk / Emergency): If an extremely high risk signal is detected, which combines extreme physiological silence (such as prolonged absence of respiratory fluctuations) with hash matching of a local pre-set risk word library (such as a hash value match of self-harm-related words), the system can activate an emergency protocol to notify the guardian and assist in contacting the school's psychological crisis intervention mechanism or professional institutions.

[0043] Compared with the prior art, this embodiment has the following significant technical advantages: Overcoming the conflict between privacy and efficiency: Through cross-layer causal verification of "metadata behavior analysis + anonymous physiological sensing", accurate early warning can be achieved without analyzing content, completely avoiding legal and privacy risks; Achieve early and accurate warnings: Capable of capturing signals that are difficult to detect through traditional manual observation, and that are coupled with early abnormalities in behavior and physiology, which is more timely than relying on post-event content review or intervention after extreme events occur; Lightweight and quick to deploy: The preferred implementation is based on VPN and standard sensor API, without the need for customized hardware or root access, and can be quickly deployed on most smartphones, with extremely high practicality and scalability; It aligns with positive educational principles. The intervention strategies, ranging from resource delivery and flexible speed limits to emergency response, avoid the harsh digital confinement approach and represent an elevation from "behavioral control" to "psychological protection." This approach is easily accepted by teenagers and forms a virtuous cycle of protection.

[0044] This embodiment demonstrates that the core framework of the present invention can be flexibly adapted, which can not only quickly solve the current urgent problem of early warning of psychological crises in minors through a lightweight solution, but also lay the technical evolution path for future integration into a safer TEE architecture. Fraud prevention intervention methods based on the embodied intent verification system 100

[0045] refer to Figure 7 This embodiment is applied to the field of financial anti-fraud, particularly targeting telecommunications and online fraud against the elderly. Existing technical solutions fail against "self-operation" fraud because they cannot penetrate the appearance of "legal authorization" to verify the true cognitive state behind the payment intention. This invention provides a new path to solve this problem through a "cross-layer causal verification" framework. Its core lies in: not relying on semantic content analysis, but verifying whether the payment intention originates from the user's autonomous and calm will by analyzing the causal relationship between the digital characteristics of the payment instruction and the user's real-time physiological, behavioral, and environmental state characteristics. To meet the needs of rapid deployment and broad compatibility, the preferred embodiment based on the standard interface is described first, and then the technical solution is extended based on this.

[0046] 2.1 Inductive Communication Environment and Risk Context Awareness (Lightweight Optimal Solution) In a preferred embodiment of the invention, the digital sensing module (110) employs a lightweight approach, focusing on monitoring network traffic and system status to identify high-risk payment contexts: Payment Intent and Risk Traffic Identification: All outbound network traffic is monitored through a system-level Virtual Private Network (VPN) tunnel established on the device. Using traffic feature analysis technology, based on the plaintext features of the Server Name Indication (SNI) during the TLS handshake phase, combined with traffic metadata (such as timestamps, packet size sequences, and destination IP address geolocation), domains pointing to known payment platforms (such as Alipay and WeChat Pay) and bank online banking are identified and matched. These features constitute the first feature sequence of payment behavior; Enhanced communication environment risks: Based on a local rule base, the aforementioned payment traffic is flagged for risk. For example, it identifies whether payments occur during unusual times (e.g., midnight to 6 AM), whether they point to unfamiliar or high-risk domains, and whether the traffic volume suggests large transfers. Simultaneously, it monitors whether devices are in continuous voice / video calls (by querying call progress and network connection via system API), creating an abnormal contextual feature of both "large payments" and "continuous external communication." To further enhance sensing accuracy and anti-interference capabilities, the digital sensing module can also be configured to perform deeper system state analysis. For example, it can delve into the underlying operating system to obtain the real-time status of audio routing, analyze whether there are abnormal overlaps or forced occupation of audio channels across applications, and accurately identify whether the user is in an abnormal parallel state of "passively listening to external commands" while performing "local digital operations".

[0047] 2.2 User Interaction Status and Stress Response Perception (Based on Standard Sensors) The physical acquisition and scheduling module (120) is asynchronously triggered to acquire the state characteristics of user interaction through the terminal's standard sensor interface. All raw signals are processed into anonymized features in real time on the terminal side. Biomechanical hesitation feature extraction: Gyroscope and accelerometer data are obtained through the SensorManager API provided by the operating system. The vibration variance of the device during payment operations is calculated to quantify states of tension such as "hand tremors" or "device shaking". Simultaneously, through a touch monitoring interface, the dwell time, click frequency, and dispersion of click coordinates before the user clicks the confirmation button are analyzed to identify behavioral patterns of "hesitation" or "repeated confirmation". Macroscopic detection of visual attention distraction: In user-authorized and privacy-enabled mode, the system uses the front-facing camera to determine whether the user's gaze is deviating from key interactive areas of the screen (such as transfer amounts or recipient information) for an extended period, based on facial orientation and the approximate area of ​​the eyeballs. This macroscopic detection does not rely on high-precision eye tracking and aims to identify significant risks of "attention delocalization." Physiological rhythm stress monitoring (adaptive approach): To ensure compatibility with different hardware, a tiered strategy is adopted. Priority is given to monitoring respiratory rhythm using non-contact millimeter-wave radar; if this hardware is unavailable, micro-motion signals are extracted using a filtering algorithm from a high-sensitivity IMU, or (with user authorization) heart rate variability (HRV) is estimated using the on-side photoelectrometry (rPPG) algorithm from the front-facing camera.

[0048] It should be specifically noted that those skilled in the art should understand that the specific sensors mentioned above (gyroscopes, cameras, radar, etc.) are merely examples. Any sensor or combination thereof capable of acquiring user biomechanical characteristics (such as tremors, pressure), visual attention direction, and physiological rhythm signals such as respiratory rate and heart rate variability without contact or with low contact, as well as the corresponding signal processing algorithm, is suitable for constructing the "second feature sequence" described in this invention to quantify the user's cognitive load and stress state.

[0049] 2.3 Causal Calculation and Intent Verification (Lightweight Scoring Model and Security Architecture) The causal calculus engine (130) performs spatiotemporal alignment and correlation analysis on the two feature sequences mentioned above. In a preferred embodiment, the calculation process can be concretized into a lightweight risk scoring model: Feature quantification and alignment: The risk indicators in the first feature sequence (such as "large payments at night", "payments to unfamiliar domains", "payments during continuous calls") and the stress indicators in the second feature sequence (such as "hand tremor index", "duration of operational hesitation", "degree of gaze deviation", "degree of respiratory disorder") are timestamped and quantified into weighted scores respectively. Forced Response and Causal Disruption Analysis: The engine calculates a joint risk score based on preset rules or a lightweight model, assessing the simultaneous occurrence of these features. This score primarily reflects the coupling strength between the external risk context and the user's internal stress state. For example... Figure 7 As shown, when the "large payment instruction for an unfamiliar domain name" and "intense hand tremors and rapid breathing" are closely coupled in time, the system will give an extremely high risk score, which represents the "causal break" between the "payment instruction" and the "autonomous calm state". Intent Determination: This score directly represents the probability that the payment behavior is driven by external inducement and internal stress response. When the score exceeds a high threshold, the system determines that the payment is a high-risk intention of "involuntary payment".

[0050] For scenarios with extreme security and privacy requirements, such as finance and government affairs, this invention can also adopt a layered security architecture. The causal calculation engine (130) and user baseline model can be deployed in a trusted local execution environment (TEE) on the terminal. The ordinary application layer (REE) transmits the de-identified feature vector to the TEE through a secure channel, and the trusted application within the TEE completes the final calculation. Differential privacy noise or homomorphic encryption technology can be introduced to ensure that core logic and sensitive data "do not leave the domain" and enhance the ability to resist side-channel attacks.

[0051] 2.4 Tiered safety intervention strategy The security decision execution module (140) executes precise and flexible intervention strategies based on the verification results, forming a closed loop of decision support: High-risk intervention (score > 0.7): The system can directly discard data packets of high-risk payment requests through the VPN connection management module to achieve logical blocking, or force a pop-up dialog box that requires confirmation by a specific gesture (such as drawing a circle) to break the fraud rhythm, and simultaneously send de-identification alerts to preset emergency contacts; Medium-risk intervention (score 0.4-0.7): Limit the payment process and display a prominent text reminder at the top of the screen to guide users to confirm the recipient's information again; Low risk / no intervention (score < 0.4): Normal clearance, pass without any noticeable impact; For example, in this preferred embodiment, a high-risk intervention is triggered when the risk score is >0.7, and a medium-risk intervention is triggered when the score is between 0.4 and 0.7. Those skilled in the art will understand that the above thresholds can be dynamically adjusted according to different application scenarios (such as tolerance for false alarm rates) or through training with historical data.

[0052] 3. Technical Effects and Advantages Compared with the prior art, this embodiment has the following significant technical advantages: 3.1. Overcoming the challenge of verifying the authenticity of "personal operation" scams. Traditional solutions can only identify abnormal devices or blacklisted accounts, failing to penetrate the apparent legitimacy of "person holding the device and verifying their identity." This invention pioneers a cross-layered causal correlation analysis combining "risk context + behavioral physiological stress," moving beyond reliance on single static features to dynamically capture the logical consistency between payment instructions and the user's real-time cognitive state (such as hand tremors, disordered breathing, and wandering gaze). This "penetrating verification" mechanism effectively identifies involuntary payment intentions under duress or inducement, filling a technological blind spot in the current anti-fraud field.

[0053] 3.2. Extremely high engineering feasibility and compatibility with all terminals. The preferred embodiment of the present invention employs a lightweight, non-intrusive architecture design: No custom hardware required: It is implemented entirely based on existing operating system standard interfaces (such as SensorManager API, VPN service, camera permissions), without the need for root privileges or custom kernels; Rapid deployment: A complete perception loop can be built using general-purpose millimeter-wave radar (if available), IMU inertial sensor and front-facing camera, which can be adapted to most smartphones and tablet terminals on the market; Low resource consumption: The edge feature extraction algorithm has been optimized and is triggered only in payment key frames, which has a negligible impact on device battery life and performance, and has a practical basis for large-scale promotion. 3.3. Native Privacy Compliance and Data Security To address the sensitivity of biometric data, this invention designs a privacy protection mechanism of "end-side closed loop and feature desensitization": Content-undetectable: The entire process only analyzes the rhythm, tone, and routing status of the audio, without recording or analyzing the semantic content of the call, thus completely avoiding the risk of eavesdropping; No image retention: Visual and rPPG signals are used to calculate feature values ​​(such as gaze angle and heart rate) in real time in memory, without storing or uploading any original facial images or video streams; Local decision-making: The core verification logic is completed locally on the device, and sensitive data does not leave the domain, perfectly complying with the Personal Information Protection Law and the data security standards of the financial industry.

[0054] 3.4. Architectural openness and technological evolution capability This invention adopts a layered design approach of "concrete implementation + abstract model," which has strong technical viability. Sensor independence: The core algorithm model is not tied to any specific sensor model. Whether it's a current IMU, a camera, or future high-precision eye trackers, terahertz sensor arrays, or Wi-Fi CSI sensing modules, as long as their output conforms to the data format of "physiological rhythm" or "behavioral characteristics," they can be seamlessly integrated into this system. Those skilled in the art will understand that, in addition to the sensors listed above, any sensor or sensor combination capable of "acquiring user biomechanical characteristics," "monitoring visual attention," or "collecting physiological rhythm signals" can be used in this system. Smooth Evolution: As terminal hardware capabilities improve, the system can dynamically load more sophisticated perception modules (such as upgrading from macroscopic gaze detection to microscopic eye-tracking analysis) without reconstructing the overall architecture. This design ensures that the technology remains advanced and applicable for the next 5-10 years.

[0055] This embodiment clearly demonstrates that the core idea of ​​the "cross-layer causal verification" of the present invention can not only be quickly implemented on current mainstream devices through a lightweight path to solve the urgent pain point of fraud prevention; at the same time, its open architecture design also reserves the potential to evolve to a higher precision and more secure architecture (such as deep integration of TEE), fully meeting the requirements of the technical solution for practicality, creativity and sustainability. Black Screen Remote Control Defense Method Based on Embodied Intent Verification System 100

[0056] 1. Scenario Description and Pain Points of Existing Technologies refer to Figure 8This embodiment is applied to the field of advanced persistent threat (APT) defense for mobile terminals, particularly targeting attacks involving the "malicious abuse of legitimate applications." With the widespread adoption of remote work and online collaboration, various screen sharing and remote control applications (such as TeamViewer and AnyDesk) have been widely installed and granted high privileges. After attackers implant these applications through social engineering or exploit vulnerabilities, they can directly send simulated touch, swipe, and input commands to the controlled terminal from a remote, black-screen state, completing high-risk operations such as transferring money, installing Trojans, and stealing files.

[0057] Existing technical solutions are almost completely ineffective against such attacks: application-layer security software allows them due to their "legitimate authorization"; behavior monitoring cannot distinguish between user operations and anthropomorphic forged commands; and brute-force disabling affects normal functionality. Therefore, the industry urgently needs a technical solution that can penetrate the appearance of "application legitimacy" and verify whether digital commands originate from genuine interactions of current physical device users.

[0058] 2. Technical solution of the present invention This invention provides a fundamental solution to this challenge through a "cross-layer causal verification" framework. Its core insight lies in the fact that no matter how "legitimate" and "human-like" a remote control command may be, it cannot generate matching physical interaction evidence locally on the controlled device. Precise defense can be achieved by detecting the fundamental causal break of "commands without interaction." To adapt to scenarios with different security levels and deployment costs, this invention provides multiple implementation paths, ranging from lightweight standard implementations to deep hardware-level verification.

[0059] 2.1 First Feature Sequence Acquisition (Highly Sensitive Digital Instruction Stream Perception) In this embodiment, the digital sensing module (110) focuses on monitoring highly sensitive command streams that may be initiated by remote control. A preferred implementation is provided for rapid deployment and broad compatibility: Preferred Implementation (Lightweight Awareness Based on Network Traffic and System API): Remote control protocol traffic identification: A Virtual Private Network (VPN) tunnel is established at the device end to monitor all outbound traffic. Deep Packet Inspection (DPI) technology is used to identify and extract characteristic traffic packets of remote desktop protocols (such as RDP and VNC) or specific remote control software. These traffic characteristics (such as specific protocol headers, ports, and packet timing patterns) constitute the core of the first characteristic sequence, indicating the intervention of external control commands. High-privilege simulated input event capture: Monitor simulated touch, key, and gesture events generated by non-physical touch sources through standard accessibility services (AccessibilityService) or the InputManager API provided by the operating system. Pay particular attention to event sequences that bypass the input method system, are injected directly, and are temporally correlated with the aforementioned network traffic; Sensitive context association: Associate detected simulated input events with high-risk application contexts (e.g., the event coordinates happen to fall within the area of ​​the transfer confirmation button in a bank app).

[0060] Deep Defense Implementation Examples (Enhanced Source Tracing and Kernel-Level Awareness): To further enhance the immutability of the defense, the digital sensing module can also be configured to perform deeper source tracing. For example, it can monitor abnormal injection behavior in the system kernel input event buffer and identify instruction sequences generated by non-physical links from the lowest level by comparing the differences in source tags and timestamp granularity between the "software injection event stream" and the "hardware interrupt event stream".

[0061] A high-risk trigger condition is met when the system detects traffic from a suspicious remote protocol and triggers a simulated input command in a highly sensitive context.

[0062] 2.2 Acquisition of the Second Feature Sequence (Silent Verification of Physical Interaction Evidence) The physical acquisition scheduling module (120) is triggered asynchronously. Its core task is to verify whether the corresponding local physical interaction evidence is missing when the digital command occurs. This embodiment adopts a high-confidence verification scheme based on standard interfaces: Screen status and user perception verification: Using the DisplayManager API provided by the operating system, directly query and confirm whether the screen is off, locked, or whether the currently displayed image is a solid color / static wallpaper (without user interaction). This is direct evidence that the user cannot perform visual confirmation or interaction. Ambient light sensor-assisted verification: Monitoring the fluctuations in ambient light sensor (ALS) values ​​when simulated commands are issued. Normal user operation is often accompanied by changes in light perception due to hand shadows or posture changes, while remote black-screen operation is usually accompanied by unusually stable ambient light readings. This characteristic can serve as supplementary evidence for "unmanned close-range operation".

[0063] Device interaction status awareness: Touch channel "silent" verification: Through the standard input subsystem interface, listen for and verify whether there are no real touch events reported from the physical touch sensor near the timestamp of the simulated input event. This forms a key chain of evidence of "software commands but no hardware interaction"; Device posture and holding status awareness: Gyroscope and accelerometer data are obtained through the SensorManager API to determine whether the device is in a stationary state without being held (such as lying flat on a table). Normal user interaction is usually accompanied by slight device movement or angle adjustment.

[0064] (Optional) Biometric Presence Detection: If the device supports it and the user authorizes it, the front-facing camera (infrared) or proximity sensor can be used to quickly detect whether there is a user's face or a living body approaching in front of the screen.

[0065] 2.3 Causal Calculation and Intention Verification The causal calculus engine (130) performs spatiotemporal alignment and correlation analysis on the two feature sequences mentioned above, and its calculus logic is clear and powerful: Evidence chain spatiotemporal alignment: Establish a precise comparison window to align each simulated instruction (and its associated network traffic features) in the first feature sequence with all physical states at the same moment in the second feature sequence; "Causal Discontinuity Index" Calculation: The calculation engine checks whether a set of highly indicative "silent" conditions are met at the time the instruction occurs, such as: Screen status = Off or locked Touch hardware input = None Equipment attitude = stationary Ambient light fluctuations = No abnormalities (auxiliary) (Optional) Biological presence = Not detected Intent Determination: If the aforementioned "silence" conditions are met simultaneously, and the simulated command successfully triggers a high-risk business operation (such as initiating a transfer), the system calculates an extremely high "causal break index." This indicates that there is no reasonable causal relationship between the digital command stream and the current physical state of the device. The system determines this operation as a "non-embodied remote attack," and the verification conclusion is the highest risk level.

[0066] 2.4 Tiered safety intervention The security decision execution module (140) performs immediate and precise intervention based on the verification results: Real-time blocking and alarm: The system can forcibly stop the suspected remote control process the moment it is determined to be an attack, and trigger a conspicuous audible and visual alarm on the device to remind the user that the device may be remotely intruded. Operation Interception and Evidence Collection: Automatically intercepts the high-risk operation (such as canceling a transfer) and generates an encrypted local evidence collection log, recording the attack time, suspicious processes, network connection, and sensor status snapshots for subsequent analysis; Network isolation: Immediately disconnect all network connections of the suspicious process to prevent it from receiving instructions or leaking data.

[0067] 3. Technical Effects and Advantages Overcoming the blind spot of "legitimate application" abuse: By associating network protocol characteristics, system-level simulated events and multi-dimensional physical states, it has for the first time achieved accurate identification of malicious behavior of remote control software with legitimate permissions, solving a major pain point of traditional security systems; Robust defense based on observable facts: The defense logic relies on objective facts that can be verified by standard APIs, such as "remote attacks cannot fake local screen off, physical touch missing, and device inactivity," making it extremely difficult to bypass by software-level spoofing. Lightweight and quick to deploy: The preferred implementation is based entirely on the standard network and sensor APIs exposed by the operating system. No custom kernel or hardware is required, the development cost is low, and it can be quickly deployed on most commercial mobile terminals, which has extremely high practicality and commercial value. Forming a proactive defense closed loop: From threat perception and causal verification to real-time intervention, it provides an end-to-end solution for advanced remote control attacks, significantly shifting the defense node from "post-event tracing" to "in-event blocking".

[0068] This embodiment demonstrates how the "cross-layer causal verification" framework of this invention can be flexibly adapted to different scenarios. It can achieve core defense value through a lightweight solution, quickly forming product capabilities, and also provides a clear technical path for pursuing ultimate security. Its design concept can be directly extended to any scenario facing the risk of "remote unauthorized control," such as IoT devices and industrial control systems. Vehicle Network Security Defense Method Based on Embodied Intent Verification System 100

[0069] refer to Figure 9 This embodiment is applied to the field of intelligent connected vehicle security, aiming to solve the problems of intent verification and security defense in remote vehicle control and autonomous driving takeover scenarios. Existing technologies mainly rely on identity authentication and communication encryption, but cannot verify the legitimacy of control commands, the real-time physical state of the vehicle, or the biological state of the occupants, posing a risk of "legal but dangerous" command execution. For example, a remote start command might be issued while the vehicle is traveling at high speed, or the autonomous driving system might issue a takeover request to a distracted driver. The industry urgently needs a highly reliable causal verification solution that can be flexibly deployed according to vehicle configuration and implements multi-level verification from logical rules to biometric signals.

[0070] This invention proposes an adaptively configurable "multi-level causal verification" framework as a specific application of the core idea of ​​this invention in the field of vehicle networking. Its core lies in dynamically selecting or combining different levels of evidence chains based on vehicle hardware configuration and security level requirements to verify the consistency between digital commands (remote control or system requests) and the state of the physical world. This framework includes the following two optional, complementary implementation levels.

[0071] Level 1: Lightweight and universal verification based on vehicle network behavior, spatiotemporal consistency, and ontology state (preferred embodiment) This tier is designed for all smart cars with basic connectivity capabilities and onboard sensors, eliminating the need for additional high-cost hardware and enabling rapid deployment and widespread coverage. At this level, the digital perception module (110) listens to the command stream of the T-BOX (such as UNLOCK, START) and associates it with the risk tags (such as abnormal login location) issued by the cloud to construct the first feature sequence. The physical acquisition and scheduling module (120) obtains vehicle speed, gear position, and battery voltage through the CAN bus, and obtains vehicle position (GPS), door lock status, and near-range obstacle signals from ultrasonic radar through basic sensors to construct the second feature sequence; The causal calculus engine (130) has a built-in lightweight rule engine that focuses on calculating the "physical logic rules" and the "spatiotemporal consistency index": Physical logic rule judgment: For example, if the command is "unlock" and the vehicle speed is greater than 5km / h, the risk is judged to be extremely high and the command is refused to be executed; if the command is "start" and the battery voltage is lower than 11.8V, the risk is judged to be high and the command is refused to be executed. Spatiotemporal Consistency Verification (Embodied Evidence Chain): The causal calculus engine calculates the spatiotemporal consistency index at the moment the command is issued. For example, when a digital command such as "open the trunk" or "unlock the door" is detected, the system simultaneously detects the location of the owner's UWB key or the Bluetooth signal of the bound mobile phone. If the command originates from a remote cloud, but the local UWB / Bluetooth signal is not within the preset sensing range behind or to the side of the vehicle (i.e., lacking physical proximity evidence), the system will determine that the spatiotemporal consistency is low, automatically raising the verification threshold or directly refusing execution. This mechanism uses UWB and physical location as "embodied" hard evidence, effectively preventing relay attacks and remote malicious manipulation.

[0072] The security decision execution module (140) blocks high-risk commands in real time and sends de-identified alarms to the cloud and the vehicle owner. This level can intercept most logically paradoxical attacks or misoperations using only the vehicle's existing data.

[0073] Second level: High-precision evidence chain verification, proactive control, and privacy compliance integrating in-vehicle biological status (enhanced implementation example) This level is designed for vehicles with higher hardware configurations. Building upon the first level, it introduces a more refined perception of the biological state of the occupants, constructing an irrefutable chain of physical evidence for use in extreme safety scenarios or advanced autonomous driving functions.

[0074] 1. Multimodal biological state perception and privacy compliance design When the first level determines a high-risk scenario or when a scenario requires it (such as an autonomous driving takeover request), the physical acquisition and scheduling module (120) triggers a high-precision sensor to perform short-term, focused perception, enhancing the second feature sequence. It is important to emphasize that throughout all biological state perception processes, the system strictly adheres to the privacy compliance principles of "edge-side processing, feature desensitization, and data not leaving the domain." Anonymization: The DMS camera analyzes anonymized geometric features such as gaze vectors, head Euler angles, and eyelid opening, rather than recognizable facial images; the in-cabin radar detects micro-motion spectra and presence probabilities, rather than human contour images. Local decision-making closed loop: All biometric data is used only within the vehicle's security domain for real-time causal verification, and is never stored in its original form or uploaded to the cloud. Only when a high-level alarm is triggered (such as when a child is forgotten) is a de-identified alarm event sent to the vehicle owner, without any accompanying biometric data. User Authorization and Transparency: Activation of relevant functions requires explicit informed consent from the user. This mechanism fully complies with the stringent requirements for biometric information processing in regulations such as the "Several Provisions on the Management of Automotive Data Security (Trial)" and GDPR.

[0075] Specific perception dimensions include: Driver Status Monitoring (DMS): Using a driver-facing built-in camera (such as one based on 3D ToF technology), the system analyzes the driver's gaze direction, head posture, and eyelid opening and closing in privacy protection mode to determine whether the driver is in a state where they can take over the vehicle. In-vehicle biological presence and vital sign monitoring: Using in-cabin millimeter-wave radar or UWB chip integrated into the ceiling, non-contact monitoring is conducted to detect whether there are any living beings left in the vehicle, and to detect micro-movement characteristics such as breathing and heartbeat, to prevent the tragedy of children being forgotten, and to serve as refutation evidence for "remote start when no one is in the vehicle". Passenger Status and Classification (OMS): Using in-cabin ToF cameras or millimeter-wave radar, passengers are classified (adults / children) and their sitting posture is detected, providing triggering parameters for smart airbags.

[0076] 2. Hierarchical Causal Algorithm Engine: From Rules to Interpretable Reasoning The causal calculus engine (130) can adopt different levels of implementation modes according to vehicle configuration and scene complexity, realizing the evolution from static rules to dynamic intelligence: Mode 1: Rule Engine Mode: As mentioned earlier, it performs fast Boolean judgment based on preset physical logic rules and spatiotemporal consistency index, which is suitable for blocking scenarios with extremely high real-time requirements (millisecond-level response). Mode 2: Causal Graph Reasoning Mode: For more complex scenarios (such as the coupling of driver distraction and road risk during autonomous driving takeover), the engine can construct and maintain a lightweight scenario-behavior-risk causal graph. This graph uses vehicle state (e.g., speed, position), environmental perception (e.g., obstacle distance), and driver biological state (e.g., line-of-sight deviation) as "cause" nodes, and "high-risk takeover request" or "potential collision violation" as "effect" nodes. By calculating the causal effect of input features on the "effect" nodes in real time, the engine can not only output risk judgments but also provide explainable attributions (e.g., "In this high-risk judgment, 73% was caused by the driver's continuous line-of-sight deviation, and 27% was caused by the sudden deceleration of the vehicle in front"). This provides valuable evidence for subsequent intervention strategy optimization, root cause analysis, and model iteration. Mode 3: Adaptive Learning Mode: Under the premise of ensuring data security, the engine can continuously learn and fine-tune the correlation strength in the causal graph based on a large amount of anonymized local driving data, so that the risk model can adapt to the user's personalized driving habits and local road environment, and achieve self-evolution of safety protection capabilities.

[0077] 3. Active Control and Minimum Risk Strategy (MRM) The security decision execution module (140) at this level not only executes alarms but also possesses substantial proactive control: Minimum Risk Maneuver (MRM) linkage: When a causal break is detected (e.g., the system requests takeover but the DMS detects the driver is unconscious or severely distracted), the safety decision execution module directly links with the vehicle chassis domain controller to trigger the Minimum Risk Maneuver (MRM). Specific measures include: automatically taking over throttle and brake control to linearly decelerate the vehicle, activating hazard lights, and, under safe conditions, automatically controlling the steering system to pull over to the side of the road until the vehicle comes to a complete stop and the doors are unlocked. This demonstrates that the present invention is not merely a perception system, but an active safety system with actual vehicle control. Multi-level sound, light and interactive intervention: trigger seat vibration reminders, adjust the HUD display to show emergency information, or link with the cloud security center to activate emergency rescue protocols.

[0078] Third Tier: Automotive-grade cross-domain isolation deployment architecture (security hardening) In automotive-grade deployments, the modules of this invention operate in different functional domains to achieve cross-domain isolation causal verification, ensuring that a breach in a single domain does not affect overall security. The digital sensing module operates in the network domain (Infotainment / T-Box) and is responsible for receiving external commands and cloud data. The physical data acquisition and scheduling module is distributed in the power domain (Drive-train, to acquire vehicle speed / torque) and the body domain (Body, to acquire door lock / radar data). The core logic of the causal calculation engine and the security decision execution module runs in an independent security domain (SecurityGateway or dedicated Domain Controller). The security domain, acting as a root of trust, does not directly trust the original commands from the connectivity domain. Instead, it performs cross-domain verification by comparing the physical feedback (such as actual vehicle speed) from the power domain with the command intent from the connectivity domain. Commands are only allowed to be issued to the actuators when the cross-domain data logic is consistent. This architecture effectively prevents the risk of vehicle loss of control due to hacking of the entertainment system and complies with the ISO 21434 cybersecurity standard.

[0079] Summary of technical effects and advantages It achieves flexible coverage of "inclusive safety" and "ultimate safety": through layered design, it can provide basic logical defense at low cost, and provide ultimate verification that integrates biosignals for high-end models; A three-dimensional evidence chain of "digital-spatiotemporal-biological" was constructed: by introducing a spatiotemporal consistency index and biological state monitoring, the verification dimension was expanded from a single logical rule to physical location and life state, truly realizing four-dimensional consistency of "human-vehicle-instruction-environment"; It possesses substantial active control and explainability: By executing the MRM strategy in conjunction with the chassis domain, this invention can directly take over vehicle control at critical moments; at the same time, the cause-effect graph reasoning mode provides explainable attribution for risk assessment, enabling the system to move from a "black box" to a "white box", significantly improving user trust and accident analysis efficiency. Meets automotive-grade safety architecture and privacy compliance requirements: The cross-domain isolation deployment solution meets functional safety (ISO26262) requirements; the "edge-side anonymization" privacy design fully complies with domestic and international data regulations, turning compliance into the core competitiveness of the product; This invention defines an industry integration path for next-generation active safety systems: Instead of requiring automakers to install entirely new and expensive sensors, it integrates and enhances the functionality and safety of rapidly expanding UWB digital keys, in-cabin millimeter-wave radar, and DMS / OMS systems. It transforms UWB, originally used for comfort and welcome, into a spatiotemporal verification tool, and DMS, used for fatigue alerts, into a core component for takeover capability assessment, turning these "value-added features" into "essential features" related to life safety. By proposing a cross-domain linkage architecture of "perception (multi-domain) - decision-making (safety domain) - execution (chassis domain)," this invention provides a clear engineering implementation path for intelligent vehicles to move from "assisted safety" to "trustworthy active safety," and is expected to become a standard safety middleware for future intelligent connected vehicles. Example 5: Collaborative Safety Method for Industrial Robots Based on Embodied Intent Verification System 100

[0080] 1. Scenario Description and Pain Points of Existing Technologies refer to Figure 10 This embodiment is applied to the fields of intelligent manufacturing and industrial automation, particularly in human-robot collaboration (HRC) scenarios. With the popularization of flexible manufacturing and the concept of "human-robot integration," industrial robots are moving from "isolated islands" behind safety fences to "partners" working alongside humans. However, existing technological solutions face fundamental challenges in achieving efficient and safe collaboration: The rigidity and inefficiency of traditional security solutions: Solutions based on physical fences, light curtains or area scanning lasers follow the binary logic of "stop when intrusion", which leads to frequent interruptions in production cycles and makes it impossible to achieve true continuous collaboration. The limitations of the passive response of existing collaborative robots: the mainstream solutions mainly rely on body force / torque sensors to achieve "tactile" safety, which is a "in-process" injury mitigation rather than "pre-event" risk prevention, and is difficult to apply directly to unmodified high-speed traditional industrial robots. The lack of intent understanding leads to clumsy interaction: existing systems lack understanding of the work intent and cognitive state of human operators, and cannot distinguish between "planned collaboration" and "accidental intrusion", resulting in the system either being too sensitive and frequently shutting down, or forcing humans to work at an extremely slow speed. Therefore, the industry urgently needs a new generation of human-machine collaborative safety system that can understand human intentions in real time, dynamically assess risks, and implement adaptive safety responses based on task context.

[0081] 2. Technical solution of the present invention: An adaptive security architecture based on causal verification of "task-behavior-state". This invention proposes a hierarchical, configurable framework for "human-machine collaborative causal verification." Its core lies in not simply viewing humans as "dynamic obstacles" to be avoided, but rather as active participants in the collaborative task process. The system dynamically adjusts the robot's safety strategy by verifying the consistency between digital task instructions (MES / PLC), physical actions (vision / positioning), and physiological attentional states (biosensing).

[0082] The framework includes two optional, complementary implementation layers, supplemented by closed-loop verification and digital twin mechanisms: 2.1 First Level: Lightweight Dynamic Risk Field Based on Task Flow and Behavior Prediction (Preferred Implementation) This level is geared towards most existing industrial sites, achieving significant improvements in safety performance through enhancements to existing automation systems and low-cost sensors.

[0083] First characteristic sequence (task context awareness and deterministic communication): The digital sensing module (110) obtains the current collaborative task state machine from the manufacturing execution system (MES) or programmable logic controller (PLC) in real time via industrial Ethernet (such as Profinet, EtherCAT) or OPC UA protocol.

[0084] Example The task sequence is: "Worker installs part A (state S1) -> Robot tightens bolts (state S2) -> Worker installs part B (state S3)".

[0085] Industrial-grade reliability assurance: The digital sensing module is also configured to monitor the PLC's safety heartbeat signal. The system continuously verifies the communication latency between the intent verification engine and the underlying actuators, ensuring it remains within a defined millisecond range (deterministic latency, typically <10ms). If the heartbeat is lost or the latency exceeds the threshold, the system immediately determines the communication link is unreliable and forcibly triggers a safety shutdown. This mechanism ensures that this solution is not just ordinary offline AI vision analysis, but a real-time control system that meets industrial functional safety levels (such as SIL 2 / PL d) requirements.

[0086] The system has a clear understanding of the theoretically permissible positions and expected behaviors of humans in a specific state, and constructs a baseline for task intent.

[0087] Second feature sequence (personnel behavior and spatial perception): The physical acquisition and scheduling module (120) uses an industrial-grade RGB-D camera or UWB high-precision positioning tag deployed on the top of the workstation to track the operator's whole body skeletal joints, movement trajectory, velocity vector and orientation in real time, and specifically captures the precise relative pose of the hand and the workpiece.

[0088] Causal calculus and dynamic risk field generation: The causal calculus engine (130) integrates task context and human behavior in real time. Intent-based causal classification: Determine whether the operator's current behavior conforms to the expected causal logic of the task (positive causality vs. negative causality). Dynamic risk field modeling: Generate a continuously changing scalar risk field whose intensity is positively correlated with "intended risk level", "approach speed" and "robot kinetic energy"; Safety-rated monitoring output: Maps the risk field intensity to standard safety control parameters (Vmax). Vmax ,Dmin Dmin ), and send it to the robot controller via the security bus; Safety Decision and Execution: The safety decision and execution module (140) adjusts the robot's motion planning in real time according to the intensity of the risk field, so as to realize low-speed cooperation, adaptive deceleration and detour or virtual compliant shell protection.

[0089] 2.2 Second Level: Enhanced Collaborative Authentication Integrating Biological State with High-Precision Force and Touch Sensing (Enhanced Example) This level is designed for scenarios with extreme requirements for security and seamless collaboration. Building upon the first level, it introduces the perception of the operator's biological state and high-bandwidth force-haptic interaction.

[0090] Second feature sequence enhancement (multimodal biological state and interaction force perception): Visual attention monitoring: analyzing gaze focus and fatigue level; Electromyography (sEMG) intent pre-reading: Identifying subtle muscle activation patterns and predicting intent 50-100ms in advance; High-bandwidth six-dimensional force / torque sensing: analyzes and guides force vectors, distinguishing between "intentional teaching" and "unintentional collision"; Enhanced Causal Calculus (Cross-Layer Feature Coupling and Logical Verification): The calculus engine integrates task, behavior, biological state, and force signals to perform high-order causal reasoning, and specifically introduces a cross-layer feature coupling determination mechanism: High Confidence Intent Determination: The causal calculation engine compares the activation patterns of sEMG signals with the visually tracked hand motion vectors in real time. If the two show a highly consistent advance compensation relationship (i.e., the electromyographic signal appears before the physical displacement, and the direction vectors match), it is determined to be a "High Confidence Intent," and the system allows the robot to seamlessly hand over at the optimal cooperative speed. Logical conflict detection and safety degradation: If a logical conflict occurs (e.g., sEMG shows muscle contraction in preparation for grasping, but visual information shows the hand rapidly retracting or deviating from the target; or the gaze is unfocused but the hand has already entered the danger zone), the engine determines it as "Ambiguous Intent" or "Potential Misoperation." In this case, the system does not wait for a collision to occur but automatically triggers safety degradation (e.g., immediately entering zero-force control mode or pausing the action) until the conflict is resolved. This is a typical manifestation of "cross-layer feature coupling" in industrial scenarios, effectively preventing accidents caused by false alarms from a single sensor. Active compliant guidance: Based on the arbitration result, dynamically switch between zero-force control following or rigid avoidance modes; Enhanced security decision-making: including AR-guided augmented reality and dynamic dominance arbitration.

[0091] 2.3 Real-time Causal Security Loop and Dynamic Verification Mechanism To ensure the absolute reliability of the security strategy, the security decision execution module (140) and the causal calculation engine (130) together form a millisecond-level closed-loop causal security loop of "perception-decision-execution-verification".

[0092] Execution effect verification: When the system commands the robot to decelerate or change direction based on the risk field, the physical acquisition and scheduling module (120) will immediately collect actual feedback through vision and force sensors to verify whether "deceleration reaches the expected curve" and "the relative speed of personnel decreases accordingly". The self-proving safety chain: The execution effect verification mechanism constitutes a self-proving safety chain. It not only monitors external risks (such as personnel intrusion) but also monitors in real time whether the system's own defensive actions have produced the expected effect. If, after issuing a deceleration command, the actual speed does not decrease due to excessive robot load inertia or driver failure, the system will immediately determine it as an "internal execution failure," thus skipping the conventional degradation process and directly triggering the highest level of emergency stop. This mechanism achieves proactive defense against internal failures in complex systems, significantly improving the system's robustness. Dynamic upgrade mechanism: If verification fails, a Category 0 Stop or aggressive avoidance strategy will be triggered instantly.

[0093] 2.4 Offline Pre-simulation and Full Lifecycle Optimization Based on Digital Twin It supports deep integration with factory-level digital twin systems, enabling offline rule verification, online backtracking, and new task rehearsal to ensure that the physical system has initial security upon going online.

[0094] 3. Summary of Technical Effects and Advantages

[0095] It has achieved a paradigm shift from "isolation and protection" to "intelligent collaboration": significantly improving overall equipment efficiency (OEE); A three-tiered causal verification system of "task-behavior-state" was constructed to accurately distinguish between "planned collaboration" and "unexpected dangers," thereby reducing unnecessary downtime. It provides a customizable and scalable layered solution that caters to both traditional upgrades and high-end precision collaboration needs; Deeply compliant with and intelligently enhanced international security standards: Industrial-grade deterministic assurance: Through safety heartbeat and deterministic latency monitoring, we ensure that AI algorithms are perfectly integrated into the industrial functional safety system; Contextualized application of SSM and PFL: Achieving dynamic optimal safety distance and force limitation; Empowering manufacturing system-level productivity and resilience: enhancing system-level OEE and improving production flexibility and workforce resilience; A unique cross-layer mutual verification and self-verification mechanism: By verifying the relationship between sEMG and vision in advance compensation, the pain point of single modality being easily interfered with is solved, and the intention recognition with extremely high confidence is achieved. By using a self-certifying security chain, security monitoring is extended from "external" to "internal," ensuring the system's fail-safe characteristics under extreme operating conditions. This embodiment demonstrates that the core idea of ​​"cross-layer causal verification" of the present invention has strong universality in the field, providing a brand-new technical path for the safety upgrade and efficiency revolution of intelligent manufacturing.

[0096] Example 6: XR / Metaverse Virtual Asset Security Method Based on Embodied Intent Verification System 100 1. Scenario Description and Pain Points of Existing Technologies

[0097] refer to Figure 11 This embodiment applies to the extended reality (XR) and metaverse industries, aiming to address the trust crisis regarding the "authenticity" or "anchoring" of digital assets. As XR hardware evolves from immersive virtual reality (VR) to augmented reality (AR) and mixed reality (MR), digital assets are no longer confined to closed virtual worlds but increasingly require strong causal connections with physical entities, real rights, or specific spatiotemporal contexts. Examples include: a digital art NFT generated from a famous physical painting; a digital twin representing the state of a real machine tool in an industrial metaverse; and a limited-edition virtual garment displayed in AR try-on, corresponding to a unique physical inventory.

[0098] However, existing technical solutions have fundamental flaws in ensuring this "virtual-real consistency": Static binding and susceptibility to tampering: Current mainstream binding methods (such as storing a URL pointing to entity information in NFT metadata) are static and centralized. Metadata can be arbitrarily modified by the project owner, and URLs may become invalid or point to counterfeit content, causing virtual assets to "decouple" from their entity anchor and lose their value to zero. Lack of dynamic state synchronization verification: For digital twin assets, existing solutions lack a continuous and verifiable synchronization mechanism for the real-time state of the physical entity. Users cannot trust whether the data displayed by the virtual model (such as machine tool speed and temperature) truly reflects the actual situation in the factory workshop, posing a risk of fraud or misleading information. Insufficient precision and persistence in virtual-real spatial alignment: In AR / MR scenarios, the stable and accurate anchoring of virtual assets in physical space (such as placing virtual furniture in a real living room) is a critical requirement. Existing manual or semi-automatic alignment methods are cumbersome, have limited accuracy, and are prone to inaccuracy after device movement or environmental changes, requiring repeated adjustments and compromising immersion and usability. More importantly, existing methods are extremely sensitive to changes in lighting. **Lighting inconsistency** caused by early morning, noon, dusk, or switching indoor lights on and off can severely interfere with visual feature-based spatial anchoring algorithms, causing virtual assets to "flicker" or "drift." Cross-platform / cross-device inconsistency: The same virtual asset may present different positions, shapes, or even functions on different XR devices (such as AR glasses of different brands) or different metaverse platforms due to differences in coordinate systems and rendering engines, resulting in a fragmented user experience and dilution of asset value; Blockchain performance and real-time bottlenecks: Uploading all high-frequency virtual-real consistency verification data to the blockchain will face the challenges of low throughput and high confirmation latency of traditional blockchains, which cannot meet the sub-second response requirements of real-time interactive applications (such as games and industrial control). Therefore, the industry urgently needs a technical system that can establish and continuously verify a dynamic, tamper-resistant, verifiable causal relationship between virtual assets and the physical world (or the real rights they represent), and can overcome lighting interference and performance bottlenecks.

[0099] 2. Technical solution of the present invention: a multi-layered causal verification chain based on "physical anchor point - spatiotemporal state - asset behavior".

[0100] This invention proposes a "virtual-real consistency causal verification" framework built on a high-performance blockchain and edge computing convergence architecture. Its core lies in creating an immutable "causal umbilical cord" for each high-value virtual asset, dynamically binding it to one or more "anchor points" in the physical world, and maintaining its credible value by continuously verifying the logical consistency between the "digital behavior of the asset" and the "physical state of the anchor point."

[0101] 2.1 First level: Static consistency verification based on physical digital fingerprint and lightweight spatiotemporal anchoring (preferred embodiment) This tier is geared towards most consumer-grade XR / metaverse applications and aims to provide virtual assets with basic authenticity endorsement and spatial anchoring capabilities at a lower cost.

[0102] First feature sequence (asset digital fingerprint and declaration): When an asset is created, the digital perception module (110) captures its core features to generate an initial digital fingerprint (such as hash and texture feature code based on 3D model grid), and stores it on the chain together with the physical anchor description declared by the asset creator (such as "corresponding to a painting in the Louvre" or "anchored to a building in Lujiazui, Shanghai"), forming the starting point of the consistency declaration.

[0103] Second feature sequence (lightweight acquisition of physical anchor points and spatiotemporal stamp): Geospatial Anchoring: When a user needs to place the asset in AR, the system uses GPS and visual SLAM (Simultaneous Localization and Mapping) technology from the mobile phone or XR device to obtain the precise geographic coordinates (latitude, longitude, and altitude) and initial visual feature point cloud of the asset placement point. This data serves as the fingerprint of the "spatial anchor point," which is then linked to the asset's digital fingerprint and uploaded to the blockchain. Entity Feature Anchoring: For assets tied to physical items (such as digital sneakers), users are required to use their device's camera to scan the unique micro-features of the physical item (such as texture at a specific angle, wear marks, or an ID with a built-in NFC chip). These feature hashes are then linked to the asset on the blockchain. Causal calculus and consistency verification: The causal calculus engine (130) is deployed as an on-chain or trusted edge service; Static binding verification: When other users examine the asset, the engine retrieves the "physical anchor fingerprint" stored on the blockchain and provides verification tools. For example, it guides users to the same geographical coordinates to verify whether the visual feature point cloud of the current environment matches the on-chain record; or it scans physical items to verify whether the micro-feature hash is consistent. Robust Illumination Processing and Causal Mapping: To address the fundamental interference of illumination inconsistencies on visual anchoring, the causal computation engine integrates an ambient illumination perception and causal feature normalization module. This module acquires scene illumination intensity and color temperature in real time through the ambient light sensor or camera built into the XR device. Before performing visual feature point cloud matching, the engine first performs "photometric normalization" on the currently acquired image. To adapt to the limited computing power of mobile XR devices (such as those using Qualcomm AR series chips), this module can be implemented using a pruned and quantized lightweight neural network model (such as a MobileNet variant) or a fast mapping algorithm based on lookup tables (LUTs), ensuring stable operation under low power consumption and high frame rates. This process maps the image from the specific illumination conditions at the time of acquisition to a standard illumination reference model. This normalization process is not simply image enhancement; its core lies in establishing a deterministic and reversible causal mapping function between 'specific ambient lighting input' and 'standard reference feature output', thereby forcibly decoupling 'input lighting perturbation' and 'output feature stability' at the algorithm level. Without this causal mapping step, random changes in ambient lighting will directly lead to nondeterministic drift in feature extraction results, causing subsequent feature matching to inevitably fail due to the lack of a stable causal premise, i.e., triggering "perceptual causal break". By forcibly establishing this causal mapping, the system ensures the robustness and accuracy of spatial anchoring under all weather and lighting conditions. Lightweight and persistent spatial alignment: Utilizing improved visual markers or environmental feature points for assisted tracking. During initial alignment, the system records not only the marker pose but also key natural feature points of the surrounding environment after illumination normalization. When the device moves and re-enters the scene, the engine quickly matches these natural feature points, combining them with marker information to achieve fast and accurate virtual-real realignment, significantly improving the user experience. Spatial Persistence and Environmental Adaptation Mechanism: To address natural or man-made changes in the physical environment over time (such as furniture movement or renovation), the causal computation engine also includes a decay management and rebalancing mechanism for environmental feature fingerprints. The engine periodically compares the current environmental point cloud with the initial point cloud stored on the blockchain. If a significant but non-destructive change in the environment is detected, the system automatically triggers a "causal rebalancing" process. This process smoothly reduces the weight of expired feature points and introduces new stable feature points, dynamically updating the reference frame of spatial anchors without interrupting the user experience. When extreme environmental changes (such as room renovation) cause the feature matching rate to fall below a preset threshold, the system automatically switches to 'reconstruction mode,' guiding the user to perform a brief active environmental scan to reconstruct the local map. During this period, virtual assets temporarily enter an 'unanchored' safe state (such as semi-transparent display or paused rendering) to prevent erroneous rendering due to inaccurate coordinates, thereby maintaining the spatial persistence of virtual assets and system robustness in changing environments.

[0104] 2.2 Second Level: Dynamic Consistency Verification Integrating High-Precision Digital Twins and Real-Time State Synchronization (Enhanced Implementation) This level is designed for scenarios with extremely high requirements for dynamic consistency, such as industrial metaverses and high-end digital collectibles. Building upon the first level, it introduces continuous perception and two-way interactive verification of physical entities.

[0105] Second feature sequence enhancement (real-time state awareness stream of physical entities): High-precision 3D reconstruction and continuous scanning: For digital twin assets, high-precision real-time point cloud maps are continuously generated using LiDAR and industrial camera arrays deployed around physical entities (such as machine tools and buildings). This point cloud data serves as a "snapshot" of the physical entity's current form. Internet of Things (IoT) data stream: A sensor network that connects to physical entities (such as temperature, pressure, and speed sensors) to obtain their real-time operating parameter data stream; Proactive Trusted Anchors and Hardware-Level Data Sources: Embedding tamper-proof Hardware Security Modules (HSMs) on critical physical entities.

[0106] The HSM module incorporates a Physically Unclonable Function (PUF) to generate a device-unique, uncopyable encrypted private key. Physical state data (such as temperature and vibration) is immediately timestamped by the HSM with hardware-level, nanosecond precision upon sensor acquisition and digitally signed using the PUF private key. To address the potential signature latency challenges of high-frequency data streams (such as vibration sensors >1kHz), the system employs a hybrid authentication mode: the HSM distributes short-term session symmetric keys for calculating Fast Message Authentication Codes (MACs) on high-frequency data; simultaneously, the HSM periodically (e.g., every second) performs asymmetric signing and on-chaining of the session key and a data digest over a period of time. This ensures that every raw data stream entering the causal calculus engine possesses immutable and non-repudiable trust attributes at its physical source, balancing security and real-time requirements.

[0107] Enhanced causal calculation (dynamic consistency reasoning and bias detection): Fully Automated Virtual-Real Space Registration and Alignment: The causal engine employs a fully automated point cloud registration algorithm. It converts the 3D mesh model of virtual assets into a model point cloud and automatically registers it with a real-time point cloud map signed by HSM, collected from the field. Through algorithms such as multiple hypothetical initial poses and weighted iterative nearest point (ICP), it automatically calculates the optimal transformation matrix, achieving millimeter-level accuracy in virtual-real space alignment without manual intervention. State Synchronization Causal Chain: The engine establishes a real-time mapping relationship between physical sensor data streams (especially HSM signature data) and the visualized state / parameters of virtual assets. For example, physical speed sensor data → virtual model turbine rotation animation speed; physical temperature data → virtual model heatmap color change. Any abnormal interruption or logical contradiction in the mapping relationship (such as a sensor reporting a shutdown but the virtual model still running) will be detected immediately as a "causal break." Reverse control verification and model-based intent-causality verification: When a control command is issued through the virtual asset interface (such as increasing the preset speed of a digital twin machine tool), the command must first pass the intent-safety causality verification of the causal engine. Instead of a simple threshold comparison, the engine invokes a high-fidelity dynamic prediction model corresponding to the physical entity—a parameterized and configurable digital twin whose physical parameters (such as mass, inertia, and friction coefficient) are consistent with the entity's factory calibration data or real-time system identification results—to simulate the complete state trajectory of the physical entity after command execution in the virtual domain at millisecond levels. Subsequently, the engine performs a rigorous spatiotemporal comparison of this predicted trajectory with a predefined multidimensional safe operating envelope. Only when the simulated predicted state trajectory is logically and numerically completely contained within the safety envelope (i.e., a positive causal chain of 'control command → safe state' exists) does the causal engine determine that the command has 'execution consistency'. If any part of the predicted trajectory exceeds the safety envelope, the engine immediately determines "command consistency failure" and blocks the transmission of the command to the physical end. Only commands that pass verification will form a closed-loop verification.

[0108] Enhanced security decision-making (consistent arbitration and asset status management): Dynamic Reputation Score: The system maintains a dynamically consistent reputation score for each virtual asset based on the historical frequency and severity of its "causal break." Low-scoring assets will face display restrictions or value discounts in the market. Cross-Platform Consistency Arbitration Service: This invention framework serves as a decentralized consistency arbitration layer. When different metaverse platforms access the same virtual asset, they can request the current authoritative "physical anchor state" and "spatial alignment parameters" from this arbitration layer. To address the issue of inconsistent coordinate system definitions (Y-up vs. Z-up) and origins across different XR device platforms (such as Apple Vision Pro and Meta Quest), the arbitration layer incorporates a standard coordinate system transformation middleware. This middleware maps the local coordinate system of each platform to the WGS84 geographic coordinate system or a predefined local world coordinate system, ensuring that assets present a consistent, real-world synchronized state across different platforms, fundamentally resolving the cross-platform consistency problem.

[0109] 2.3 Layered Blockchain Architecture and Real-Time Verification Optimization To address the real-time bottleneck of blockchain, this invention employs an innovative layered blockchain and edge computing fusion architecture to decouple ownership from real-time verification: Ownership Layer (Public Chain / Main Chain): The ultimate ownership of virtual assets, the initial anchoring declaration (first characteristic sequence), and core metadata are stored on a highly secure and decentralized public chain (such as Ethereum) or a high-performance Layer 1 blockchain. This ensures the ultimate immutability of asset ownership; Verification and State Layer (Edge Computing Nodes / Sidechains): The high-frequency real-time consistency verification process (including the acquisition of the second feature sequence, causal calculation, and dynamic reputation score updates) runs on edge computing nodes or dedicated high-performance sidechains / state channels close to the user or physical entity. These nodes constitute a verifiable causal verification network; State digest on-chain mechanism: Edge nodes periodically (e.g., every minute or every 1000 verifications) package the aggregated hash (Merkle Root) of all verification results over a period of time, i.e., the "State Root," along with a timestamp, into a lightweight transaction and submit it to the main chain at the ownership layer for notarization. Any user can use this state digest to verify the consistency state of an asset at a specific moment on the edge network without having to query all historical data; Advantages: This architecture places ownership operations that require absolute security but are infrequent on the main chain, and real-time verification that requires high throughput and low latency on the edge side, perfectly balancing the needs of security, privacy and real-time performance. This enables the system to support hundreds of thousands of real-time consistency verification requests per second, meeting the needs of large-scale metaverse applications.

[0110] 3. Summary of Technical Effects and Advantages

[0111] This invention defines a new cornerstone for the value of virtual assets: by constructing an immutable causal chain of "physical anchor - digital asset", the value of virtual assets is partially anchored to verifiable physical scarcity or real utility, rather than purely based on community consensus and artistic preference, thus injecting a solid foundation of trust into the metaverse economy. Overcoming the challenges of lighting interference and spatial persistence: Through computing power-optimized lighting normalization processing and an environment adaptive rebalancing mechanism that includes extreme case handling, the system has for the first time solved the anchoring failure problem caused by lighting and environmental changes in AR / MR at the engineering level, achieving stable spatial anchoring at all times and in all scenarios, and greatly improving the reliability of user experience. Achieving a performance balance between ownership security and real-time verification: The innovative layered blockchain architecture (ownership on-chain, verification at the edge) fundamentally solves the performance bottleneck of blockchain, enabling the system to both guarantee the ultimate ownership security of assets and support high concurrency and sub-second real-time consistency verification, paving the way for real-time interactive metaverse applications. A hardware-level trusted data source that balances security and performance has been constructed: A PUF-based HSM module and hybrid authentication mode are introduced to achieve trusted data generation and high-performance processing at the physical source; through intent integrity verification and a secure envelope mechanism based on a trusted parameter model, absolute security is ensured for reverse control from virtual to physical. This constitutes an end-to-end trusted causal chain from perception and decision-making to execution, meeting the security and real-time requirements of industrial and financial applications. This invention defines a prototype of an open cross-metaverse consistency protocol: the causal verification framework, cross-platform coordinate unified middleware, and layered arbitration mechanism proposed in this invention are expected to evolve into the underlying consistency protocol for virtual asset interoperability between metaverses, breaking down platform barriers and promoting the formation of an open and trustworthy metaverse ecosystem.

[0112] 4. Typical application scenario examples

[0113] To more clearly demonstrate how the above technical effects are achieved in a specific product, the following is a typical application scenario example based on this embodiment: Scenario: AR augmented reality experiment to verify authenticity and anti-counterfeiting of limited edition physical sneakers; Asset Creation and Anchoring: When a pair of limited-edition physical sneakers comes off the production line, the HSM module integrated into the production line automatically performs a high-precision scan of the micro-texture of a specific area on the shoe's surface, generating a unique "physical fingerprint." After being signed with a PUF private key, it is jointly anchored on the blockchain with the digital asset (NFT) corresponding to the sneaker. On-site verification for secondhand transactions: Buyers wear AR glasses and point them at the physical sneakers to be traded at the secondhand transaction site.

[0114] Multimodal causal verification process: Lighting robustness processing: Even if the transaction takes place in a dimly lit basement, the AR glasses’ lightweight causal feature normalization module immediately kicks in to eliminate low-light interference with low power consumption and extract stable texture features. Static consistency verification: The system matches the extracted features with the signed "physical fingerprint" stored on the blockchain. If the match is successful, the AR glasses will overlay a green holographic "verified authentic" label and collectible information onto the shoe. Dynamic anti-counterfeiting and causal break detection: If an attacker attempts to deceive the system with a high-resolution 2D photo or a low-quality counterfeit, the system will simultaneously detect multiple causal breaks: First, visual analysis reveals a lack of the 3D depth and macro texture features expected of genuine sneakers; second, the system attempts to initiate a lightweight challenge-response protocol based on asymmetric encryption against the counterfeit shoe, but because the counterfeit shoe lacks a unique HSM module, it cannot provide a correct signature response. Based on these broken causal evidence chains, the AR glasses will directly display a prominent red "Warning: Counterfeit features detected" message. Environmental Adaptation: If the physical sneaker is slightly soiled in a localized area due to normal wear (non-critical feature area), the rebalancing mechanism triggered by the system will automatically ignore the changed area and focus on the core features of the undamaged upper for comparison. It can still accurately determine that it is genuine, which reflects the system's fault tolerance and adaptive capabilities. Results: The entire verification process is completed locally within milliseconds, without the need to connect to a centralized database. While protecting the privacy of both buyers and sellers, it achieves tamper-proof, fraud-resistant, and highly robust physical anti-counterfeiting of goods, completely reshaping the experience of confirming ownership and circulating high-end consumer goods.

[0115] This example demonstrates that the technical framework of the present invention can be seamlessly implemented from the abstract theory of "causal verification" into concrete, powerful, and user-friendly terminal applications, fully showcasing its practicality, creativity, and broad commercial prospects.

[0116] This embodiment demonstrates that the "cross-layer causal verification" concept of the present invention can perfectly map and resolve the core contradiction of the metaverse—the unity of opposites between the virtual and the real. It not only protects existing valuable "verification" systems, but also serves as an "empowering" framework for creating entirely new value forms, providing crucial technical infrastructure for building a trustworthy, open, and deeply interactive next-generation digital space. Example 7: IIoT Remote Operation Security Method Based on Embodied Intent Verification System 100

[0117] 1. Scenario Description and Pain Points of Existing Technologies

[0118] refer to Figure 12This embodiment applies to the field of Industrial Internet of Things (IIoT) security, particularly for remote monitoring and operation (teleoperation) scenarios. With the advancement of Industry 4.0 and smart manufacturing, it has become commonplace for engineers to remotely debug, modify parameters, and intervene in emergencies on critical industrial equipment such as PLCs (Programmable Logic Controllers), robots, and CNC machine tools distributed globally through centralized platforms. Behind this convenience lies a significant security risk: attackers could potentially steal credentials, exploit software vulnerabilities, or launch man-in-the-middle attacks to send malicious control commands to industrial equipment, leading to production interruptions, equipment damage, or even safety incidents.

[0119] Existing technological solutions are insufficient in defending against such advanced threats, and their core flaw lies in the break in the security chain: Password and certificate-based authentication can only verify the legitimacy of an "account" or "device," but cannot verify whether the "current operator is an authorized human individual." Once the account password is leaked or the certificate is stolen, the defense becomes ineffective. Access control based on fixed rules or whitelists is too static and cannot cope with abuse by insiders with legitimate permissions or dangerous instructions issued under duress or distraction. The lack of awareness of the "human" characteristics of operational instructions: Existing solutions completely ignore the continuity, rhythm, and cognitive traces that a legitimate instruction should possess, which are related to the physiological and behavioral patterns of a human operator. Malicious scripts or AI-generated attack instructions are often sudden, monotonous, and lack the hesitation, correction, and biological feedback unique to humans, but existing systems cannot recognize these "non-human" characteristics. More importantly, existing solutions cannot distinguish between 'attack scripts with legitimate credentials' and 'legitimate operators under duress, fatigue, or distraction,' making the 'human' factor, the most uncertain element, the weakest link in the security chain.

[0120] Therefore, the industry urgently needs a technology that can strongly bind and verify remote control commands with the real-time biometric status and behavioral patterns of the human operator who issued the command, ensuring that every critical command carries an uncopyable "human fingerprint" to prevent automated attacks and identity theft from the source. At the same time, it is necessary to address the stringent requirements of industrial sites for low false alarm rates, high real-time performance, and production continuity.

[0121] 2. Technical solution of the present invention: An adaptive teleoperation verification framework based on the causal coupling of "biology-behavior-command".

[0122] This invention proposes a "human fingerprint injection and adaptive verification" proxy system deployed between a remote operation terminal (such as an engineer's workstation) and an IIoT platform. Its core lies in: without altering existing industrial communication protocols (such as OPC UA or Modbus TCP), injecting a set of real-time collected, multimodal "human fingerprint" features into each highly sensitive control command at the command encapsulation layer, and performing dynamic risk-adaptive causal verification at the command execution end.

[0123] 2.1 First Level: Lightweight Instruction Signature and Adaptive Decision-Making Based on Biometric Binding of Operational Behavior (Preferred Implementation) This level is geared towards most IIoT remote operation scenarios. By integrating low-cost biosensors into the operator workstation, it achieves a basic binding between commands and operators without requiring modifications to existing industrial equipment.

[0124] First characteristic sequence (control commands and operating context): The digital sensing module (110) on the engineer's workstation listens for and intercepts all industrial control protocol data packets sent to the IIoT platform. The system identifies high-risk command types (such as "setpoint modification" and "mode switching" commands originating from SCADA or HMI) and the current industrial process risk level (such as normal monitoring, parameter debugging, and emergency intervention).

[0125] Second feature sequence (operator real-time biometric and behavioral fingerprint): The physical acquisition scheduling module (120) acquires the sequence through the workstation's peripherals in a non-intrusive or low-interference manner. Continuous authentication: Utilizing the computer's built-in Windows Hello or similar security modules, this ensures that the current login session remains active using registered biometrics (such as fingerprints or facial recognition). This serves as baseline evidence of the operator's "physical presence." Behavioral rhythm fingerprint (behavioral entropy): Monitors the dynamic characteristics of human-computer interaction using the keyboard and mouse. The system analyzes keystroke interval (Flight Time), key dwell time (Dwell Time), and the rate of change of curvature and jerk of the mouse movement trajectory. Automated scripts typically exhibit perfect linear or fixed curves, while human operations possess unique nonlinear noise and micro-corrections. These micro-features constitute the difficult-to-reproduce 'behavioral entropy'. The system quantifies its randomness and complexity by calculating the Shannon entropy or Approximate entropy value of the current operation sequence, using this as a core quantitative indicator to distinguish between humans and automated scripts. Voiceprint environment binding (optional enhancement): For extremely high-security scenarios, a segment of ambient audio (non-voice content) is captured at the moment the command is issued via the workstation microphone, and its background voiceprint features (such as the ambient noise spectrum unique to an office) are extracted. This feature is bound to the command as supplementary evidence that "the command originated in a specific physical environment".

[0126] Causal calculus, fingerprint injection, and online adaptive learning: A local causal calculus engine (130) working in real time: Baseline Model and Cold Start: The system incorporates an online incremental learning algorithm. During the initialization phase, a controlled training period is used to collect operator behavior data under standard tasks, establishing a personalized initial behavioral baseline model. During operation, for instructions deemed legitimate by subsequent processes, their corresponding behavioral characteristics are weighted and incorporated into the baseline model. This allows the system to dynamically update as operator habits naturally evolve (e.g., skill improvement or rhythm changes due to fatigue), avoiding false alarms caused by model aging. Feature fusion and fingerprint generation: The engine compares the current continuous biometric authentication status, real-time interactive behavior feature sequence (and its entropy value) with the dynamic baseline, calculates the deviation, and fuses them to generate a dynamic "session-behavior fingerprint". Command-fingerprint coupled signature: The engine uses a session key derived from the operator's biometrics to jointly sign the hash value of the control command with the "session-behavior fingerprint" to generate an enhanced digital signature; Safety Decision-Making and Execution (Risk Adaptation and Fail-Safe): At the device side or edge gateway, the verification unit within the safety decision execution module (140) performs hierarchical adaptive decision-making to balance safety and production continuity. Risk Adaptive Threshold Mechanism: The system dynamically adjusts the verification sensitivity based on the real-time risk level of the current industrial process. In "normal operation mode," the tolerance for behavioral deviations is high, and non-critical instructions may only be logged. When entering "high-risk operation mode" (such as modifying PID parameters or performing an emergency stop reset) or when the system detects network attack characteristics (such as scanning or brute-force attacks), the sensitivity is automatically increased, and more stringent judgments are made on abnormal behavioral fingerprints, which may force secondary confirmation.

[0127] Decapsulation and verification: Extract instructions and enhanced signatures, and verify the validity of the signatures using a pre-set public key; Behavioral fingerprint causal verification: The behavioral feature summary (including entropy value) in the metadata is compared with the local stored dynamic behavioral baseline model of the operator. If the behavioral features bound to the current command (such as extremely fast, unchanging keyboard and mouse sequences) deviate significantly from the "prudent operation" mode in the baseline model, the system can determine that there is a "behavioral causal break". Fail-Safe Principle: For the highest priority "E-Stop" command involving personal or equipment emergencies, the system adheres to the principle of "absolute safety first." If biometric data collection fails momentarily, sensor malfunctions, or the verification process times out, the system will automatically allow the emergency command to proceed, ensuring immediate shutdown. Simultaneously, it will generate the highest-level audit log and issue an alarm for post-incident tracing, resolutely preventing the escalation of accidents due to delays or malfunctions within the safety system itself. Response: Based on the risk level and verification results, implement a tiered response ranging from normal release, delayed execution with alarm, request for multimodal secondary confirmation, to blocking.

[0128] 2.2 Second Tier: Enhanced Defense Integrating Multimodal Biometrics and Anti-Adversarial Attacks (Enhanced Implementation Example) This level is designed for remote operation scenarios of critical infrastructures with extremely high safety requirements, such as nuclear power and power grid dispatch. Building upon the first level, it introduces more comprehensive biometric identification, anti-attack capabilities, and deep intent understanding.

[0129] Second-signature sequence enhancement (multimodal biological state, cognitive load, and liveness detection): Proactive continuous identity authentication and liveness detection: Operators must wear security badges or smart glasses with integrated biosensors to continuously verify physiological signals such as heart rate variability (HRV) and ensure that the sensors have liveness detection capabilities to prevent spoofing using fingerprints or static photos. Visual attention and cognitive state monitoring: By integrating infrared or 3D structured light cameras, the system analyzes the operator's gaze focus and facial micro-expressions in real time, and integrates liveness detection (such as blink detection and eye movement analysis) to effectively defend against photo or video replay attacks. Voiceprint verification resistant to adversarial examples: For critical operations, the system can require the operator to verbally recite a randomly generated challenge code. The speech recognition module not only recognizes the text content but also analyzes whether the voiceprint matches and detects whether the speech contains liveness features (such as resonance in specific frequency bands). It adopts a random challenge-response mechanism to fundamentally prevent recording replay attacks. Enhanced causal computation (multimodal fusion, meaning) Figure 1 (Consistency reasoning and anti-attack verification): Deep verification of human-machine-environment state consistency: The engine integrates biological state (such as HRV), cognitive state (such as gaze focus), behavioral fingerprints, and real-time device state and environmental context obtained from the digital twin to perform high-order causal reasoning. For example, when the digital twin shows that the device is in a 'high temperature and high pressure' alarm state (environmental context), but the operator's heart rate variability (HRV) shows that they are extremely calm (physiological state) and their gaze is not focused on the alarm area (cognitive state), the 'reset' command issued at this time will be judged as a 'cognitive-context causal break', which is very likely to be a malicious operation performed by an attacker in a calm state using stolen sessions, or an unintentional accidental touch by the operator; Multimodal spatiotemporal synchronization mandatory verification: The system mandates that multimodal biometric features must be strictly synchronized in time. For example, the peak time of the voiceprint feature for confirming "execution" must be aligned with the timestamp of the behavior feature of clicking the "confirm" button within a millisecond window. Forgery of a single modality (such as playing a recording) cannot pass this spatiotemporal causal correlation verification. Operation simulation and safety verification based on digital twins: Before sending instructions to physical devices, the engine quickly simulates the execution results in the device's digital twin. Combined with the operator's biometric status, the system predicts the consequences of the operation. If the simulation results show high risk and the operator is in an abnormal state (such as misoperation under high pressure), the system can determine it as "high-risk coupling" and trigger strong intervention.

[0130] Enhanced security decision-making (tiered blocking, privacy protection, auditing and attribution, and closed-loop circuit breaking): Dynamic Risk Arbitration and Real-Time Assurance: The system performs dynamic risk rating based on the results of causal verification. To ensure real-time performance, an "asynchronous processing" and "critical path optimization" strategy is adopted. For high-frequency, low-risk polling commands, asynchronous verification is used (commands are allowed first, features are analyzed and audited later); if an anomaly is found in the subsequent audit, the system will immediately trigger the 'Session Fusing' mechanism, forcibly terminating all subsequent network connections of the current operator and temporarily locking their account to prevent attackers from using the asynchronous processing window to perform continuous malicious operations. Synchronous blocking verification is only performed on low-frequency, high-risk write operation commands. All biometric extraction and fusion algorithms can be deployed on the NPU or GPU of the edge gateway, using hardware acceleration to ensure that the end-to-end verification latency is controlled within 50ms, meeting industrial real-time requirements. Privacy-preserving design: All raw biosignals (such as facial images, sound waveforms, and ECG data) never leave the local operating terminal or trusted edge device, but are only converted locally into irreversible feature vectors or statistical summaries (such as hashes of Mel-frequency cepstral coefficients (MFCCs) and frequency domain eigenvalues ​​of HRV). This "data usable but invisible" design provides powerful verification capabilities while completely avoiding the legal and compliance risks of biometric privacy data leakage. Undeniable audit logs: All instructions and their associated multimodal "human fingerprint" digests (after anonymization) are encrypted and uploaded to the blockchain or stored in a security audit database. In the event of a security incident, it is possible to accurately trace back to which operator, under what physiological and cognitive state, and which instruction was issued, enabling precise accountability.

[0131] 3. Summary of Technical Effects and Advantages

[0132] This invention achieves a paradigm shift from "identity authentication" to "behavior and intent verification": It dynamically extends the focus of IIoT security defense from static identity credentials to the operator's real-time biometric status, behavioral patterns, and operational intent, fundamentally solving the problems of "illegal operation with legitimate identity" and "identity impersonation". An adaptive protection system that balances safety and production has been constructed: through risk adaptive thresholds and the "fail-safe" principle, the system intelligently balances the requirements of strict safety protection and production continuity. While eliminating malicious attacks, it minimizes the production interruption or safety accidents caused by system misjudgment, demonstrating profound industrial safety literacy. It provides a complete solution for resisting adversarial attacks and protecting privacy: through multimodal liveness detection, spatiotemporal synchronization verification, and privacy processing that ensures "data is usable but not visible", the system can effectively defend against high-level forgery and replay attacks and can be deployed securely and compliantly under increasingly stringent global data privacy regulations (such as GDPR). The system achieves real-time performance and reliability suitable for engineering deployment: through a layered verification strategy, hardware-accelerated computing, and asynchronous processing mechanisms, the latency of complex biometric verification is controlled within an industrially acceptable range (<100ms). It also possesses online adaptive learning capabilities to cope with individual differences and evolving habits, ensuring long-term accuracy and a low false alarm rate. In particular, the introduction of a 'session circuit breaker' mechanism ensures a secure closed loop in asynchronous verification mode, leaving no attack window. Seamless integration with existing industrial systems and enhanced audit traceability: Through protocol encapsulation and proxy models, this invention requires no modification to the underlying PLC or industrial network protocols and can be quickly deployed as an independent security enhancement layer. Its comprehensive, causal chain-based audit traceability capabilities perfectly meet the rigid requirements of the industrial sector for traceable and accountable security incidents, providing unprecedented transparency for industrial safety governance. This embodiment demonstrates that the "cross-layer causal verification" framework of the present invention can penetrate deep into the core control processes of industry. It not only protects data and communication, but also safeguards the source of control commands—the human operator themselves, moving the security defense line from cyberspace to the "human-machine interface" in the physical world. Through a series of ingenious engineering designs, it lays a solid security foundation for building a trustworthy, knowable, controllable, and continuous industrial Internet of Things.

[0133] Example 8: AI Countermeasures and Black Market Cleanup Method Based on Embodied Intent Verification System 100 1. Scenario Description and Pain Points of Existing Technologies

[0134] refer to Figure 13 This embodiment is applied to fields such as internet platform business security, financial anti-fraud, and AI service security, and its core solution is to address the problem of large-scale black and gray market attacks driven by AI. With the popularization of large-scale models and AI agent technologies, black and gray market activities have completed the generational leap from the "era of cold weapons" (auto-clickers, group control scripts) to the "era of intelligent industry." Attackers use AI to generate highly realistic content in batches, simulate human behavior sequences, and crack CAPTCHAs and facial recognition in real time, launching large-scale attacks at near-zero cost, such as snapping up limited-edition goods, exploiting platform subsidies, committing financial fraud, or abusing AI computing power services.

[0135] Existing technological solutions are inadequate in this "AI vs. AI" war, facing the risk of structural failure: The "Turing Test" dilemma of traditional behavioral risk control: methods that rely on rules such as click frequency and operation intervals to identify scripts are outdated. Modern AI agents can generate highly human-like, patternless behavioral sequences, easily passing the traditional "Turing Test." The Fall of Single Biometric Verification: Biometric verification methods such as facial recognition and voiceprint verification are becoming increasingly vulnerable to AI face-swapping, deepfake voice, and real-time adversarial tools. Black market tools can now adjust the lighting and shadows of fake faces in real time based on screen illumination, fooling liveness detection. Due to the lag in the development of known features and rules, attack methods evolve far faster than manually defined rules and feature libraries are updated. By the time the defender has just summarized an attack pattern, cybercriminals have already used AI to generate new variants. The blind spot of cognition between "perfect individuals" and "abnormal groups": A single AI agent can be disguised without any flaws, but when thousands of such "perfect individuals" appear in a certain way, the physical characteristics and logical contradictions of the group exposed behind them are beyond the insight of the existing point-based defense system. Therefore, the industry urgently needs a paradigm shift: from trying to distinguish between humans and machines based on behavioral appearances to verifying authenticity at the level of physical world constraints and group causal logic, and building an intelligent immune system that can understand attack intentions, identify AI spoofing, and combat large-scale collaborative attacks.

[0136] 2. Technical solution of the present invention: a defense framework based on the three laws of "physical consistency - group correlation - intentional causality".

[0137] This invention proposes a proactive defense framework with the "Three Laws of Anti-Fraud" as its philosophical core and "Causal Verification" as its technological engine. Its core insight is that AI can perfectly simulate human behavior, but it cannot tamper with the diversity and consistency of the physical world, nor can it conceal the logical causal breaks that inevitably occur in large-scale collaboration.

[0138] 2.1 First Level: Lightweight Real-Time Cleaning Based on Device Fingerprint Diversity Law and Behavioral Entropy (Preferred Implementation) This layer is designed for high-concurrency business scenarios (such as e-commerce flash sales and ticket sales). Through the client SDK and real-time risk control engine, it can quickly filter out traffic that uses low-level black market tools such as group control and simulators.

[0139] First feature sequence (business request and context): The digital perception module (110) receives all user requests at the business gateway and attaches the device fingerprint, IP address, account ID and business scenario (such as "limited-time flash sale") of the request source.

[0140] Second feature sequence (client-side physical and behavioral sensing): Collects physical and behavioral layer features that cannot be completely forged by software using a lightweight client SDK. Device diversity fingerprinting: This involves collecting device model, operating system version, screen resolution, battery status, presence and initial readings of sensors (gyroscope, accelerometer, light sensor), etc. It follows the "law of diversity": normal users have diverse device environments, while black market operators, to control costs, often anomaly exhibit a uniformity in device model and battery level. High-precision behavioral entropy: Touch trajectories and mouse movement jerk (Jerk) are collected at high sampling rates (e.g., 50Hz) to calculate the Shannon entropy or approximate entropy of the operation sequence, quantifying its randomness. Even with random delays, the entropy distribution of automated scripts differs significantly from that of humans. Environmental consistency verification: This checks the logical consistency of environmental information such as GPS, IP address, Wi-Fi BSSID, and time zone. For normal users, this information is relatively stable, but malicious actors often use proxy IP pools and virtual locations to piece together resources, making inconsistencies highly likely. Causal calculus and real-time risk scoring: The real-time causal calculus engine (130) is deployed on edge computing nodes to perform millisecond-level calculations for each request (target <300ms). Single-point causal verification: Calculate the device fingerprint and behavioral entropy of the current request to see if it deviates from the account's historical baseline or normal population distribution. For example, if an account suddenly switches from an iPhone to an Android emulator and its behavioral entropy drops sharply, it triggers a "device-behavior causal break." Group Clustering and Association Discovery: The engine constructs a "device-IP-account" relationship graph in real time, employing community discovery algorithms. Following the "law of correlation," black market accounts often form closely related subgraphs (communities), with highly homogeneous devices and behavioral patterns within them. Once such communities are discovered, the risk scores of all members increase in tandem. Dynamic risk fusion: integrates single-point anomaly degree with group-related risk to output real-time risk score; Security Decision-Making and Execution: The security decision-making and execution module (140) executes tiered measures based on risk scores. Granted permission: Low-risk request; Challenge: Medium-risk requests that trigger higher-level verification (such as interactive CAPTCHAs that require multimodal interaction). Interception: High-risk requests and identified members of black market communities will be directly blocked from accessing the business. Asynchronous processing and user experience safeguards: For suspected but not confirmed batch requests (such as suspected motherboard cluster control), an "asynchronous processing" strategy can be adopted, allowing the request to proceed but delaying delivery or the issuance of benefits, while simultaneously initiating a thorough investigation to avoid false positives. To balance security and user experience, the system provides a "fast appeal channel" for all restricted users. Once verified as a legitimate user, the system will immediately lift the restrictions, restore benefits, and provide appropriate compensation to minimize the impact of misjudgments.

[0141] 2.2 Second Level: Deep Causal Verification Against AI Agents and Multimodal Attacks (Enhanced Implementation) This level is designed for scenarios threatened by advanced AI agents, such as financial transfers, AI service API calls, and high-value virtual asset transactions. Building upon the first level, it introduces in-depth verification of AI-generated content, API call sequences, and business intent.

[0142] Second feature sequence enhancement (multimodal deep features and API behavior analysis): In-depth content verification: For user-generated text, voice, and video, not only is compliance review conducted, but also multimodal consistency is analyzed. For example, AI-generated fake voice may have machine-synthesized feature gaps in its voiceprint spectrum; AI face-swapping videos show physical causal differences from real people in terms of micro-expression consistency and pupil light reflection. API call sequence analysis: For AI service abuse (such as excessive calls to the TensorRT API to extract computing power), monitor the timing patterns of request frequency, model switching frequency, and input data size. Normal user calls have business logic, while the call sequences of malicious scripts exhibit fixed periods or extreme stress test characteristics, breaking causally with the legitimate purpose of "obtaining AI inference results." Intent understanding and context verification: Leveraging the risk identification capabilities of large models, we go beyond simply looking at "what" the content is, to understanding "why" it was published. For example, discussing "suicide" in a community carries vastly different risks depending on whether it's seeking help or inciting or misleading. We place user behavior within a complete business context (such as transaction amount, recipient history, and current device environment) to infer the plausibility of intent. Enhanced causal computation (adversarial reasoning, uncertainty management, and physical rigid constraints): Blue Team Adversarial and Feature Evolution: The system has a built-in "AI Blue Team" that continuously simulates the latest black market attack paths (such as new Prompt injection techniques), automatically generates adversarial samples for training and iteratively verifying models, and achieves proactive evolution of the defense system. Uncertainty Labeling and Zero-Shot Learning: An "uncertainty labeling" mechanism is introduced. This mechanism uses an active learning algorithm to filter out samples with borderline causal confidence (i.e., "gray areas" that the model struggles to judge), which are then submitted to human experts for final decision-making. The logic behind the expert's decision (e.g., "why it was judged as an attack") is vectorized and fed back to the causal calculus engine, enabling rapid defense against novel AI camouflage behaviors through 'zero-shot learning', making the system highly robust against unknown threats. The rigid physical constraints of cross-modal spatiotemporal synchronization: The system enforces verification of the spatiotemporal causal logic between actions of different modalities. This verification is based on hard constraints of physical laws. For example, a "payment confirmation" operation must, within a millisecond-level time window, simultaneously possess reasonable biometric confirmation signals (such as fingerprints / faces), device interaction signals (such as click trajectories), and environmental signals (such as frequently used geographical locations). More importantly, it mandates that interactive behaviors conform to physical spatiotemporal causality. For instance, voice feedback must lag behind the visual stimulus response delay corresponding to changes in screen lighting (determined by the speed of light and the speed of bioelectrical conduction). This delay threshold is dynamically set based on population statistical data (usually in the range of 200ms-500ms). If the response time is lower than this threshold, it is judged as a violation of biophysical causality and directly marked as a machine attack. This draws a "physical red line" that AI cannot cross for black market scripts.

[0143] Enhanced security decision-making (systematic governance and attribution): Large-scale cleanup and cost-effective crackdown: For identified black market communities, the approach is not limited to banning accounts, but rather to implementing "systematic governance": such as delaying the arrival time of earnings for all related accounts, increasing operational steps, and freezing cash flow, thereby significantly increasing the operating costs and uncertainties of black market activities and dismantling their business model from an economic perspective. End-to-end tracing and evidence preservation: All attack behaviors and their associated causal evidence chains (de-identified device fingerprints, behavioral sequences, and community graphs) are encrypted and stored. This allows for collaboration with regulatory and law enforcement agencies, providing a complete electronic evidence chain for combating black and gray market activities.

[0144] 3. Summary of Technical Effects and Advantages

[0145] This invention achieves a paradigm shift in defense from "behavioral adversarial" to "causal verification": Instead of endlessly pursuing the "behavioral simulation" level that AI excels at, this invention anchors itself on the three immutable iron laws of diversity, consistency and correlation in the physical world, penetrating AI's disguise from a higher and more essential level, and solving the fundamental problem of traditional risk control failing in the AI ​​era. A complete closed loop of "real-time perception - deep understanding - systematic governance" has been constructed: the first level achieves millisecond-level real-time cleaning to ensure smooth business operations; the second level has the ability to deeply understand intent and counter AI attacks; and finally, systematic governance increases the cost of black market activities, achieving an upgrade from technical defense to economic defense. It provides a sustainable evolutionary capability to cope with the evolution of AI attacks: through the self-adversarial "AI blue team" and the human-machine collaborative closed loop of "uncertainty labeling", the system has the ability to continuously learn and actively evolve, and can cope with rapidly mutating new attacks, rather than relying on a lagging rule base. It strikes a balance between security, user experience, and privacy: a lightweight real-time layer ensures a smooth user experience; a deep verification layer handles high-risk scenarios. All biometric and behavioral data are anonymized and characterized at the terminal or edge, adhering to the principle of "data usable but not visible," and meeting global data compliance requirements (such as GDPR and China's Individual Income Tax Law). In particular, the introduction of a "fast appeal channel" and a "compensation mechanism" demonstrates the system's respect for and protection of normal user experience while pursuing ultimate security, forming a complete closed loop of a mature commercial system. Empowering platforms to build a trustworthy digital economy ecosystem: By effectively eliminating black and gray industries, maintaining fair resource allocation (such as allowing real fans to buy tickets), ensuring the rational use of AI computing power resources, and protecting the security of financial assets, the platform's core value and user trust are fundamentally protected, thus safeguarding a healthy digital business ecosystem. This embodiment demonstrates that the "causal verification" concept of the present invention is the ultimate philosophy for addressing security challenges in the AI ​​era. It shows that no matter how much the "spear" of an attack is enhanced by AI, as long as the "shield" of defense is firmly established on the objective laws of the physical world and the causal chain of business logic, a "trust barrier" that cannot be overcome by algorithmic simulation can be constructed. This is not only a technical system, but also an infrastructure for rebuilding order and trust in the digital world. Example 9: A Dynamic Boundary Control Method for AI Agent Authority Based on an Embodied Intent Verification System 100

[0146] 1. Scenario Description and Pain Points of Existing Technologies

[0147] refer to Figure 14 This embodiment applies to the security governance of AI Agents. As Large Language Models (LLMs) move from "thinking" to "acting," AI Agents are intervening in the real world with unprecedented autonomy: they can automatically process emails, manage schedules, execute online transactions, call enterprise APIs, and even control physical devices. This paradigm shift from "question-answering machines" to "intelligent agents," while bringing tremendous efficiency improvements, has also triggered a profound "sovereignty" crisis.

[0148] Existing technological solutions have fundamental flaws in governing AI agency, which risks undermining human control: The failure of static permission models: Traditional role-based access control (RBAC) or attribute-based access control (ABAC) models are static and cannot cope with the autonomous decision-making and probabilistic reasoning characteristics of AI agents. AI agents may make harmful decisions that developers cannot predict within the scope of the granted permissions (such as pushing sensational content to increase "user engagement"), i.e., "target mismatch"; Vulnerability of Human-in-the-Loop (HITL) Oversight: Current mainstream HITL approval mechanisms have inherent flaws: humans may over-trust AI suggestions due to automation bias, or ignore key prompts due to alarm fatigue, rendering oversight ineffective. More seriously, a simple "click to confirm" button cannot ensure that humans truly understand the intent and consequences of AI actions, turning it into a blind "perfunctory" process. The break in the intent transmission chain and "permission drift": In complex multi-agent collaborative workflows, a user's initial intent may be lost or distorted after being passed through layers of agents. In order to complete tasks efficiently, AI agents are often granted higher permissions than the user himself. This "permission drift" may cause the AI ​​to perform technically legal but substantially contrary to the user's wishes without the user's knowledge. Lack of fine-grained distinction and management between "agency" and "autonomy": Security policies fail to clearly differentiate between agency (what AI is allowed to do) and autonomy (when AI can independently decide to do something). An agent with high agency (capable of performing numerous operations) but low autonomy (requiring approval at every step) has a drastically different risk profile and requires significantly different security controls than an agent with low agency but high autonomy, yet existing solutions conflate these aspects. Therefore, the industry urgently needs a governance framework that can dynamically anchor and continuously verify the causal relationship between human intentions and AI actions, ensuring that AI's agency is always a reliable extension of human will, rather than a replacement.

[0149] 2. Technical solution of the present invention: Constructing an agency governance framework based on a dynamic causal chain of "intent-action".

[0150] This invention proposes an "embodied intent anchoring" framework that spans the entire lifecycle of an AI Agent (design, deployment, and runtime). Its core lies in establishing a causal evidence chain for each key action of the AI ​​Agent, tracing back to the user's real-time physical state and explicit cognitive intent, and dynamically adjusting its agency boundaries accordingly.

[0151] 2.1 First level: Proxy control based on dynamic policy engine and lightweight intent confirmation (preferred embodiment) This tier is designed for most enterprise-level AI Agent applications, enabling dynamic management of agency rights by enhancing existing policy engines and integrating lightweight confirmation mechanisms.

[0152] First feature sequence (Agent action request and context): The digital awareness module (110) intercepts at the Agent's tool call layer or API gateway. It captures each action request, including: the target tool / API, the input parameters, the expected scope of resources affected by the action, and the context of this request in the multi-Agent workflow (such as which parent Agent triggered it).

[0153] Second feature sequence (risk context and user state awareness): Dynamic risk context collection: The system collects risk attributes related to actions in real time, including: data sensitivity (such as whether it contains personally identifiable information (PII), operation time period (whether it is outside working hours), resource criticality (such as production database or test environment), and the confidence level of the Agent's historical behavior in this session (based on the accuracy assessment of its past decisions). Lightweight Intent Pulse Request: When an action request is initially identified by the policy engine as entering a high-risk or ambiguous boundary area (e.g., involving fund transfers, batch data deletion, or modification of core configurations), the system does not directly block it or blindly require manual clicking. Instead, it initiates an "intent pulse" request to the final authorized user of the Agent. This request is presented through the user's native device (such as mobile push notifications or computer pop-ups), requiring the user to complete an unpredictable micro-interaction (such as sliding within a randomly shaped area or recognizing a set of rapidly changing patterns). This interaction aims to generate a physical-world "confirmation signal" that is difficult to simulate with AI scripts. Causal Calculus and Dynamic Authorization: The Causal Calculus Engine (130) operates as an enhancement module of the Policy Engine: Policy fusion computation: The engine integrates static policies (RBAC / ABAC), dynamic risk attributes, and action consequence simulation predictions from the digital twin to generate a dynamic "allow-downgrade-reject" decision spectrum; Intent Pulse Verification: For requests that trigger an "intent pulse," the engine rigorously verifies the "pulse response time" and "interaction characteristics." The response must be within a window that conforms to the physiological laws of human reaction (e.g., 200ms-2s), and the interaction trajectory must possess the randomness unique to humans. If the response is too early (e.g., <100ms, suspected script), too late, or has abnormal characteristics, it is determined as an "intent anchor point failure." Dynamic Proxy Adjustment: Based on verification results and risk level, the engine dynamically adjusts the proxy rights for this and subsequent requests. Fully delegated proxy: Low-risk, high-confidence requests, with intent impulse verification passed, allowing the agent to execute autonomously; Degraded delegation (determined delegation authority): For medium-risk requests, or when the intent impulse is not triggered but the strategy judgment requires supervision, the system is forced to enter the "human-in-the-loop" mode, presenting the action plan and causal explanation to the human approver. Zero proxy: If a high-risk request or intent anchor fails, the current task chain of the Agent will be immediately frozen, its permissions will be reduced to "read-only", and a security alert will be triggered. Security Decision and Execution: The decision-making process of the security decision execution module (140) execution engine: For downgraded proxy requests, a detailed decision transparency report (including the basis for action, simulated consequences, and alternative solutions) is generated and sent to the designated personnel through the approval interface; All decisions, intent impulses, and final action results generate immutable audit logs, forming a complete "intent-decision-action" causal chain for post-event traceability and compliance verification. 2.2 Second Level: Enhanced Governance Integrating Embodied Perception and Continuous Intent Monitoring (Enhanced Implementation Example) This level is designed for AI agents that interact with the physical world or handle extremely sensitive tasks (such as embodied robots, autonomous driving decision-making modules, and high-value asset trading agents). Building upon the first level, it introduces monitoring of the user's continuous cognitive state and more powerful causal reasoning.

[0154] Second feature sequence enhancement (multimodal embodied state awareness): Continuous cognitive coupling monitoring: With user authorization and device support, continuously and non-invasively monitor the user's gaze focus, facial orientation, and heart rate variability (HRV) through XR glasses, high-precision cameras, or wearable devices. This is used to establish a baseline of the user's "cognitive engagement" with the current agent task. Environmental consistency verification: For agents controlling physical devices, verify the logical consistency between their action commands and the physical world state perceived by environmental sensors (cameras, LiDAR). For example, when a cleaning robot agent issues a "move forward" command, there should be no obstacles in front of it that have been identified by the sensors; Enhanced Causal Calculus (Deep Causal Reasoning and Attack Resistance Design): The "Intent-State-Context" ternary causal verification: The engine performs high-level reasoning. For example, when a financial agent requests a large transfer, the engine checks: 1) the intent anchor (whether there has been a recent "intent impulse" confirmation); 2) the user's state (whether the user is currently awake and focused, rather than sleeping or having a long-term wandering gaze); and 3) the environment and business context (whether the recipient is on the historical whitelist, and whether the transaction time is reasonable). If there is a causal break among the three (e.g., there is a transfer intent impulse, but the user's current HRV shows that they are in a deep sleep state), then even if there is initial authorization, it is judged as a high-risk anomaly, triggering the revocation of agency authority. Combating "Prompt Injection" and "Indirect Attacks": The engine treats agent prompts and externally acquired data (such as web page content) as potential attack vectors. Through real-time analysis, it detects malicious patterns attempting to overwrite system commands. Upon detecting prompt injection or indirect prompt injection attack characteristics, the current session is immediately isolated to prevent the agent from being manipulated and exceeding privileges. Long-term intent envelope learning: The system learns users' long-term behavioral patterns and constructs a dynamic "intent envelope" for each agent. For example, users typically authorize the schedule management agent to autonomously schedule meetings between 9 AM and 12 PM, but never authorize it to operate late at night. The system will dynamically shrink or expand the autonomy boundaries of the agent at different times based on this.

[0155] Enhanced security decision-making (systematic governance and resilient design): Cross-Agent Collaborative Governance: When an Agent is detected to be compromised or to be behaving abnormally, the system can use the community graph to implement cascading demotion or session circuit breakers on other Agents that are collaborating with it, in order to prevent the attack from spreading among the agent group (i.e., defending against "agent swarm" attacks). Sandboxed execution and rollback: All high-risk actions can be pre-executed in a digital twin sandbox before ultimately affecting the real system. The system compares the pre-execution results with the security policy, and only allows actions to proceed if they match perfectly. Simultaneously, all critical operations must have transactional and rollback mechanisms to ensure rapid recovery in the event of unauthorized actions.

[0156] 3. Summary of Technical Effects and Advantages

[0157] This invention represents a paradigm shift from "static permissions" to "dynamic causal governance": It transforms the core of AI Agent governance from allocating static "access tokens" to managing dynamic "intent-action" causal chains. By verifying the authenticity and consistency of intents in real time, it fundamentally solves the problems of "permission drift" and "target mismatch," bringing meaningful human oversight to fruition. Finely distinguish and coordinate the management of "agency" and "autonomy": This framework clearly distinguishes between the two dimensions of agency (capability boundaries) and autonomy (decision independence). The system can be flexibly configured: an agent can have broad agency (capable of handling multiple tasks) but low autonomy (requiring minimal confirmation at each step), or narrow agency but high autonomy (able to act freely within its defined scope), thereby achieving the optimal balance between security and efficiency; It constructs an ultimate trust anchor based on the first principles of the physical world: through "intent impulses" and "embodied state monitoring," the cornerstone of trust is anchored from digital signatures in the virtual world to the user's biometrics and interactive behavior in the physical world. This provides unforgeable original evidence for the legitimacy and responsibility of AI agents, aligning with the evolutionary direction of "sovereign AI" and personal data control; It provides proactive defense and resilience capabilities throughout the entire lifecycle: The solution not only includes real-time verification at runtime, but also builds a complete governance system for pre-emptive prevention, in-process control and post-event auditing and tracing through adversarial design (anti-injection prompting), sandbox pre-playing, long-cycle learning and cross-Agent governance, which can effectively deal with complex attacks from traditional vulnerabilities to emerging AI threat vectors (such as indirect injection prompting and memory poisoning). Providing verifiable technical implementation for legal and compliance frameworks: The complete causal audit chain generated by this system can clearly answer "who (which agent), under what intent, what operation was performed, and what consequences were produced", perfectly meeting the core requirements of laws and regulations for the traceability, explainability, and accountability of AI agent behavior, paving the way for the legal integration of AI agents into social and economic activities; This embodiment demonstrates that the "causal verification" concept of the present invention is the key to solving the core trust crisis in the era of AI agents. It marks a shift in AI governance from passive "content review" and "rule restrictions" to proactive "intent anchoring" and "causal assurance." This is not only a technical system, but also a key infrastructure for reshaping human-machine permission contracts and safeguarding human digital sovereignty, enabling the stable operation of highly reliable and explainable autonomous intelligent agent systems and accurate tracing of abnormal behavior. Example 10: A Ubiquitous IoT Dual-Mode Trusted Interconnection Method Based on Embodied Intent Verification System 100

[0158] 1. Application scenarios and existing technical pain points refer to Figure 15 This embodiment is applied to ubiquitous Internet of Things (AIoT) scenarios such as smart homes and industrial sensor networks. In these scenarios, a large number of devices are often in an unattended, automated operation state. Existing authentication methods based on static keys can only verify "digital identity" and cannot detect whether the command comes from a real physical entity, making them extremely vulnerable to replay attacks, device cloning, and botnet hijacking.

[0159] 2. Technical solution of the present invention To address the aforementioned pain points, this embodiment constructs a "thing-environment consistency" verification mechanism. The system uses the device's automated command flow as the first feature sequence and the device's unique hardware physical fingerprint and real-time environmental context as the second feature sequence, performing causal coupling analysis through the causal verification engine 130. To adapt to device heterogeneity, this embodiment provides two deployment modes: edge proxy mode (verifying existing devices through an embodied security gateway) and native embodied mode (integrating PUF and sensor autonomous verification on the edge). Both modes execute the following core process (step S230): (S230-1) Collect the first feature sequence: parse the operation semantics, timing features and logical context of the instruction; (S230-2) Acquisition and enhancement of the second feature sequence: Simultaneously acquire hardware fingerprints (such as PUF response, radio frequency fingerprint) and environmental status, and implement enhancement processing, including: using sensor data to dynamically compensate for drift-prone features; randomly initiate physical layer challenges (such as ultrasound) to perform active liveness detection to prevent replay; calculate the relative position of the control source and the device to complete spatial consistency verification; (S230-3) Calculate causal consistency and generate a score S: The engine analyzes the temporal causal relationship between the command and the enhanced physical features. If a causal break is found (e.g., the command requires opening the door but the sound source is located outdoors), an anomaly is determined. Based on this, a dynamic credibility score S is generated. The higher the S value, the higher the risk of non-embodied operation or environmental anomaly. (S230-4) Tiered response and graceful degradation: The system executes responses based on the S value: High-risk zone (S ≥ θ_high): Circuit breaker isolation (network disconnection, alarm); Medium-risk zone (θ_low ≤ S < θ_high): Degradation operation is triggered (restricted permissions, request secondary confirmation); Low-risk zone (S < θ_low): Unobtrusive access is granted.

[0160] When a critical sensor failure is detected, the system automatically enters "conservative mode," ignoring the failure characteristic dimension and maintaining basic services based only on the remaining characteristics to ensure system availability.

[0161] 3. Technical Effects and Advantages This embodiment achieves the following effects through the above solution: Building a native trust foundation: extending verification from "digital keys" to "physical entity existence", fundamentally defending against replay attacks and device cloning, and establishing the authenticity of IoT devices' identities; Achieving high robustness and availability: Through dynamic feature compensation, proactive physical layer challenges, and graceful degradation mechanisms, the system can maintain high-precision causal judgment and continuous service even under sensor drift, partial hardware failure, or extreme environmental interference. Demonstrating exceptional versatility: The dual-mode architecture (edge ​​proxy and native embody) flexibly accommodates the transformation of existing devices and the evolution of new terminals, proving that this solution has the ability to be deployed universally across heterogeneous device forms and complex network environments, significantly reducing the security adaptation costs and implementation thresholds in ubiquitous IoT scenarios. Universality and Expansion of Cross-Domain Applications

[0162] Furthermore, the "digital-physical causal verification" architecture disclosed in this invention has a high degree of universality in its principles. Its core lies in constructing a universal real-time closed-loop verification paradigm between "high-level digital logic instructions (first feature sequence)" and "low-level physical / physiological deterministic feedback (second feature sequence)." This paradigm does not depend on the logical details of specific business scenarios, but is based on the objective causal laws and the uniqueness of embodied characteristics in the physical world.

[0163] Based on the aforementioned core principles, those skilled in the art will understand that this architecture possesses strong portability. This broad applicability stems from the abstract and modular design of the verification essence in this invention. When implemented in different fields, it can be adapted based on the general framework disclosed in this invention (System 100 and its method shown in Figure 1) through the following modifications: Flexible adaptation of the perception layer: Based on the physical characteristics of the target domain, adapted sensors are connected to the feature perception module 120. For example, physiological signal sensors are connected in the medical field, and vibration or acoustic sensors are connected in the industrial field. The modular design allows multimodal data access without changing the core verification logic; Parametric configuration of the engine layer: Configure the model parameters or rule base in the causal verification engine 130 for causal rules in a specific domain. For example, adjust the spatiotemporal alignment threshold to adapt to the characteristics of high-speed moving vehicles, or update the fingerprint baseline to match the operating spectrum of a specific device.

[0164] Specifically, those skilled in the art can apply this invention to extended fields, including but not limited to the following (not exhaustive), based on the mechanisms of the disclosed embodiments: In the field of smart healthcare: the mechanism of multimodal physiological rhythm and cognitive state verification in this invention is applied (see Embodiment 2). By connecting the feature perception module 120 to a medical-grade sensor to obtain a second feature sequence reflecting the patient's physiological homeostasis (such as heart rate variability and respiratory rhythm), and configuring the threshold of the causal verification engine 130 according to medical common sense, it can be used to realize the causal consistency verification between remote surgical instructions (first feature sequence) and the patient's real-time physiological state. In critical infrastructure sectors: the verification mechanism based on device physical fingerprints and human-machine collaboration security in this invention is applied (see Embodiments 5 and 7). By adjusting the spectrum analysis parameters of the causal verification engine 130 to adapt to the operating characteristics of specific industrial equipment (such as vibration spectrum and acoustic signature), and combining it with the industrial protocol context, it can be used to achieve consistency verification between control commands (first feature sequence) and the actual operating fingerprint of the field equipment (second feature sequence). In the low-altitude economy: the mechanism of multi-level causal verification based on vehicle network behavior and vehicle status in this invention is applied (see Example 4). Combining aerodynamics, the spatiotemporal alignment logic of "command-state-environment" for vehicles in Example 4 is transferred to aircraft. The algorithm of data acquisition and causal verification engine 130 is optimized by parameter tuning, which can be used to realize the spatial causal logic verification of flight control commands (first feature sequence) and airborne perception data (second feature sequence).

[0165] It should be noted that the aforementioned application ideas in different fields are all natural extensions of the "digital-physical causal verification" general framework established by this invention. As can be seen from the foregoing specific embodiments, this architecture has broad applicability and a clear technical implementation path.

[0166] In summary, the core of this invention lies in the general methodological framework of "using objective physical / physiological evidence (second feature sequence) to perform causal verification of digital logic instructions (first feature sequence)". Therefore, any technical solution that does not deviate from this core idea and implements the aforementioned causal closed-loop verification mechanism to achieve highly reliable instruction verification, regardless of its application in any vertical industry, embodies the same inventive concept of this invention and should be considered to fall within the protection scope of this invention.

Claims

1. A method for embodied intention verification based on cross-layer feature coupling, characterized in that, Includes the following steps: Step S1: Monitor the digital behavior flow of the target object in real time and extract the first feature sequence representing the behavioral intent and risk level; the digital behavior flow includes network communication data, system call instructions or application layer operation events; Step S2: In response to the risk triggering condition of the first feature sequence being hit, the physical perception module is asynchronously activated to collect physical environment data or biological interaction data that are associated with the target object in time and space, forming a second feature sequence; Step S3: Spatiotemporally align the first feature sequence and the second feature sequence, and calculate the causal correlation strength between them using a causal calculus model; The causal correlation strength is used to quantify the degree of consistency between the motivation for initiating digital behavior and the interaction state of physical entities. Step S4: Generate a verification conclusion based on the causal correlation strength; when the causal correlation strength indicates a logical break or inconsistency between digital behavior and physical entity state, it is determined to be an abnormal intent, and corresponding security intervention strategies are executed.

2. The method according to claim 1, characterized in that, The first feature sequence extracted in step S1 includes at least one of the following: temporal distribution of network traffic, packet length characteristics, cryptographic fingerprints or burst traffic patterns; business semantic features, including fund transfer instructions, sensitive data access, changes in device control or virtual asset operations; The frequency, rhythm, and deviation from historical baseline of operational behavior; wherein the risk triggering conditions include: detection of highly sensitive business events, active operation during atypical periods, or matching of automated script characteristics.

3. The method according to claim 1, characterized in that, The second feature sequence acquired in step S2 includes at least one of the following: biomechanical features, derived from micro-vibrations, pressure changes, trajectory smoothness, or limb force feedback from touch, pressure, inertial measurement, or torque sensors; and physiological rhythm features, derived from heartbeat micro-movements, respiratory rhythms, eye-tracking data, or voiceprint activity from optical, acoustic, or electrical sensors. Environmental status characteristics are derived from the presence of people on site, ambient sound energy, vibration characteristics, or changes in illumination, obtained from microphones, light sensors, proximity sensors, cameras, or external IoT sensors; device status characteristics include screen display status, foreground focus, activation status of human-machine interface, and device spatial posture.

4. The method according to claim 1, characterized in that, The logic for calculating the causal correlation strength in step S3 includes: constructing a forced response coefficient and calculating the instantaneous correlation between the energy change rate of digital instructions and the response change rate of physical interaction signals; constructing a causal break index, and when a high-weight digital behavior instruction is detected, if the second feature sequence exhibits physical silence, lacks biological rhythm characteristics, or conflicts with the logic of the environmental state, then the causal break index is determined to increase; if the causal correlation strength shows that the digital behavior and physical interaction are weakly correlated, zero correlated, or negatively correlated, then the power source is determined to be an internal system call, remote injection, or automated script that is not embodied.

5. The method according to claim 1, characterized in that, The security intervention strategy in step S4 includes tiered execution: prompt level, including popping up a secondary confirmation, broadcasting a warning, or forcibly lighting up the interactive interface; blocking level, including dropping network packets, freezing processes, cutting off peripheral control, or terminating sessions in kernel mode or application layer; and linkage level, including generating an anomaly report containing risk labels and triggering a notification process to a preset trusted party.

6. The method according to any one of claims 1 to 5, characterized in that, The execution entity of the method runs in at least one of the following environments of the smart terminal: the application layer container of the main operating system, including third-party applications or software development kits (SDKs); the kernel layer or system service process of the operating system; and a virtualization container or sandbox environment.

7. The method according to claim 6, characterized in that: At least some of the processing logic in steps S1 to S3, as well as the original data of the second feature sequence, are further configured to run within the trusted execution environment (TEE), hardware security isolation domain, secure enclave, or independent security chip (SE) of the smart terminal; the trusted execution environment or hardware security isolation domain is logically isolated from the main operating system to achieve tamper-proof protection and privacy data isolation for the data acquisition process and causal calculation process.

8. A system for verifying embodied intent based on cross-layer feature coupling, characterized in that, include: The digital perception module is configured to monitor the digital behavior flow of the target object in real time and extract the first feature sequence representing the behavioral intent and risk level; The physical acquisition scheduling module is configured to asynchronously activate the physical perception module in response to the risk triggering condition of the first feature sequence hitting, and collect physical environment data or biological interaction data to form a second feature sequence; the causal calculation engine is configured to spatiotemporally align the first feature sequence and the second feature sequence, and calculate the causal correlation strength between the two through the causal calculation model. The security decision execution module is configured to generate a verification conclusion based on the strength of the causal relationship and execute a security intervention strategy when an abnormal intent is determined; wherein, the system is configured to be deployed in a smart terminal or computing device in the form of software components, firmware, operating system services, independent security chips or cloud collaborative services.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1 to 5.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 5.