Agent output correction and cause-effect chain storage method and system based on physical constraints
Patent Information
- Application Number
- CN202610849082.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]本申请提供了基于物理约束的智能体输出修正与因果链存证方法及系统,针对解决现有技术中智能体修正过程不可追溯、责任难以认定、证据效力不足等技术问题
本申请首先接收智能体待下发的原始指令,解析动作参数,为后续动作参数修正提供初始数据;其次,将构建物理约束沙盒,并将原始指令作为输入数据进行影子仿真,生成状态演化时间序列,获取原指令执行的数据;再次,根据状态演化时间序列,计算各项安全指标与对应预设安全阈值之间的偏差,得到偏差序列,并计算综合风险系数,将综合风险系数用于后续数据修正,提高修正的准确率,为后续数据溯源提供归因的环节;进一步地,对原始指令中动作参数进行安全灵敏度分析,计算影响梯度,通过灵敏度反映微小扰动量对动作参数的影响,为后续数据溯源提供归因环节;进一步地,根据偏差序列、综合风险系数和影响梯度,迭代进行影子仿真,优化求解得到修正后的动作参数,提高安全性的同时,保证修正的准确性;进一步地,根据修正后的动作参数生成修正指令,下发执行并采集执行反馈,便于后续对修正的动作参数进行因果分析;最终,将原始指令、修正指令以及执行反馈按时间顺序组织成因果链,并对因果链进行存证,构建完整的存证链条,为事故溯源、责任认定及功能安全认证提供可审计的客观证据。
Smart Images

Figure CN122655842A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a method and system for correcting intelligent agent output and storing causal chain evidence based on physical constraints. Background Technology
[0002] Intelligent agents based on deep reinforcement learning, imitation learning and other methods have demonstrated excellent decision-making capabilities in complex control tasks and are increasingly being applied in safety-critical fields such as industrial robots, autonomous driving, and smart grids.
[0003] However, the training of existing intelligent agents usually relies on virtual simulation environments. Their output instructions often only pursue the optimization of a single task objective and lack systematic modeling and verification of rigid constraints in the real physical layer. Directly sending the original instructions generated by the intelligent agent to physical devices for execution can easily cause equipment to operate beyond its limits, structural overload, thermal runaway, or even catastrophic accidents. This leads to technical problems such as the untraceability of the intelligent agent's correction process, difficulty in determining responsibility, and insufficient evidentiary value. Summary of the Invention
[0004] This application provides a method and system for intelligent agent output correction and causal chain evidence preservation based on physical constraints, which addresses the technical problems in the prior art such as the untraceability of the intelligent agent correction process, the difficulty in determining responsibility, and the insufficient evidentiary value.
[0005] In view of the above problems, this application provides a method and system for intelligent agent output correction and causal chain evidence storage based on physical constraints.
[0006] In a first aspect, this application provides a method for correcting agent output and storing causal chain evidence based on physical constraints, the method comprising: Receive the raw instructions to be issued by the intelligent agent and parse the action parameters in the raw instructions; The original instructions are input into a physical constraint sandbox for shadow simulation to predict the evolution of the physical device after the original instructions are executed, and a state evolution time series is generated. Based on the state evolution time series, the deviation between each safety indicator and the corresponding preset safety threshold is calculated to obtain the deviation sequence, and the comprehensive risk coefficient is calculated based on the deviation sequence. Perform a safety sensitivity analysis on the action parameters in the original instruction, and calculate the gradient of the impact of changes in each action parameter on the safety index. Based on the deviation sequence, the comprehensive risk coefficient, and the influence gradient, with the goal of minimizing the action parameter offset and the constraint that the safety index does not exceed the preset safety threshold, the corrected action parameters are obtained through iterative optimization. Based on the modified action parameters, a correction instruction is generated, the correction instruction is issued and executed, and execution feedback is collected. The original instruction, the corrected instruction, and the execution feedback are organized into a causal chain in chronological order, and the causal chain is stored as evidence.
[0007] Secondly, this invention provides a system for correcting intelligent agent outputs and storing causal chain evidence based on physical constraints, the system comprising: The action parameter acquisition module is used to receive the original instructions to be issued by the intelligent agent and parse the action parameters in the original instructions; The simulation evolution module is used to input the original instructions into the physical constraint sandbox for shadow simulation, predict the evolution process of the physical device after the original instructions are executed, and generate a state evolution time series. The risk coefficient calculation module is used to calculate the deviation between each safety indicator and the corresponding preset safety threshold according to the state evolution time series, obtain the deviation sequence, and calculate the comprehensive risk coefficient based on the deviation sequence. The influence gradient calculation module is used to perform safety sensitivity analysis on the action parameters in the original instruction and calculate the influence gradient of the change of each action parameter on the safety index. The iterative optimization module is used to obtain the corrected action parameters by iteratively optimizing the deviation sequence, the comprehensive risk coefficient and the influence gradient, with the objective of minimizing the action parameter offset and the constraint that the safety index does not exceed the preset safety threshold. The correction feedback module is used to generate correction instructions based on the corrected action parameters, issue the correction instructions for execution, and collect execution feedback. The feedback and evidence storage module is used to organize the original instruction, the correction instruction, and the execution feedback into a causal chain in chronological order, and to store the causal chain as evidence.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application first receives the original instructions to be issued by the intelligent agent, parses the action parameters, and provides initial data for subsequent action parameter correction. Second, it constructs a physical constraint sandbox and uses the original instructions as input data for shadow simulation, generating a state evolution time series to obtain data on the execution of the original instructions. Third, based on the state evolution time series, it calculates the deviation between various safety indicators and their corresponding preset safety thresholds, obtaining a deviation sequence, and calculates a comprehensive risk coefficient. This comprehensive risk coefficient is used for subsequent data correction to improve the accuracy of the correction and provide an attribution step for subsequent data tracing. Furthermore, it performs a safety sensitivity analysis on the action parameters in the original instructions, calculates the influence gradient, and uses the sensitivity... This process reflects the impact of minute disturbances on action parameters, providing an attribution step for subsequent data tracing. Furthermore, based on the deviation sequence, comprehensive risk coefficient, and influence gradient, iterative shadow simulation is performed to optimize the solution and obtain the corrected action parameters, improving safety while ensuring the accuracy of the correction. Further, corrected instructions are generated based on the corrected action parameters, issued for execution, and execution feedback is collected, facilitating subsequent causal analysis of the corrected action parameters. Finally, the original instructions, corrected instructions, and execution feedback are organized into a causal chain in chronological order, and the causal chain is documented to construct a complete evidence chain, providing auditable and objective evidence for accident tracing, liability determination, and functional safety certification. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating the physical constraint-based intelligent agent output correction and causal chain evidence preservation method provided in this application. Figure 2 This is a schematic diagram of the structure of the physical constraint-based intelligent agent output correction and causal chain evidence storage system provided in this application.
[0010] In the attached diagram, the components represented by each number are as follows: The module includes: motion parameter acquisition module 11, simulation evolution module 12, risk coefficient calculation module 13, influence gradient calculation module 14, iterative optimization module 15, correction feedback module 16, and feedback storage module 17. Detailed Implementation
[0011] This application provides a method for correcting intelligent agent output and storing causal chain evidence based on physical constraints, which specifically solves the technical problems in the prior art such as the untraceability of the intelligent agent correction process, the difficulty in determining responsibility, and the insufficient evidentiary value.
[0012] The present invention will now be described in detail with reference to the accompanying drawings.
[0013] Example 1, as Figure 1 As shown, this application provides a method for intelligent agent output correction and causal chain evidence preservation based on physical constraints, the method comprising: S10: Receive the original instruction to be issued by the intelligent agent and parse the action parameters in the original instruction; In this embodiment, the original instructions generated by the intelligent agent are first captured in real time through the instruction receiving interface. Key action parameters are then extracted from the instructions according to the instruction parsing protocol.
[0014] Step S10 in the method provided in this embodiment of the invention includes: The original instructions to be issued generated by the intelligent agent are obtained through the instruction receiving interface. The original instructions include the target device identifier, action type and action parameter set. The action parameter set includes at least one of the following: target value, execution speed, duration and execution path. According to the preset instruction parsing protocol, action parameters are extracted from the original instruction and converted into a unified data structure.
[0015] In this embodiment, the instruction receiving interface is a software communication endpoint deployed at the edge or center of the entity; the target device identifier is an identification code used to specify the physical device that the original instruction intends to control; the action type is an abstract classification of the operation performed by the physical device, used to distinguish different control modes; the action parameter set is an ordered combination of a set of quantifiable control variables associated with the action type.
[0016] In the action parameter set, the target value is the final physical quantity value expected to be achieved after the action is executed; the execution speed is the rate of change when the current state transitions to the target value; the duration is the length of time taken to execute the action, used to limit timeliness or to plan periodic actions; and the execution path is the trajectory from the starting point to the target point in the state space.
[0017] Specifically, the previously configured instruction receiving interface is first invoked. After the agent issues a decision, the original instruction is captured and the original instruction is structured and parsed to extract the target device identifier field, which indicates the specific device to which the instruction should be sent; the action type field and the action parameter set field; and then the instruction is stored.
[0018] Secondly, after obtaining the raw instructions, a pre-configured instruction parsing protocol is loaded first. This protocol may include certain mapping rules; through the field mapping table in the protocol, specific key-value pairs are extracted from the raw instruction data, and their values are assigned to internal variables.
[0019] Assume the predefined instruction parsing protocol includes the following mapping rules: string fields map to the string type of internal variables; V, T, and S under a field map to the execution speed, duration, and execution path of the internal variables, respectively. Based on the instruction parsing protocol, the corresponding key-value pairs are extracted, data type conversion is performed, and then the values are filled into predefined message instances and converted into a unified data structure. The predefined instruction parsing protocol consists of predefined and embedded syntax and semantic mapping rules within the system.
[0020] In this embodiment, the original instructions to be issued generated by the intelligent agent are directly obtained through the instruction receiving interface, reducing information omissions and misalignments. Then, according to the preset instruction parsing protocol, action parameters are extracted from the original instructions and converted into a unified data structure, providing a sufficient data foundation for subsequent data analysis.
[0021] S20: Input the original instructions into the physical constraint sandbox for shadow simulation, predict the evolution process of the physical device after the original instructions are executed, and generate a state evolution time series. In this embodiment, the physical constraint sandbox is a simulation environment that operates independently of real physical devices. It integrates physical law models and dynamic characteristic models of specific devices, and does not send any control signals to real devices. Therefore, it is called shadow simulation. Shadow simulation is a process of using the physical sandbox to rehearse or test the potential consequences of instructions before they are issued and executed.
[0022] Specifically, after obtaining the action parameters of the original command, the acquired action parameters are loaded into the physical constraint sandbox. Using the real-time state of the current physical device as the initial condition, and based on the collaborative work of the internal general physics engine and domain-specific model, the action parameters are used as control inputs to perform time-integrated calculations. Finally, a list of state vectors is output as a state evolution time series.
[0023] Step S20 in the method provided in this embodiment of the invention includes: A physical constraint sandbox is constructed, which includes a general physics engine, a domain-specific model, and a parameter configuration layer. The general physics engine is used to simulate at least one physical process among rigid body motion, thermodynamic change, or fluid flow. The domain-specific model is used to simulate the dynamic response characteristics, material properties, or environmental change laws of physical devices. The parameter configuration layer is used to set initial conditions. The parsed action parameters are input into the physical constraint sandbox, and the physical constraint sandbox configures the current physical device's position parameters, state parameters, and environmental parameters through the parameter configuration layer as the current state; The physical constraint sandbox uses the current state as the initial condition and simulates the state changes over a future period of time according to the action parameters to generate a state evolution time series.
[0024] In this embodiment, the physical constraint sandbox is first initialized and constructed. During the construction process, three core components are loaded and instantiated: a general physics engine, a domain-specific model, and a parameter configuration layer. The general physics engine can be constructed based on rigid body kinematics or physical principles, and internally implements the mass, inertia, collision detection, and joint constraint solution of the rigid body. The domain-specific model can be a dedicated model file loaded internally by the system. The domain-specific model characterizes the individual characteristics of the device, such as the torque-speed characteristic curve of the motor, the elastic deformation coefficient of the robotic arm link, and the open-circuit voltage-state-of-charge relationship table of the battery, so that the simulation results approximate the real device file. The parameter configuration layer can be obtained through instantiation. After the three components are constructed, the initialized physical constraint sandbox is obtained.
[0025] Specifically, the physical constraint sandbox is a virtual simulation environment that operates independently of real physical devices and is constrained by physical laws; the general physics engine is a low-level simulation core module applicable to physical phenomena, encapsulating universal physical law mathematical models and providing the skeleton of physical simulation; the domain-specific model is a customized mathematical model built according to a specific type of physical device or specific environmental conditions; the parameter configuration layer is a special module in the sandbox, responsible for receiving and setting the initial condition parameters of the simulation before each simulation starts.
[0026] Initial condition parameters are configured uniformly through a parameter configuration layer. Based on the current state vector and action parameters, the general physics engine solves Newtonian mechanics, heat conduction, or fluid dynamics equations, outputting changes in fundamental physical quantities. These fundamental physical quantities, output by the general physics engine, are then passed as input to the domain-specific model. This model is modified according to the device's unique dynamic response characteristics to calculate device-specific state changes. The general physics engine and the domain-specific model work collaboratively through sequential solving; within each time step, the two models execute sequentially and exchange data, avoiding the difficulties of solving complex simultaneous equations while maintaining the causal consistency of the physical process.
[0027] Secondly, the parameters are configured through the parameter configuration layer to obtain the initial condition parameters of the physical constraint sandbox, including the position parameters, state parameters, and environmental parameters of the physical device. Among them, the current state is the value of the actual situation of the physical device in the real world when the shadow simulation starts; the position parameter is the variable of the physical device's pose in space; the state parameter is the variable describing the kinematic or dynamic state of the physical device; and the environmental parameter is the variable describing the external environmental conditions of the physical device.
[0028] Specifically, a status read request is sent to the target physical device via a real-time data acquisition bus to obtain the device's actual sensor data at the current moment. Taking a drone as an example, the angle values of each joint encoder are read to obtain position parameters; the rotational speed feedback values of each joint servo driver are read to obtain status parameters; and the temperature sensor values installed at the joints are read to obtain status parameters.
[0029] Then, the parameter configuration layer interface of the physical constraint sandbox is called, and the read real data is used as initial condition parameters and passed to the state variables inside the sandbox. The parameter configuration layer fills these values into the initial state structure of the sandbox. After configuration, it is ensured that the internal state of the sandbox is consistent with the current state of the real physical device. Subsequently, the action parameters are input into the physical constraint sandbox.
[0030] Secondly, the physical constraint sandbox uses the current conditions as the initial conditions, and then starts the simulation. The simulation duration is adaptively determined according to the action type. The time step is set according to the dynamic characteristics of the system. After the simulation loop starts, at each time step, the current state is obtained based on the current state and action parameters, and the corresponding time is recorded. After the simulation ends, a state evolution time series with timestamps is output. The state evolution time series is a state vector recorded at a fixed sampling interval from the initial time to the predicted time domain endpoint.
[0031] For example, taking a drone acceleration command as an example, the sequence will include the drone's position (X,Y,Z) and speed (Vx,Vy,Vz) and other motion parameters for each millisecond within the next 5 seconds.
[0032] In this embodiment, a physical constraint sandbox is first constructed based on three key components: a general physics engine, a domain-specific model, and a parameter configuration layer, to obtain the basic structure of the simulation. Then, the current state of the physical device is obtained to ensure the realism and accuracy of the simulation evolution. Subsequently, the analyzed action parameters are input into the physical constraint sandbox for simulation deduction to simulate the state changes over a period of time in the future, providing a data foundation and reference for subsequent comparative analysis.
[0033] S30: Based on the state evolution time series, calculate the deviation between each safety indicator and the corresponding preset safety threshold to obtain the deviation sequence, and calculate the comprehensive risk coefficient based on the deviation sequence; In this embodiment, the safety index is a specific physical quantity extracted from the state evolution time series to quantify the safe operation boundary of the equipment; the preset safety threshold is an upper or lower limit value of the physical quantity predefined by the equipment manufacturer, industry safety standards, or operation and maintenance experts; the deviation sequence is a numerical sequence obtained by comparing the safety index value of each time step with its corresponding threshold and arranging them in chronological order; and the comprehensive risk coefficient represents the overall risk level faced by the execution of the original instruction.
[0034] Specifically, the process iterates through each time step in the state evolution time series and extracts predefined safety indicators. For each safety indicator, a corresponding preset safety threshold is obtained. Then, the difference between the indicator value and the threshold or the excess ratio at each time step is calculated, forming a deviation sequence arranged by time. Based on the deviation sequence, the safety risk assessment model outputs a comprehensive risk coefficient.
[0035] Step S30 in the method provided in this embodiment of the invention includes: The safety index for each time step is extracted from the state evolution time series, wherein the safety index includes at least one of the following: equipment movement speed, position distance, temperature, pressure, and power. For each security indicator, obtain a pre-determined preset security threshold; Calculate the difference between the safety index and the corresponding preset safety threshold at each time step to obtain the deviation sequence; The deviation sequence is input into a pre-trained safety risk assessment model, and the comprehensive risk coefficient is output.
[0036] In this embodiment, the state evolution time series is first obtained, and then the corresponding speed, distance, temperature, pressure and power values for each time step of the state evolution time series are read. The data are then integrated in chronological order to obtain the safety index for each time step.
[0037] Specifically, the time step is the time interval between two adjacent state updates during the shadow simulation process. The size of the time step is determined by the trade-off between simulation accuracy and computational efficiency, with a typical value range of 1 millisecond to 100 milliseconds; the device speed is the rate at which the device or its key components move in space, and for rotating components, it can also refer to the angular velocity; the position distance is the distance between the current position of the device and the restricted area; and the temperature is the current temperature value of the key components of the device.
[0038] Secondly, the preset safety threshold can be set according to the technical specifications provided by the equipment manufacturer, industry safety standards, physical constraints of the operating environment, or the damage limit obtained through experimental calibration. When in use, the threshold data corresponding to the current safety indicator can be read from the system configuration database, local cache, or remote configuration center.
[0039] The system iterates through each security metric name, searching the configuration storage using the security metric as the retrieval criterion. Upon successful search, the system associates the preset security threshold with the corresponding security metric and stores the information.
[0040] Next, based on the safety indicators and the corresponding preset safety thresholds, the difference between the safety indicator and the corresponding preset safety threshold at each time step is calculated.
[0041] For upper limit thresholds, the difference calculation formula is: Deviation = Safety Index - Upper Limit Value. When the index value is less than or equal to the upper limit, the deviation is negative or zero, indicating safety; when the safety index is greater than the upper limit, the deviation is positive, indicating exceeding the limit. For lower limit thresholds, the difference calculation formula is: Deviation = Lower Limit Value - Safety Index. When the safety index is greater than or equal to the lower limit, the deviation is negative or zero, indicating safety; when the safety index is less than the lower limit, the deviation is positive, indicating exceeding the limit. For interval thresholds, the deviations to the upper and lower bounds can be calculated separately. Then, the deviation sequences are generated by arranging the deviation values calculated at all time steps in chronological order.
[0042] Finally, the obtained deviation sequence is preprocessed by data cleaning and outlier removal; then the preprocessed deviation sequence is input into the pre-trained safety risk assessment model, where matrix multiplication and nonlinear activation are performed, and finally the corresponding comprehensive risk coefficient is output.
[0043] A comprehensive risk coefficient of 0 indicates no risk, with all indicators within the threshold and no dangerous trend; a comprehensive risk coefficient between 0 and 1 indicates low to medium risk; and a comprehensive risk coefficient above 1 indicates high risk. The pre-trained safety risk assessment model is trained using machine learning methods and is a function mapper that maps the deviation sequence to the comprehensive risk coefficient.
[0044] For example, the deviation sequence of the drone, from t=1s to t=3s, a total of 200 time steps, each time step containing three values: speed deviation, temperature deviation, and altitude deviation, is input into the safety risk assessment model. After calculation, the comprehensive risk coefficient is output as 1.35.
[0045] In step S30 of the method provided in this embodiment of the invention, the process of constructing the security risk assessment model includes: Multiple sets of historical deviation sequences are collected to form a training dataset, and the safety risk level label corresponding to each set of historical deviation sequences in the training dataset is obtained to form a supervision label set; A security risk assessment model is constructed using neural networks; Using the training dataset as input and the supervision label set as supervision signal, the security risk assessment model is trained in a supervised manner until the verification convergence is achieved, thus obtaining the trained security risk assessment model.
[0046] In this embodiment, the historical deviation sequence is a sample of deviation sequences calculated from the state evolution time series generated by shadow simulation during historical operation; the training dataset is a collection of multiple historical deviation sequences; and the safety risk level label is a quantified value of the true risk level, either manually labeled or derived from rules for each set of historical deviation sequences.
[0047] First, deviation sequences were collected from shadow simulations of 5000 agent commands recorded over the past three months. Each deviation sequence contained 200 time steps, with each time step recording the corresponding safety indicator value. These 5000 deviation sequences were used as a training dataset. Subsequently, a corresponding safety risk level label was assigned to each deviation sequence.
[0048] Safety risk level labels can be independently generated by several domain experts or professional labelers who review historical operational data. Based on the actual consequences of accidents, historical deviation sequences after historical operations are retrospectively labeled. Risk levels are divided into 1-5 according to industry safety standards, where Level 1 is no risk, Level 2 is very low risk, Level 3 is moderate risk, Level 4 is high risk, and Level 5 is severe risk. The average score from three experts is used as the safety risk level label. After labeling, deviation sequences are paired with labels to form a training dataset and a supervision label set.
[0049] For example, 10,000 sets of historical deviation sequences of drone flights are collected, with each set containing 200 time steps, and the corresponding safety indicators are recorded at each step. At the same time, a training dataset and a label set are formed based on the final results of each flight and the battery health status detection data after the flight.
[0050] Secondly, the safety risk assessment model is a function mapper with a neural network as its main architecture. The input dimension matches the representation of the deviation sequence, and the output dimension is 1.
[0051] Based on the deviation sequence, a neural network architecture is selected to construct a security risk assessment model. It is assumed that each sample in the historical deviation sequence has a time step of 200, 4 features, and a 3-dimensional dimension. Given the moderate sequence length and temporal dependencies, a Long Short-Term Memory (LSTM) network is chosen as the basic architecture to capture long-term temporal dependencies and mitigate the gradient vanishing problem.
[0052] The first layer is the input layer, used to receive time-series data; the second layer is an LSTM layer with 64 hidden units, and it is set to output only the hidden state of the last time step; the third layer is a fully connected layer with 32 neurons and the ReLU activation function, used to extract higher-order features; the fourth layer is a dropout layer with a dropout rate of 0.2, used to prevent overfitting; the fifth layer is the output layer with 1 neuron, no activation function, and directly outputs the comprehensive risk coefficient.
[0053] Furthermore, the training dataset and the supervised label set are divided into training set, validation set and test set in a ratio of 7:2:1. The division adopts stratified sampling to ensure that the samples of various risk levels are distributed consistently in each subset.
[0054] The optimizer is Adam, with an initial learning rate of 0.001 and a mean squared error (MSE) loss function. In each training epoch, the training set is randomly divided into multiple mini-batches of 32 samples each. For each mini-batch, the bias sequence tensor is input into the neural network model, and the predicted risk coefficient is calculated through forward propagation. Then, the loss function value is calculated, using the MSE loss to calculate the loss between the true label and the prediction error. Subsequently, the gradient of the loss with respect to each model parameter is automatically calculated through backpropagation, and the optimizer is invoked to update the parameters. After completing a batch, the model performance is evaluated on the validation set. Training continues until the validation loss fails to decrease for 10 consecutive batches, triggering early stopping, and the model parameters with the lowest validation loss are saved as the final model. Finally, the trained model is evaluated on the test set to confirm that its generalization performance meets the requirements, such as a mean absolute error (MAE) of less than 0.1 on the test set. After evaluation, the model is exported and deployed to the safety risk coefficient calculation module for subsequent real-time inference.
[0055] For example, an LSTM model was trained using 10,000 sets of drone deviation sequences. The training process consisted of 200 batches, and the loss was minimized in the 145th batch, triggering early stopping. On the test set, the mean absolute error was 0.067, and the Pearson correlation coefficient was 0.93. The model was saved and deployed to edge computing nodes. For subsequent input deviation sequences, the model could output a comprehensive risk coefficient within 2 seconds for real-time decision correction.
[0056] In this embodiment, a safety index for each time step is extracted from the state evolution time series. Then, the difference between the safety index and the corresponding preset safety threshold is calculated to obtain a deviation sequence for subsequent evaluation. The two preset safety thresholds are allocated from the system's historical files to provide matching raw comparison data. Subsequently, a neural network structure is used to construct and train a safety risk assessment model. The deviation sequence is used as input data for model calculation, and a comprehensive risk coefficient is output. This safety risk model improves the accuracy of risk assessment and enables timely response to risk assessments.
[0057] S40: Perform a safety sensitivity analysis on the action parameters in the original instruction, and calculate the gradient of the impact of the change of each action parameter on the safety index; In this embodiment, safety sensitivity analysis is a method to quantitatively assess the strength of causal relationships by fine-tuning the input variables and observing the magnitude of changes in the output variables; the influence gradient is a matrix or vector whose elements represent the degree and direction of the influence of a change in a certain action parameter on a certain safety indicator.
[0058] Specifically, keeping other parameters constant, a small positive perturbation is applied to the action parameters to be analyzed. The perturbed parameters are then re-input into the physical constraint sandbox for a second shadow simulation to obtain the perturbed safety index. Next, the sensitivity coefficient between the original safety index and the perturbed safety index is calculated. Sensitivity analysis is performed on all action parameters and all safety indices, and the results are finally combined into an influence gradient matrix.
[0059] Step S40 in the method provided in this embodiment of the invention includes: For each action parameter, the current value of the action parameter is multiplied by a preset small offset ratio to obtain a small perturbation amount; The small perturbation is applied based on the current value of the action parameter to generate the perturbed command; The perturbed command is input into the physical constraint sandbox for shadow simulation to obtain the perturbed safety index value; Calculate the difference between the disturbed safety index value and the original safety index, and divide the difference by the range of the safety index to obtain the dimensionless safety index change rate. Divide the minute disturbance by the range of the motion parameter to obtain the dimensionless rate of change of the motion parameter; Divide the rate of change of the safety index by the rate of change of the action parameter to obtain the sensitivity coefficient of the action parameter to the safety index; The sensitivity coefficients of all action parameters to all safety indicators are combined into an influence gradient matrix, which serves as the influence gradient.
[0060] In this embodiment, the preset micro-offset ratio is a predetermined positive number much less than 1, used to apply a micro-perturbation to the action parameters. The ratio needs to be small enough to ensure that the perturbed parameters remain within a physically reasonable range and do not change the basic nature of the command; at the same time, it needs to overcome numerical noise in the simulation and ensure the numerical stability of the sensitivity calculation. A typical value is 0.005; the micro-perturbation amount is the numerical change obtained by multiplying the current value of the action parameters by the preset micro-offset ratio.
[0061] First, all motion parameters to be analyzed are obtained from the motion parameter set. Then, a pre-configured small offset ratio, such as 0.005, is obtained. Subsequently, for each set of motion parameters, the corresponding small perturbation is calculated, which is obtained by multiplying the corresponding preset small offset ratio.
[0062] For example, for a target horizontal velocity of 15 m / s, the calculated small perturbation is 15 × 0.005 = 0.075. If the current value of a certain motion parameter is zero, to avoid zero multiplied by a proportional value still being zero, an absolute small offset, such as 0.001, is used. Several corresponding small perturbations are obtained for all motion parameters.
[0063] Next, the calculated minute perturbation is added to the corresponding action parameters to obtain the perturbed command. Several perturbed commands are generated by adding the minute perturbation to the corresponding action parameters. All generated perturbed commands are then integrated and sequentially input into the physical constraint sandbox for simulation. For example, for the target horizontal velocity parameter in the UAV command, the original value is 15 m / s, and the minute perturbation is 0.075 m / s, then the perturbed velocity is 15.075 m / s, thus generating a perturbed command.
[0064] Next, each perturbation command is sequentially input into the previously constructed physical constraint sandbox. For each perturbation command, the real-time state of the current physical device is reread through the parameter configuration layer. Then, using the perturbation-induced action as the control input, the simulation is executed with the exact same duration and time step settings as the original command simulation, generating a perturbation-induced state evolution time series. Subsequently, each safety index value is extracted from the state evolution time series.
[0065] Furthermore, the difference between the safety index extracted in the second step and the safety index obtained from the original command simulation is calculated. Then, the difference is divided by the range of the safety index to obtain the dimensionless rate of change of the safety index. The range of the safety index is the range of values that the safety index can take physically, that is, the difference between the maximum and minimum values. For example, the range of joint angular velocity can be 0-120, so the range is 120.
[0066] Furthermore, the calculated minute disturbance is divided by the range of the motion parameter to obtain the dimensionless rate of change of the motion parameter, where the range of the motion parameter is the difference between the maximum and minimum values of the motion parameter.
[0067] Furthermore, dividing the dimensionless rate of change of the safety index by the corresponding dimensionless rate of change of the motion parameter yields the sensitivity coefficient of the motion parameter to the safety index. The sensitivity coefficient quantifies the relative change in the safety index when the motion parameter changes by a unit relative amount. A positive sensitivity coefficient indicates that increasing the motion parameter will lead to an increase in the safety index and an increase in risk; a negative coefficient indicates that increasing the motion parameter will lead to a decrease in the safety index and a decrease in risk, meaning that the parameter is negatively correlated with risk.
[0068] Ultimately, the gradient matrix is a two-dimensional matrix, and the matrix elements are the calculated sensitivity coefficients.
[0069] Specifically, all calculated sensitivity coefficients are inserted into the corresponding positions in the matrix according to a uniform index order, with motion parameters as rows and safety indicators as columns. An influence gradient matrix is then created. Taking the robotic arm as an example, there are two motion parameters and two safety indicators; a 2×2 influence gradient matrix is constructed.
[0070] In this embodiment, a small perturbation is calculated for each action parameter. This small perturbation is then added to the original data to generate a perturbed command. By calculating the perturbed command, the degree of impact of the perturbation on the command is analyzed. Through shadow simulation, the perturbed safety index value is obtained for comparison. Accurate evaluation of the perturbed image through perturbation simulation facilitates subsequent corrections and improves simulation accuracy. Subsequently, the rate of change of the safety index and the rate of change of the action parameters after perturbation are calculated, and the ratio is used to obtain the sensitivity coefficient. Finally, an influence gradient matrix is constructed based on the sensitivity coefficient as the influence gradient.
[0071] S50: Based on the deviation sequence, the comprehensive risk coefficient, and the influence gradient, with the goal of minimizing the action parameter offset and the constraint that the safety index does not exceed the preset safety threshold, the corrected action parameters are obtained through iterative optimization. In this embodiment, the motion parameter offset is the amount of change in the corrected motion parameter relative to the motion parameter in the original instruction.
[0072] Specifically, the deviation sequence, comprehensive risk coefficient, and influence gradient are used as three analytical dimensions. With the goal of minimizing the deviation of action parameters, and under the constraint that the safety index does not exceed the preset safety threshold, the action parameters are recalculated to obtain new candidate parameters. The new candidate parameters are then input into the physical constraint sandbox for a new round of shadow simulation to update the deviation sequence. In this way, the action parameters are iteratively optimized to obtain the optimal corrected action parameters.
[0073] Step S50 in the method provided in this embodiment of the invention includes: Use the action parameters in the original instruction as the current action parameters; Repeat the following iterative steps until all security metrics meet the preset security threshold: Based on the influence gradient, determine the adjustment direction of each action parameter, and calculate the adjustment step size of each action parameter based on the deviation sequence and the comprehensive risk coefficient. The current motion parameters are corrected according to the adjustment direction and the adjustment step size to obtain new motion parameters; The new action parameters are input into the physical constraint sandbox for shadow simulation, and the state evolution time series and deviation series are updated. If the updated deviation sequence shows that all safety indicators meet the preset safety threshold, the iteration is terminated and the new action parameters are output as the corrected action parameters. Otherwise, the new action parameters are used as the current action parameters, and the next iteration continues.
[0074] In this embodiment, the action parameters of the original instruction are used as the action parameters of the current simulation, and then several rounds of shadow simulation are repeated until the stopping condition is met.
[0075] First, based on the previously constructed influence gradient, the adjustment direction of the action parameters is determined. The adjustment direction is determined by the sign of the corresponding sensitivity coefficient in the influence gradient matrix: if the sensitivity coefficient is positive, increasing the parameter will increase the safety index value, so the adjustment direction is set to decrease the parameter; if the sensitivity coefficient is negative, increasing the parameter will decrease the safety index value, so the adjustment direction is set to increase the parameter.
[0076] The step size is calculated based on the deviation sequence and the comprehensive risk coefficient. The larger the deviation sequence and the comprehensive risk coefficient, the larger the step size. The adjustment step size is the absolute value by which each action parameter changes along the adjustment direction in the current iteration.
[0077] Secondly, based on the adjustment step size, corrections are made in the corresponding adjustment direction. For parameters that need to be decreased, the new motion parameter = the current motion parameter - one step size; for parameters that need to be increased, the new motion parameter = the current motion parameter + one step size.
[0078] Furthermore, the corrected action parameters are used as input data for the current iteration, input into the physical constraint sandbox for shadow simulation, perform one round of shadow simulation, and update the state evolution time series and deviation series accordingly.
[0079] Finally, the action parameters obtained after one round of correction will be used as the current action parameters for the next iteration, or, if the convergence condition is met, as the final corrected action parameters output.
[0080] Specifically, if the corrected deviation sequence shows that all safety indicators are greater than or equal to the preset safety threshold, the correction requirements are considered met, and the corrected data is within a safe range. The iteration can then be terminated immediately, and the data from the current round can be output as the correction result. If the corrected deviation sequence shows that all safety indicators are less than the preset safety threshold, the safety requirements are considered not met, and correction can continue until convergence or the maximum number of iterations is reached.
[0081] Step S50 of the method provided in this embodiment of the invention, which determines the adjustment direction of each action parameter based on the influence gradient and calculates the adjustment step size of each action parameter based on the deviation sequence and the comprehensive risk coefficient, includes: For each safety indicator that exceeds the preset safety threshold, the sensitivity coefficient of the safety indicator to each action parameter is extracted from the influence gradient; The adjustment direction of the motion parameter is determined according to the sign of the sensitivity coefficient: if the sensitivity coefficient is positive, the adjustment direction is to decrease the motion parameter; if the sensitivity coefficient is negative, the adjustment direction is to increase the motion parameter. Calculate the ratio of the excess amount of each safety indicator to the corresponding preset safety threshold as the absolute deviation ratio, and use the absolute value of the sensitivity coefficient as the weight to calculate the weighted sum of the absolute deviation ratios of each safety indicator to obtain the weighted deviation ratio. Add one to the aforementioned comprehensive risk coefficient to obtain the risk scaling factor; The adjustment step size for each action parameter is obtained by multiplying the preset base step size by the weighted deviation ratio and then by the risk scaling factor.
[0082] In this embodiment, firstly, a safety index is extracted from the deviation sequence generated by the shadow simulation of the current iteration from the influence gradient. This safety index has at least one time step where its value is greater than the upper threshold or less than the lower threshold. Then, the sensitivity coefficients of the corresponding action parameters are obtained from a column of the influence gradient matrix of the safety index. For example, if the out-of-limit index of a drone is battery temperature, the sensitivity coefficients are: target speed 0.8, climb rate 0.3, and yaw rate -0.1.
[0083] Secondly, the sign of the sensitivity coefficient is used as the adjustment direction for the action parameters. This adjustment direction determines whether each action parameter should be increased or decreased to bring the safety index below the threshold. Since parameters change in the same direction as risk, parameters need to be decreased to reduce risk; therefore, the adjustment direction for positive sensitivity coefficients is defined as decreasing. For negative sensitivity coefficients, parameters need to be increased to reduce risk; therefore, the adjustment direction for negative sensitivity coefficients is defined as increasing. For parameters with a sensitivity coefficient of zero, the adjustment direction remains unchanged.
[0084] Next, subtract the preset safety threshold from the safety index, and then divide by the preset safety threshold to obtain the absolute deviation ratio. Using the absolute values of the corresponding extracted sensitivity coefficients as weighting coefficients, a weighted sum is performed to obtain the weighted deviation ratio. Weighted deviation ratio = W1×L1 + W2×L2 + W3×L3... For example, if the drone's battery temperature exceeds the limit by 0.08 and its horizontal speed exceeds the limit by 0.05, and its speed sensitivity to temperature is 0.8 and its speed sensitivity to temperature is 1.2, then the weighted deviation ratio for speed = 0.8×0.08 + 1.2×0.05 = 0.124.
[0085] Further, the weighted deviation ratio is added to the comprehensive risk coefficient to obtain the risk scaling factor. Then, the adjustment step size is obtained by combining the scaling risk factor with the motion parameters. Adjustment step size = base step size × weighted deviation ratio × risk scaling factor. The preset base step size is a pre-defined value related to the range of the motion parameters, typically between 0.05 and 0.2. The base step size represents the maximum relative adjustment allowed in each iteration under standard conditions. For example, if the robotic arm's risk coefficient is 0.85 and the risk scaling factor is 1.85, the adjustment step size = 0.1 × 0.85 × 1.85 ≈ 0.157.
[0086] In this embodiment, the action parameters in the original instruction are used as the current action parameters to determine the adjustment direction and calculate the adjustment step size of the action parameters. Adaptive updates are performed through adaptive adjustment step size calculation. New action parameters are obtained through correction. Shadow simulation is then performed to update the state evolution time series and deviation series. Multiple iterations are performed to obtain the corrected action parameters, thereby improving the accuracy and sensitivity of the correction and avoiding inaccurate step size, which could lead to correction errors.
[0087] S60: Generate a correction command based on the modified action parameters, issue the correction command for execution, and collect execution feedback; In this embodiment, the correction instruction is a control command that repackages the solved corrected action parameters according to the same communication protocol and data structure as the original instruction.
[0088] Specifically, the corrected action parameters obtained from iterative optimization are assembled into a corrected instruction according to the format of the original instruction. Then, the corrected instruction is sent to the controller of the physical device via the real-time communication bus through the instruction delivery module. Simultaneously, a feedback acquisition listener is activated to receive the actual data stream transmitted back from the sensors after the device executes the instruction.
[0089] In this embodiment of the application, the modified action parameters are executed, and the modified action parameters are collected by the feedback acquisition device to obtain the modified operation data, so as to provide sufficient analysis data for subsequent evidence preservation and attribution.
[0090] S70: Organize the original instruction, the modified instruction, and the execution feedback into a causal chain in chronological order, and store the causal chain.
[0091] In this embodiment of the application, the causal chain is a chain-like data structure based on the blockchain concept, in which each block contains a data hash value, timestamp, and hash pointer to the previous node for an event; evidence storage is to submit the key information of the causal chain to an immutable storage medium for subsequent compliance audits, accident tracing, liability determination, or model improvement operations.
[0092] Specifically, after the instruction is issued and the execution feedback is collected, evidence is stored. First, an initial causal chain is created according to the actual chronological order of events. Then, a node is generated for each key event in sequence; for example, node 1 corresponds to receiving the original instruction and storing the hash value of the instruction content. The hash value of the previous node is recorded when each node is generated, forming a chain structure. Finally, the hash value of the root node of the entire causal chain is uploaded to the chain, and the complete causal chain data is backed up to a trusted audit node.
[0093] Step S70 in the method provided in this embodiment of the invention includes: Create causal chain nodes in chronological order. Each node contains a timestamp, event type, event data hash value, and the hash value of the previous node, forming a chain structure. The event types include at least instruction reception event, shadow simulation start event, correction decision event, instruction issuance event, and execution feedback event. The detailed data generated by the shadow simulation of the physical constraint sandbox is stored in an off-chain distributed storage system, and the Merkle root hash of the detailed data is calculated. The Merkle root hash is then stored in the extended field of the corresponding node. The detailed data includes at least the original instruction, the correction instruction, and the execution feedback. The hash value of the root node of the causal chain is stored on the blockchain for evidence, and the complete causal chain data is backed up to a trusted audit node.
[0094] In this embodiment of the application, firstly, several causal chain nodes containing timestamps, event types, event data hash values, and the hash value of the previous node are created in chronological order to form an initial causal chain to be filled. The event types include at least five types: instruction receiving event, shadow simulation start event, correction decision event, instruction issuance event, and execution feedback event. Therefore, an initial causal chain containing at least five nodes can be created.
[0095] Next, the original instruction content is uploaded to an off-chain distributed storage system. Upon successful upload, a corresponding content identifier is obtained. Then, the original instruction content, along with other potentially related data, is used to construct a Merkle tree. For example, the original instruction can be split into multiple data blocks, the hash of each block can be calculated, and these blocks can be combined layer by layer upwards to obtain the Merkle tree root hash. The Merkle tree root hash is then stored in the extended field of the causal chain node. Simultaneously, locators stored off-chain can also be stored for subsequent retrieval.
[0096] Specifically, an off-chain distributed storage system is a large-capacity storage network independent of the main blockchain or causal chain; a Merkle tree is a binary tree structure where each leaf node is the hash value of a data block, and each non-leaf node is a combined hash of the hash values of its two child nodes. The resulting root hash uniquely represents the entire dataset and allows verification of whether a data block belongs to the dataset without downloading all the data; the Merkle tree root hash is the top-level hash value obtained after constructing a Merkle tree from a set of detailed data; the extension field is an optional field reserved in the causal chain node structure for storing additional metadata.
[0097] Finally, the root hash of the entire causal chain is calculated. All nodes are concatenated sequentially with their respective hash values into a long byte sequence, and then this sequence is hashed to obtain the final chain root hash, which uniquely identifies the state of the entire causal chain. The blockchain smart contract is then invoked, uploading the root hash as a parameter to the deployed evidence storage contract. The blockchain network performs consensus processing on the transaction, packaging it into a transaction hash that includes a timestamp.
[0098] Once generated, anyone can query the records on the blockchain. The complete causal chain data is then integrated and encrypted using a strong encryption algorithm, such as AES-256, before being transmitted to multiple trusted audit nodes for backup, thus completing the evidence preservation. Subsequent auditors can obtain the complete causal chain data from the trusted audit nodes, verify the consistency of the hashes within the chain, compare it with the root hash obtained from the chain, and then verify the integrity of the detailed off-chain data using the Merkle root hash in the extended fields.
[0099] In this embodiment, causal chain nodes are created in chronological order to construct an initial causal chain, providing a structural foundation for subsequent evidence preservation. Subsequently, the detailed data generated by the shadow simulation is stored in an off-chain distributed storage system, and the Merkle tree root hash is calculated and stored in the extended field of the corresponding node to facilitate subsequent retrieval and improve retrieval efficiency. Finally, the hash value of the root node of the causal chain is stored on the chain for evidence preservation, and the complete causal chain data is backed up to a trusted audit node, reducing audit costs and forming a complete evidence loop.
[0100] The above-described specific implementation methods achieve the following technical effects: In this embodiment, the original instructions to be issued by the intelligent agent are first obtained directly through the instruction receiving interface to reduce information omissions and misalignments. Then, according to the preset instruction parsing protocol, the action parameters are extracted from the original instructions and converted into a unified data structure to provide a sufficient data foundation for subsequent data analysis.
[0101] Secondly, based on the general physics engine, the domain-specific model, and the parameter configuration layer, a physical constraint sandbox is constructed to obtain the basic structure of the simulation. Then, the current state of the physical devices is obtained to ensure the realism and accuracy of the simulation evolution. Subsequently, the analyzed action parameters are input into the physical constraint sandbox to perform simulation deduction, simulating the state changes over a period of time in the future, providing a data foundation and reference for subsequent comparative analysis.
[0102] Next, safety indicators for each time step are extracted from the state evolution time series. Then, the difference between the safety indicator and the corresponding preset safety threshold is calculated to obtain a deviation sequence for subsequent evaluation. The two preset safety thresholds are allocated from the system's historical files to provide matching raw comparison data. Subsequently, a neural network structure is used to construct and train a safety risk assessment model. The deviation sequence is used as input data for model calculation, and the output is a comprehensive risk coefficient. This safety risk model improves the accuracy of risk assessment and enables timely response to risk assessments.
[0103] Furthermore, for each motion parameter, a corresponding small perturbation is calculated. This small perturbation is then added to the original data to generate the perturbed command. By calculating the perturbed command, the degree of impact of the perturbation on the command is analyzed. Through shadow simulation, the perturbed safety index value is obtained for comparison. Accurate evaluation of the perturbed image through perturbation simulation facilitates subsequent corrections and improves simulation accuracy. Subsequently, the rate of change of the safety index and the rate of change of the motion parameters after perturbation are calculated separately, and the ratio is calculated to obtain the sensitivity coefficient. Finally, an influence gradient matrix is constructed based on the sensitivity coefficient, serving as the influence gradient.
[0104] Furthermore, the action parameters in the original instruction are used as the current action parameters to determine the adjustment direction and calculate the adjustment step size of the action parameters. Adaptive updates are performed through adaptive adjustment step size calculation. New action parameters are obtained through correction. Shadow simulation is then performed to update the state evolution time series and deviation series. Multiple iterations are performed to obtain the corrected action parameters, thereby improving the accuracy and sensitivity of the correction and avoiding inaccurate step size, which could lead to correction errors.
[0105] Furthermore, the corrected action parameters are executed, and the corrected action parameters are collected through feedback acquisition devices to obtain corrected operational data, providing sufficient data for subsequent analysis and improving the accuracy of subsequent evidence preservation and attribution.
[0106] Finally, causal chain nodes are created in chronological order to construct the initial causal chain, providing a structural foundation for subsequent evidence preservation. Subsequently, the detailed data generated by the shadow simulation is stored in an off-chain distributed storage system, and the Merkle tree root hash is calculated and stored in the extended field of the corresponding node to facilitate subsequent retrieval and improve retrieval efficiency. Finally, the hash value of the root node of the causal chain is stored on the chain for evidence preservation, and the complete causal chain data is backed up to a trusted audit node, reducing audit costs and forming a complete evidence loop.
[0107] Example 2, as Figure 2 As shown, based on the same inventive concept as the physical constraint-based agent output correction and causal chain evidence preservation method provided in Embodiment 1, this embodiment of the invention also provides a physical constraint-based agent output correction and causal chain evidence preservation system, the system comprising: The action parameter acquisition module 11 is used to receive the original instruction to be issued by the intelligent agent and parse the action parameters in the original instruction; The simulation evolution module 12 is used to input the original instructions into the physical constraint sandbox for shadow simulation, predict the evolution process of the physical device after the original instructions are executed, and generate a state evolution time series. The risk coefficient calculation module 13 is used to calculate the deviation between each safety indicator and the corresponding preset safety threshold according to the state evolution time series, obtain the deviation sequence, and calculate the comprehensive risk coefficient based on the deviation sequence. The influence gradient calculation module 14 is used to perform safety sensitivity analysis on the action parameters in the original instruction and calculate the influence gradient of the change of each action parameter on the safety index. Iterative optimization module 15 is used to obtain the corrected action parameters by iterative optimization based on the deviation sequence, the comprehensive risk coefficient and the influence gradient, with the goal of minimizing the action parameter offset and the constraint that the safety index does not exceed the preset safety threshold. The correction feedback module 16 is used to generate correction instructions based on the corrected action parameters, issue the correction instructions for execution, and collect execution feedback. The feedback and evidence storage module 17 is used to organize the original instruction, the correction instruction, and the execution feedback into a causal chain in chronological order, and to store the causal chain as evidence.
[0108] In one embodiment, the action parameter acquisition module 11 is used for: The original instructions to be issued generated by the intelligent agent are obtained through the instruction receiving interface. The original instructions include the target device identifier, action type and action parameter set. The action parameter set includes at least one of the following: target value, execution speed, duration and execution path. According to the preset instruction parsing protocol, action parameters are extracted from the original instruction and converted into a unified data structure.
[0109] In one embodiment, the simulation evolution module 12 is used for: A physical constraint sandbox is constructed, which includes a general physics engine, a domain-specific model, and a parameter configuration layer. The general physics engine is used to simulate at least one physical process among rigid body motion, thermodynamic change, or fluid flow. The domain-specific model is used to simulate the dynamic response characteristics, material properties, or environmental change laws of physical devices. The parameter configuration layer is used to set initial conditions. The parsed action parameters are input into the physical constraint sandbox, and the physical constraint sandbox configures the current physical device's position parameters, state parameters, and environmental parameters through the parameter configuration layer as the current state; The physical constraint sandbox uses the current state as the initial condition and simulates the state changes over a future period of time according to the action parameters to generate a state evolution time series.
[0110] In one embodiment, the risk coefficient calculation module 13 is used for: The safety index for each time step is extracted from the state evolution time series, wherein the safety index includes at least one of the following: equipment movement speed, position distance, temperature, pressure, and power. For each security indicator, obtain a pre-determined preset security threshold; Calculate the difference between the safety index and the corresponding preset safety threshold at each time step to obtain the deviation sequence; The deviation sequence is input into a pre-trained safety risk assessment model, and the comprehensive risk coefficient is output.
[0111] The process of constructing the security risk assessment model includes: Multiple sets of historical deviation sequences are collected to form a training dataset, and the safety risk level label corresponding to each set of historical deviation sequences in the training dataset is obtained to form a supervision label set; A security risk assessment model is constructed using neural networks; Using the training dataset as input and the supervision label set as supervision signal, the security risk assessment model is trained in a supervised manner until the verification convergence is achieved, thus obtaining the trained security risk assessment model.
[0112] In one embodiment, the gradient calculation module 14 is affected for: For each action parameter, the current value of the action parameter is multiplied by a preset small offset ratio to obtain a small perturbation amount; The small perturbation is applied based on the current value of the action parameter to generate the perturbed command; The perturbed command is input into the physical constraint sandbox for shadow simulation to obtain the perturbed safety index value; Calculate the difference between the disturbed safety index value and the original safety index, and divide the difference by the range of the safety index to obtain the dimensionless safety index change rate. Divide the minute disturbance by the range of the motion parameter to obtain the dimensionless rate of change of the motion parameter; Divide the rate of change of the safety index by the rate of change of the action parameter to obtain the sensitivity coefficient of the action parameter to the safety index; The sensitivity coefficients of all action parameters to all safety indicators are combined into an influence gradient matrix, which serves as the influence gradient.
[0113] In one embodiment, the iterative optimization module 15 is used for: Use the action parameters in the original instruction as the current action parameters; Repeat the following iterative steps until all security metrics meet the preset security threshold: Based on the influence gradient, determine the adjustment direction of each action parameter, and calculate the adjustment step size of each action parameter based on the deviation sequence and the comprehensive risk coefficient. The current motion parameters are corrected according to the adjustment direction and the adjustment step size to obtain new motion parameters; The new action parameters are input into the physical constraint sandbox for shadow simulation, and the state evolution time series and deviation series are updated. If the updated deviation sequence shows that all safety indicators meet the preset safety threshold, the iteration is terminated and the new action parameters are output as the corrected action parameters. Otherwise, the new action parameters are used as the current action parameters, and the next iteration continues.
[0114] Specifically, based on the influence gradient, the adjustment direction of each action parameter is determined, and based on the deviation sequence and the comprehensive risk coefficient, the adjustment step size of each action parameter is calculated, including: For each safety indicator that exceeds the preset safety threshold, the sensitivity coefficient of the safety indicator to each action parameter is extracted from the influence gradient; The adjustment direction of the motion parameter is determined according to the sign of the sensitivity coefficient: if the sensitivity coefficient is positive, the adjustment direction is to decrease the motion parameter; if the sensitivity coefficient is negative, the adjustment direction is to increase the motion parameter. Calculate the ratio of the excess amount of each safety indicator to the corresponding preset safety threshold as the absolute deviation ratio, and use the absolute value of the sensitivity coefficient as the weight to calculate the weighted sum of the absolute deviation ratios of each safety indicator to obtain the weighted deviation ratio. Add one to the aforementioned comprehensive risk coefficient to obtain the risk scaling factor; The adjustment step size for each action parameter is obtained by multiplying the preset base step size by the weighted deviation ratio and then by the risk scaling factor.
[0115] In one embodiment, the feedback and evidence storage module 17 is used for: Create causal chain nodes in chronological order. Each node contains a timestamp, event type, event data hash value, and the hash value of the previous node, forming a chain structure. The event types include at least instruction reception event, shadow simulation start event, correction decision event, instruction issuance event, and execution feedback event. The detailed data generated by the shadow simulation of the physical constraint sandbox is stored in an off-chain distributed storage system, and the Merkle root hash of the detailed data is calculated. The Merkle root hash is then stored in the extended field of the corresponding node. The detailed data includes at least the original instruction, the correction instruction, and the execution feedback. The hash value of the root node of the causal chain is stored on the blockchain for evidence, and the complete causal chain data is backed up to a trusted audit node.
[0116] Compared to existing technologies, this application first obtains the original instructions to be issued generated by the intelligent agent directly through the instruction receiving interface, reducing information omissions and misalignments. Then, according to the preset instruction parsing protocol, it extracts action parameters from the original instructions and converts the action parameters into a unified data structure, providing a sufficient data foundation for subsequent data analysis.
[0117] Secondly, based on the general physics engine, the domain-specific model, and the parameter configuration layer, a physical constraint sandbox is constructed to obtain the basic structure of the simulation. Then, the current state of the physical devices is obtained to ensure the realism and accuracy of the simulation evolution. Subsequently, the analyzed action parameters are input into the physical constraint sandbox to perform simulation deduction, simulating the state changes over a period of time in the future, providing a data foundation and reference for subsequent comparative analysis.
[0118] Next, safety indicators for each time step are extracted from the state evolution time series. Then, the difference between the safety indicator and the corresponding preset safety threshold is calculated to obtain a deviation sequence for subsequent evaluation. The two preset safety thresholds are allocated from the system's historical files to provide matching raw comparison data. Subsequently, a neural network structure is used to construct and train a safety risk assessment model. The deviation sequence is used as input data for model calculation, and the output is a comprehensive risk coefficient. This safety risk model improves the accuracy of risk assessment and enables timely response to risk assessments.
[0119] Furthermore, for each motion parameter, a corresponding small perturbation is calculated. This small perturbation is then added to the original data to generate the perturbed command. By calculating the perturbed command, the degree of impact of the perturbation on the command is analyzed. Through shadow simulation, the perturbed safety index value is obtained for comparison. Accurate evaluation of the perturbed image through perturbation simulation facilitates subsequent corrections and improves simulation accuracy. Subsequently, the rate of change of the safety index and the rate of change of the motion parameters after perturbation are calculated separately, and the ratio is calculated to obtain the sensitivity coefficient. Finally, an influence gradient matrix is constructed based on the sensitivity coefficient, serving as the influence gradient.
[0120] Furthermore, the action parameters in the original instruction are used as the current action parameters to determine the adjustment direction and calculate the adjustment step size of the action parameters. Adaptive updates are performed through adaptive adjustment step size calculation. New action parameters are obtained through correction. Shadow simulation is then performed to update the state evolution time series and deviation series. Multiple iterations are performed to obtain the corrected action parameters, thereby improving the accuracy and sensitivity of the correction and avoiding inaccurate step size, which could lead to correction errors.
[0121] Furthermore, the corrected action parameters are executed, and the corrected action parameters are collected through feedback acquisition devices to obtain corrected operational data, providing sufficient data for subsequent analysis and improving the accuracy of subsequent evidence preservation and attribution.
[0122] Finally, causal chain nodes are created in chronological order to construct the initial causal chain, providing a structural foundation for subsequent evidence preservation. Subsequently, the detailed data generated by the shadow simulation is stored in an off-chain distributed storage system, and the Merkle tree root hash is calculated and stored in the extended field of the corresponding node to facilitate subsequent retrieval and improve retrieval efficiency. Finally, the hash value of the root node of the causal chain is stored on the chain for evidence preservation, and the complete causal chain data is backed up to a trusted audit node, reducing audit costs and forming a complete evidence loop.
Claims
1. A method for correcting agent output and storing causal chain evidence based on physical constraints, characterized in that, The method includes: Receive the raw instructions to be issued by the intelligent agent and parse the action parameters in the raw instructions; The original instructions are input into a physical constraint sandbox for shadow simulation to predict the evolution of the physical device after the original instructions are executed, and a state evolution time series is generated. Based on the state evolution time series, the deviation between each safety indicator and the corresponding preset safety threshold is calculated to obtain the deviation sequence, and the comprehensive risk coefficient is calculated based on the deviation sequence. Perform a safety sensitivity analysis on the action parameters in the original instruction, and calculate the gradient of the impact of changes in each action parameter on the safety index. Based on the deviation sequence, the comprehensive risk coefficient, and the influence gradient, with the goal of minimizing the action parameter offset and the constraint that the safety index does not exceed the preset safety threshold, the corrected action parameters are obtained through iterative optimization. Based on the modified action parameters, a correction instruction is generated, the correction instruction is issued and executed, and execution feedback is collected. The original instruction, the corrected instruction, and the execution feedback are organized into a causal chain in chronological order, and the causal chain is stored as evidence.
2. The method for correcting agent output and storing causal chain evidence based on physical constraints according to claim 1, characterized in that, Receive the raw instructions to be issued by the intelligent agent, and parse the action parameters in the raw instructions, including: The original instructions to be issued generated by the intelligent agent are obtained through the instruction receiving interface. The original instructions include the target device identifier, action type and action parameter set. The action parameter set includes at least one of the following: target value, execution speed, duration and execution path. According to the preset instruction parsing protocol, action parameters are extracted from the original instruction and converted into a unified data structure.
3. The method for correcting agent output and storing causal chain evidence based on physical constraints according to claim 1, characterized in that, The original instructions are input into a physical constraint sandbox for shadow simulation to predict the evolution of the physical device after the original instructions are executed, generating a state evolution time series, including: A physical constraint sandbox is constructed, which includes a general physics engine, a domain-specific model, and a parameter configuration layer. The general physics engine is used to simulate at least one physical process among rigid body motion, thermodynamic change, or fluid flow. The domain-specific model is used to simulate the dynamic response characteristics, material properties, or environmental change laws of physical devices. The parameter configuration layer is used to set initial conditions. The parsed action parameters are input into the physical constraint sandbox, and the physical constraint sandbox configures the current physical device's position parameters, state parameters, and environmental parameters through the parameter configuration layer as the current state; The physical constraint sandbox uses the current state as the initial condition and simulates the state changes over a future period of time according to the action parameters to generate a state evolution time series.
4. The method for correcting agent output and storing causal chain evidence based on physical constraints according to claim 1, characterized in that, Based on the state evolution time series, the deviations between each safety indicator and the corresponding preset safety thresholds are calculated to obtain a deviation sequence. A comprehensive risk coefficient is then calculated based on the deviation sequence, including: The safety index for each time step is extracted from the state evolution time series, wherein the safety index includes at least one of the following: equipment movement speed, position distance, temperature, pressure, and power. For each security indicator, obtain a pre-determined preset security threshold; Calculate the difference between the safety index and the corresponding preset safety threshold at each time step to obtain the deviation sequence; The deviation sequence is input into a pre-trained safety risk assessment model, and the comprehensive risk coefficient is output.
5. The method for correcting agent output and storing causal chain evidence based on physical constraints according to claim 4, characterized in that, The process of constructing a security risk assessment model includes: Multiple sets of historical deviation sequences are collected to form a training dataset, and the safety risk level label corresponding to each set of historical deviation sequences in the training dataset is obtained to form a supervision label set; A security risk assessment model is constructed using neural networks; Using the training dataset as input and the supervision label set as supervision signal, the security risk assessment model is trained in a supervised manner until the verification convergence is achieved, thus obtaining the trained security risk assessment model.
6. The method for correcting agent output and storing causal chain evidence based on physical constraints according to claim 1, characterized in that, A safety sensitivity analysis is performed on the action parameters in the original instruction, and the impact gradient of the change in each action parameter on the safety index is calculated, including: For each action parameter, the current value of the action parameter is multiplied by a preset small offset ratio to obtain a small perturbation amount; The small perturbation is applied based on the current value of the action parameter to generate the perturbed command; The perturbed command is input into the physical constraint sandbox for shadow simulation to obtain the perturbed safety index value; Calculate the difference between the disturbed safety index value and the original safety index, and divide the difference by the range of the safety index to obtain the dimensionless safety index change rate. Divide the minute disturbance by the range of the motion parameter to obtain the dimensionless rate of change of the motion parameter; Divide the rate of change of the safety index by the rate of change of the action parameter to obtain the sensitivity coefficient of the action parameter to the safety index; The sensitivity coefficients of all action parameters to all safety indicators are combined into an influence gradient matrix, which serves as the influence gradient.
7. The method for correcting agent output and storing causal chain evidence based on physical constraints according to claim 1, characterized in that, Based on the deviation sequence, the comprehensive risk coefficient, and the influence gradient, with the objective of minimizing the action parameter offset and the constraint that the safety index does not exceed the preset safety threshold, the corrected action parameters are obtained through iterative optimization, including: Use the action parameters in the original instruction as the current action parameters; Repeat the following iterative steps until all security metrics meet the preset security threshold: Based on the influence gradient, determine the adjustment direction of each action parameter, and calculate the adjustment step size of each action parameter based on the deviation sequence and the comprehensive risk coefficient. The current motion parameters are corrected according to the adjustment direction and the adjustment step size to obtain new motion parameters; The new action parameters are input into the physical constraint sandbox for shadow simulation, and the state evolution time series and deviation series are updated. If the updated deviation sequence shows that all safety indicators meet the preset safety threshold, the iteration is terminated and the new action parameters are output as the corrected action parameters. Otherwise, the new action parameters are used as the current action parameters, and the next iteration continues.
8. The method for correcting agent output and storing causal chain evidence based on physical constraints according to claim 7, characterized in that, Based on the influence gradient, the adjustment direction for each action parameter is determined, and the adjustment step size for each action parameter is calculated based on the deviation sequence and the comprehensive risk coefficient, including: For each safety indicator that exceeds the preset safety threshold, the sensitivity coefficient of the safety indicator to each action parameter is extracted from the influence gradient; The adjustment direction of the motion parameter is determined according to the sign of the sensitivity coefficient: if the sensitivity coefficient is positive, the adjustment direction is to decrease the motion parameter; if the sensitivity coefficient is negative, the adjustment direction is to increase the motion parameter. Calculate the ratio of the excess amount of each safety indicator to the corresponding preset safety threshold as the absolute deviation ratio, and use the absolute value of the sensitivity coefficient as the weight to calculate the weighted sum of the absolute deviation ratios of each safety indicator to obtain the weighted deviation ratio. Add one to the aforementioned comprehensive risk coefficient to obtain the risk scaling factor; The adjustment step size for each action parameter is obtained by multiplying the preset base step size by the weighted deviation ratio and then by the risk scaling factor.
9. The method for correcting agent output and storing causal chain evidence based on physical constraints according to claim 1, characterized in that, Organizing the original instruction, the corrected instruction, and the execution feedback into a causal chain in chronological order, and storing evidence of the causal chain, including: Create causal chain nodes in chronological order. Each node contains a timestamp, event type, event data hash value, and the hash value of the previous node, forming a chain structure. The event types include at least instruction reception event, shadow simulation start event, correction decision event, instruction issuance event, and execution feedback event. The detailed data generated by the shadow simulation of the physical constraint sandbox is stored in an off-chain distributed storage system, and the Merkle root hash of the detailed data is calculated. The Merkle root hash is then stored in the extended field of the corresponding node. The detailed data includes at least the original instruction, the correction instruction, and the execution feedback. The hash value of the root node of the causal chain is stored on the blockchain for evidence, and the complete causal chain data is backed up to a trusted audit node.
10. A system for correcting agent output and storing causal chain evidence based on physical constraints, characterized in that, The system is used to implement the physical constraint-based agent output correction and causal chain evidence storage method according to any one of claims 1-9, the system comprising: The action parameter acquisition module is used to receive the original instructions to be issued by the intelligent agent and parse the action parameters in the original instructions; The simulation evolution module is used to input the original instructions into the physical constraint sandbox for shadow simulation, predict the evolution process of the physical device after the original instructions are executed, and generate a state evolution time series. The risk coefficient calculation module is used to calculate the deviation between each safety indicator and the corresponding preset safety threshold according to the state evolution time series, obtain the deviation sequence, and calculate the comprehensive risk coefficient based on the deviation sequence. The influence gradient calculation module is used to perform safety sensitivity analysis on the action parameters in the original instruction and calculate the influence gradient of the change of each action parameter on the safety index. The iterative optimization module is used to obtain the corrected action parameters by iteratively optimizing the deviation sequence, the comprehensive risk coefficient and the influence gradient, with the objective of minimizing the action parameter offset and the constraint that the safety index does not exceed the preset safety threshold. The correction feedback module is used to generate correction instructions based on the corrected action parameters, issue the correction instructions for execution, and collect execution feedback. The feedback and evidence storage module is used to organize the original instruction, the correction instruction, and the execution feedback into a causal chain in chronological order, and to store the causal chain as evidence.