Method and system for multimodal perception trajectory chaining and responsibility definition of physical robots
Patent Information
- Application Number
- CN202610960447.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
然而,该方案面向智能网联汽车场景,采用全量行驶数据上链的方式,存储成本极高,且缺少多机器人间的感知交叉验证机制来防止单点故障或数据篡改,也未考虑机器人特有的关节力矩、决策意图等多模态感知数据在责任界定中的作用,难以直接应用于实体机器人场景
[0018]本发明公开的一种实体机器人的多模态感知轨迹上链与责任界定方法及系统,通过在机器人机载芯片部署轻量级多模态人工智能模型实现感知数据的实时监控和异常事件触发,通过关键帧智能裁剪仅提取包含时间戳、环境特征向量及决策意图的语义关键帧,大幅降低了上链数据量;通过本地点对点网络实现多机器人间的感知交叉验证,利用感知分歧自动标记与优先上链机制防止单机故障或被攻击导致的误判;通过默克尔根哈希锚定区块链获取可信时间戳,确保证据的不可篡改性和司法公信力;通过智能合约根据预设判定规则自动执行责任判定与保险赔付,实现了从事故发生到保险理赔的全流程自动化。
Smart Images

Figure CN122807876A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of robotics, multimodal artificial intelligence and blockchain technology, and more specifically, to a method and system for on-chaining multimodal perception trajectory and liability determination of physical robots. Background Technology
[0002] With the widespread application of robots in industrial manufacturing, logistics and distribution, medical rehabilitation, and home services, the interaction between robots and humans and the environment is becoming increasingly frequent, leading to a rise in safety accidents and liability disputes. Accurately reconstructing the accident process and determining liability when robots collide, malfunction, or cause other safety incidents has become a core issue that urgently needs to be addressed in the industrialization of robotics.
[0003] Existing robot behavior recording solutions mostly rely on local log storage on the robot, making the data susceptible to tampering or deletion, and difficult to use as judicial evidence. Some solutions attempt to store all sensor data on the blockchain to leverage its immutability; however, the sheer volume of sensor data generated during robot operation presents a dual bottleneck of excessively high storage costs and insufficient blockchain throughput. Furthermore, existing liability determination schemes largely depend on post-event manual analysis, lacking automated judgment mechanisms, resulting in low efficiency and high subjectivity.
[0004] For example, Chinese patent CN115686778A discloses a blockchain-based decentralized swarm robot system framework. This solution mainly employs loss detection and reputation mechanisms to achieve decentralized collaboration and secure communication among swarm robots. However, this solution does not address the intelligent cropping and filtering of keyframes in robot perception data, nor does it solve the storage cost problem of uploading all data to the blockchain. Furthermore, it lacks automatic determination of accident liability and insurance compensation mechanisms, failing to meet the needs of evidence preservation and rapid liability determination in robot accident scenarios.
[0005] For example, Chinese patent CN118108785A discloses a system and method for storing and determining liability in intelligent connected vehicle accidents. This solution uses a blockchain network module and a smart contract module to store and determine liability for vehicle accident data. However, this solution is designed for intelligent connected vehicle scenarios and uses the method of putting all driving data on the blockchain, resulting in extremely high storage costs. It also lacks a cross-verification mechanism between multiple robots to prevent single-point failures or data tampering, and it does not consider the role of multimodal perception data such as joint torques and decision-making intentions unique to robots in liability determination, making it difficult to directly apply to physical robot scenarios.
[0006] Therefore, there is an urgent need for a method and system for storing key perception data on the blockchain, ensuring the credibility of evidence through cross-validation by multiple robots, and automatically determining accident liability and insurance compensation, which enables the on-chain storage of multimodal perception trajectories of physical robots. Summary of the Invention
[0007] In view of the above problems, the purpose of this invention is to provide a method and system for on-chaining and liability determination of multimodal perception trajectory of physical robots. It can reduce the cost of on-chain storage by intelligent keyframe clipping, ensure the credibility of evidence by cross-verification of multi-robot perception, and realize automatic liability determination and insurance compensation through smart contracts.
[0008] The first aspect of this invention provides a method for multimodal perception trajectory uplinking and responsibility delineation for physical robots, comprising: A lightweight multimodal artificial intelligence model is deployed on the robot's onboard chip to monitor the multimodal data collected by sensors in real time. The multimodal data includes visual images, sound signals, joint torques, and distance sensing data. When the lightweight multimodal AI model detects an abnormal event, it triggers intelligent keyframe cropping to extract semantic keyframes containing timestamps, environmental feature vectors, and decision intent. Cross-validation of perception among multiple robots is achieved through a local peer-to-peer network. Each validating robot independently evaluates the semantic keyframes and returns the validation results. The semantic keyframe data is constructed into a Merkle tree structure, and the Merkle root hash is anchored to the blockchain to obtain a trusted timestamp; Smart contracts deployed on the blockchain automatically determine liability and pay insurance claims based on preset judgment rules, combined with joint torque, safety thresholds, and collision signals.
[0009] In this solution, when the lightweight multimodal AI model detects an abnormal event, it triggers intelligent keyframe cropping, which specifically includes: Set multi-level anomaly detection thresholds, including a first-level early warning threshold and a second-level emergency threshold; When the sensor data exceeds the first-level warning threshold, the pre-recording buffer is activated, and multimodal data segments within the preset time window before and after are extracted as primary keyframes. When sensor data exceeds the secondary emergency threshold, it is marked as a high-priority event, which improves the extraction frequency and accuracy of key frames, and records the peak joint torque and collision signal intensity at the moment the event is triggered.
[0010] In this solution, the implementation of cross-validation of perception among multiple robots via a local point-to-point network specifically includes: Once keyframe intelligent cropping is triggered, the semantic keyframe is broadcast to neighboring robots within a preset range via a local peer-to-peer network. Each neighboring robot performs independent anomaly evaluation on the semantic keyframes based on its own deployed lightweight multimodal artificial intelligence model, and generates verification conclusions. The verification conclusions of all neighboring robots are summarized, and the verification consistency ratio is calculated. If the consistency ratio is lower than the preset consensus threshold, it is determined to be a perceived divergence event, which is automatically marked and prioritized for on-chain processing.
[0011] In this scheme, if the consensus ratio is lower than a preset consensus threshold, it is determined to be a perceived divergence event, automatically marked, and prioritized for on-chain processing. Specifically, this includes: When a perception divergence event is determined, a divergence evidence package is generated, which includes the original semantic keyframe, the verification conclusions of each neighboring robot, and the quantification value of the degree of divergence. The divergence evidence package is marked as the highest priority on-chain data, skipping the regular on-chain queuing sequence, and directly performing Merkle tree construction and hash anchoring operations; The discrepancy evidence package records the sensor channel identifier that caused the discrepancy and the corresponding environmental feature vector differences.
[0012] In this solution, constructing the semantic keyframe data into a Merkle tree structure and anchoring the Merkle root hash to the blockchain specifically includes: Based on the data importance of semantic keyframes and storage resource constraints, keyframe data is divided into on-chain anchored data and off-chain stored data. Calculate hash digests for off-chain stored data, construct a Merkle tree using all hash digests as leaf nodes, and calculate the Merkle root hash. Only the Merkle root hash and metadata are anchored to the blockchain, while off-chain data is stored in a distributed storage network; Establish an index mapping relationship between on-chain hashes and off-chain data to ensure that the integrity of off-chain data can be located and verified through on-chain credentials during subsequent evidence collection.
[0013] In this solution, the smart contract deployed on the blockchain automatically performs liability determination and insurance compensation based on preset judgment rules, combined with joint torque, safety threshold, and collision signals. Specifically, this includes: The smart contract has a pre-built multi-dimensional liability determination rule base, which includes joint torque threshold range, safety distance threshold and collision signal level determination conditions; The Merkle root hash and metadata corresponding to the semantic keyframes anchored on the receiving chain are used to extract joint torque data, collision signal data and timestamp data. The extracted data is matched according to the multidimensional liability determination rule base to determine the liability level, which includes full liability, partial liability, and force majeure. Based on the determined level of liability, a claim request is automatically generated and sent to the associated insurance platform interface, triggering the insurance payout process.
[0014] A second aspect of the present invention provides a multimodal perception trajectory uplinking and responsibility delineation system for a physical robot, comprising a memory and a processor. The memory includes a method program for multimodal perception trajectory uplinking and responsibility delineation for a physical robot. When the processor executes the method program for multimodal perception trajectory uplinking and responsibility delineation for a physical robot, it performs the following steps: A lightweight multimodal artificial intelligence model is deployed on the robot's onboard chip to monitor the multimodal data collected by sensors in real time. The multimodal data includes visual images, sound signals, joint torques, and distance sensing data. When the lightweight multimodal AI model detects an abnormal event, it triggers intelligent keyframe cropping to extract semantic keyframes containing timestamps, environmental feature vectors, and decision intent. Cross-validation of perception among multiple robots is achieved through a local peer-to-peer network. Each validating robot independently evaluates the semantic keyframes and returns the validation results. The semantic keyframe data is constructed into a Merkle tree structure, and the Merkle root hash is anchored to the blockchain to obtain a trusted timestamp; Smart contracts deployed on the blockchain automatically determine liability and pay insurance claims based on preset judgment rules, combined with joint torque, safety thresholds, and collision signals.
[0015] In this solution, when the lightweight multimodal AI model detects an abnormal event, it triggers intelligent keyframe cropping, which specifically includes: Set multi-level anomaly detection thresholds, including a first-level early warning threshold and a second-level emergency threshold; When the sensor data exceeds the first-level warning threshold, the pre-recording buffer is activated, and multimodal data segments within the preset time window before and after are extracted as primary keyframes. When sensor data exceeds the secondary emergency threshold, it is marked as a high-priority event, which improves the extraction frequency and accuracy of key frames, and records the peak joint torque and collision signal intensity at the moment the event is triggered.
[0016] In this solution, the implementation of cross-validation of perception among multiple robots via a local point-to-point network specifically includes: Once keyframe intelligent cropping is triggered, the semantic keyframe is broadcast to neighboring robots within a preset range via a local peer-to-peer network. Each neighboring robot performs independent anomaly evaluation on the semantic keyframes based on its own deployed lightweight multimodal artificial intelligence model, and generates verification conclusions. The verification conclusions of all neighboring robots are summarized, and the verification consistency ratio is calculated. If the consistency ratio is lower than the preset consensus threshold, it is determined to be a perceived divergence event, which is automatically marked and prioritized for on-chain processing.
[0017] A third aspect of the present invention provides a computer-readable storage medium including a method program for multimodal sensing trajectory uplinking and responsibility delineation of an entity robot. When the method program is executed by a processor, it implements the steps of the method program for multimodal sensing trajectory uplinking and responsibility delineation of an entity robot as described in any of the preceding claims.
[0018] This invention discloses a method and system for on-chaining and liability determination of multimodal perception trajectory of physical robots. It achieves real-time monitoring of perception data and triggering of abnormal events by deploying a lightweight multimodal artificial intelligence model on the robot's onboard chip. Through intelligent keyframe cropping, only semantic keyframes containing timestamps, environmental feature vectors, and decision intentions are extracted, significantly reducing the amount of data on-chain. It enables cross-verification of perception among multiple robots via a local peer-to-peer network, and utilizes automatic marking and priority on-chaining mechanisms for perception discrepancies to prevent misjudgments caused by single-machine failures or attacks. It obtains trusted timestamps by anchoring the blockchain using Merkle root hashes, ensuring the immutability of evidence and judicial credibility. Through smart contracts, it automatically executes liability determination and insurance payouts according to preset judgment rules, achieving full automation from accident occurrence to insurance claim settlement. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope.
[0020] Figure 1 A flowchart of a method for multimodal perception trajectory uplinking and responsibility delineation for a physical robot according to the present invention is shown; Figure 2 The flowchart of a keyframe intelligent cropping method provided by an embodiment of the present invention is shown; Figure 3 A flowchart of a multi-robot perception cross-validation method provided by an embodiment of the present invention is shown; Figure 4 A block diagram of a multimodal perception trajectory chaining and responsibility delineation system for a physical robot according to the present invention is shown. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Unless otherwise defined, all terms (including technical and scientific terms) used in embodiments of this invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in a common dictionary shall be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and not as being interpreted in an idealized or highly formalized sense, unless expressly defined in this embodiment of the invention.
[0023] The terms "first," "second," and similar words used in the embodiments of this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "an," "a," or "the" do not indicate a quantity limitation, but rather indicate the presence of at least one. Similarly, terms such as "including" or "comprising" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The steps preceding or following the steps in the method of the embodiments of this invention are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0024] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0025] Figure 1 The flowchart of a multimodal perception trajectory uplink and responsibility delineation method for a physical robot according to the present invention is shown.
[0026] like Figure 1 As shown, the first aspect of this invention discloses a method for multimodal perception trajectory uplinking and responsibility delineation for physical robots, the method comprising: S102, Deploy a lightweight multimodal artificial intelligence model on the robot's onboard chip to monitor the multimodal data collected by the sensors in real time. The multimodal data includes visual images, sound signals, joint torques, and distance sensing data. S104, when the lightweight multimodal artificial intelligence model detects an abnormal event, it triggers keyframe intelligent cropping to extract semantic keyframes containing timestamps, environmental feature vectors and decision intentions; S106, realize the perception cross-verification among multiple robots through the local point-to-point network, and each verification robot independently evaluates the semantic keyframe and returns the verification result; S108, construct the semantic keyframe data into a Merkle tree structure, and anchor the Merkle root hash to the blockchain to obtain a trusted timestamp; S110, a smart contract deployed on the blockchain, automatically executes liability determination and insurance compensation based on preset judgment rules, combined with joint torque, safety threshold and collision signal.
[0027] It should be noted that in this embodiment, the lightweight multimodal model refers to an inference model that, after pruning and quantization compression, can run on the robot's onboard edge chip, integrating visual, auditory, and tactile perception channels. Input data includes camera footage, sound picked up by the microphone array, torque sensor readings, and distance measurements from LiDAR or ultrasound. Keyframe pruning means that when an anomaly is detected, the full data stream is not stored; only the most representative segment is extracted. Each keyframe is accompanied by a precise timestamp, environmental characteristic parameters at that time, and a description of the robot's decision intention at that moment. The peer-to-peer network is a decentralized communication channel where multiple robots in the same area are directly connected via wireless LAN or short-range communication protocols, exchanging data without relying on a central server. The Merkle tree structure obtains a unique root hash value through hierarchical hashing of the underlying data. Any modification to the underlying content will change the root hash, thus enabling the detection of tampering. Throughout the entire process, accident-related evidence can be secured without storing the full data, significantly reducing the storage overhead on the blockchain while ensuring the integrity of the evidence.
[0028] Figure 2 A flowchart of a keyframe intelligent cropping method provided by an embodiment of the present invention is shown.
[0029] like Figure 2 As shown in the embodiment of the present invention, when the lightweight multimodal artificial intelligence model detects an abnormal event, triggering intelligent keyframe cropping specifically includes: Set multi-level anomaly detection thresholds, including a first-level early warning threshold and a second-level emergency threshold; When the sensor data exceeds the first-level warning threshold, the pre-recording buffer is activated, and multimodal data segments within the preset time window before and after are extracted as primary keyframes. When sensor data exceeds the secondary emergency threshold, it is marked as a high-priority event, which improves the extraction frequency and accuracy of key frames, and records the peak joint torque and collision signal intensity at the moment the event is triggered.
[0030] It should be noted that in this embodiment, the division of the two threshold levels directly determines the keyframe extraction behavior pattern. Level 1 warning corresponds to situations where there is a slight deviation but not yet an emergency, typically such as joint torque exceeding the normal range but still within the safety limit, or an obstacle appearing in front of the user but still far away. When the level 1 threshold is reached, the system first activates the pre-recording buffer, capturing a multimodal data segment approximately two seconds before and after the trigger, marking it as a primary keyframe, and temporarily storing it in the local cache for later use. Level 2 emergency threshold represents situations where a real accident has occurred; a sudden surge in joint torque exceeding the safety limit or a collision sensor triggering the system directly falls into this category. At level 2, the system increases the priority, increasing the keyframe extraction frequency from one frame per second to multiple frames per second, and correspondingly improving the sampling accuracy. Simultaneously, it records the peak torque and collision signal strength as important references for future liability determination. This gradient design between the two levels ensures that storage space is not wasted during normal periods, while providing sufficiently detailed evidence when a real accident occurs.
[0031] Figure 3 A flowchart of a multi-robot perception cross-validation method provided by an embodiment of the present invention is shown.
[0032] like Figure 3 As shown in the embodiment of the present invention, the step of realizing cross-verification of perception among multiple robots through a local point-to-point network specifically includes: Once keyframe intelligent cropping is triggered, the semantic keyframe is broadcast to neighboring robots within a preset range via a local peer-to-peer network. Each neighboring robot performs independent anomaly evaluation on the semantic keyframes based on its own deployed lightweight multimodal artificial intelligence model, and generates verification conclusions. The verification conclusions of all neighboring robots are summarized, and the verification consistency ratio is calculated. If the consistency ratio is lower than the preset consensus threshold, it is determined to be a perceived divergence event, which is automatically marked and prioritized for on-chain processing.
[0033] It's important to note that in this embodiment, the purpose of cross-validation is to prevent perceptual biases or malfunctions of a single robot from affecting the credibility of the evidence. After keyframe pruning is triggered, the robot sends the extracted semantic keyframes to its neighbors within a certain range via a peer-to-peer network. Neighbors don't need to run the exact same task, but they must install lightweight models of the same specifications to ensure comparable conclusions. Each neighbor independently evaluates the results using its own model, determining whether the result is abnormal or normal. The system then calculates the percentage of neighbors returning abnormal conclusions out of the total number of participants; this percentage is called the validation consistency ratio. If the ratio is lower than a preset consensus threshold (usually set at around 70%), it indicates a significant difference in opinion on the matter, and the system marks this event as a perceptual disagreement and assigns it the highest priority for on-chain processing. Disagreements may arise from several reasons, such as sensor malfunctions on the requesting party, model modifications, or extremely harsh environmental conditions making it difficult to determine the cause. With this cross-validation mechanism, it becomes much more difficult for a single machine to malfunction or for malicious attacks to insert false evidence.
[0034] According to an embodiment of the present invention, if the consensus ratio is lower than a preset consensus threshold, it is determined to be a perceived divergence event, automatically marked, and prioritized for on-chain processing, specifically including: When a perception divergence event is determined, a divergence evidence package is generated, which includes the original semantic keyframe, the verification conclusions of each neighboring robot, and the quantification value of the degree of divergence. The divergence evidence package is marked as the highest priority on-chain data, skipping the regular on-chain queuing sequence, and directly performing Merkle tree construction and hash anchoring operations; The discrepancy evidence package records the sensor channel identifier that caused the discrepancy and the corresponding environmental feature vector differences.
[0035] It's important to note that in this embodiment, while traditional approaches often treat perceived discrepancies as errors or noise, in the context of robot liability delineation, the discrepancy itself is valuable information. It can reflect whether a sensor is faulty, whether the model has been tampered with, or whether the environment is so harsh that no one can make a reliable judgment. In addition to the original semantic keyframes, the discrepancy evidence package includes the verification conclusions of all neighboring robots, regardless of consistency, along with the precise consistency ratio and the difference in confidence scores between each robot. This package is marked as the highest priority, skipping the normal queue and directly uploaded to the blockchain, receiving immediate timestamp protection. The package also records which sensor channel caused the discrepancy and where the differences in environmental characteristics lie, facilitating rapid root cause identification during post-analysis. By transforming discrepancies from defects into usable evidence, the completeness and fairness of the entire liability delineation process are significantly improved.
[0036] According to an embodiment of the present invention, the step of constructing the semantic keyframe data into a Merkle tree structure and anchoring the Merkle root hash to the blockchain specifically includes: Based on the data importance of semantic keyframes and storage resource constraints, keyframe data is divided into on-chain anchored data and off-chain stored data. Calculate hash digests for off-chain stored data, construct a Merkle tree using all hash digests as leaf nodes, and calculate the Merkle root hash. Only the Merkle root hash and metadata are anchored to the blockchain, while off-chain data is stored in a distributed storage network; Establish an index mapping relationship between on-chain hashes and off-chain data to ensure that the integrity of off-chain data can be located and verified through on-chain credentials during subsequent evidence collection.
[0037] It's important to note that in this embodiment, the combination of Merkle tree and hash anchoring addresses the issue of the blockchain being unable to store the large amount of keyframe data generated during robot operation. Writing every frame directly to the chain would be prohibitively expensive in terms of storage and throughput. The actual solution involves hierarchically processing keyframes according to their importance. Only the Merkle root hash and essential metadata such as timestamps, robot IDs, and event categories are uploaded to the chain, while the original frame data resides in an off-chain distributed storage network. The construction process is as follows: first, a hash digest is calculated for each keyframe, treating it as a leaf node. Adjacent leaf nodes are paired and their parent node hashes are calculated, merging upwards layer by layer until only a single unique root hash remains. During forensics, the root hash on the chain is used to locate the off-chain data, and the hash is recalculated and compared to determine if the data has been tampered with. This process reduces on-chain overhead by several orders of magnitude, essentially eliminating the storage bottleneck.
[0038] According to an embodiment of the present invention, the smart contract deployed on the blockchain automatically performs liability determination and insurance compensation based on preset judgment rules, combined with joint torque, safety threshold, and collision signal, specifically including: The smart contract has a pre-built multi-dimensional liability determination rule base, which includes joint torque threshold range, safety distance threshold and collision signal level determination conditions; The Merkle root hash and metadata corresponding to the semantic keyframes anchored on the receiving chain are used to extract joint torque data, collision signal data and timestamp data. The extracted data is matched according to the multidimensional liability determination rule base to determine the liability level, which includes full liability, partial liability, and force majeure. Based on the determined level of liability, a claim request is automatically generated and sent to the associated insurance platform interface, triggering the insurance payout process.
[0039] It should be noted that in this embodiment, the smart contract is an on-chain program that executes automatically once preset conditions are met, requiring no human intervention throughout the process. A multi-dimensional liability determination rule base is pre-embedded in the contract. The rules cover determination conditions such as joint torque range division, safe distance limits, and collision signal levels, and are pre-defined by security experts based on industry standards and past accident cases. When the contract executes, it pulls torque data, collision signals, and timestamps from off-chain and matches them against the rule base. For example, if the torque data exceeds the safe operating range and the collision signal indicates an active impact, full liability is determined; if the torque is normal but the collision signal indicates a passive impact and there are unforeseen obstacles in the environment, partial liability or force majeure is determined. Once the liability level is determined, the contract automatically generates a claim application specifying the accident overview, liability level, and compensation amount, and sends it to the connected insurance platform via the on-chain interface to initiate the compensation process. From on-chain processing to liability determination to compensation, the entire process is automated.
[0040] According to an embodiment of the present invention, the method further includes a local autonomous evidence storage mode under communication interruption scenarios, specifically including: When the local peer-to-peer network is unavailable, the robot switches to local autonomous evidence storage mode and caches semantic keyframe data in the local secure storage area. Apply anti-tampering tags to cached data, the anti-tampering tags including an incrementing sequence number and a hash chain check value, to ensure the temporal integrity and content integrity of cached data; Once the network is restored, the cached data is uploaded to the blockchain in batches and integrity verification is performed. After successful verification, the cached data is cleared from the local secure storage area.
[0041] It should be noted that in this embodiment, local autonomous evidence storage is a fallback solution specifically designed for situations where the network is completely disconnected. When the robot cannot connect to its neighbors due to network failure, signal blockage, or communication module damage, the system automatically switches to local mode to continue operating. In local mode, sensor monitoring and keyframe extraction continue as usual. The difference is that all keyframes are temporarily cached in a secure storage area with hardware-level isolation, preventing external programs from modifying its contents. Each cached data entry carries an incrementing sequence number to ensure that the order cannot be forged. A hash chain is also maintained, with the hash of each data entry concatenated with the previous one. Modification of any intermediate chain will cause all subsequent checksums to mismatch. Once the network is restored, the cached data is uploaded to the chain in batches, and the local hash chain and the newly constructed Merkle tree are checked to ensure that the data has not been tampered with or lost during the network outage. This fallback solution ensures that the chain of evidence remains intact even under the worst communication conditions.
[0042] According to an embodiment of the present invention, the method further includes the robot actively requesting cross-validation, specifically including: When the robot's confidence in its own perception results is lower than a preset self-test threshold, it actively sends a verification request to neighboring robots. After receiving the verification request, the neighbor robot reviews the semantic keyframes of the requester and returns the review conclusion. If the review conclusion is consistent with the requester's own judgment, the credibility weight of the semantic keyframe is increased; if they are inconsistent, the perception disagreement processing procedure is triggered.
[0043] It should be noted that in this embodiment, the conventional process involves the robot detecting an anomaly before initiating a verification request. However, sometimes the robot's judgment is not so certain. For example, a sudden change in lighting might cause a drop in visual confidence, or a sensor reading might fluctuate around a critical value. In such cases, it's better to confirm with neighbors beforehand than to wait for the anomaly to be confirmed. The self-check trigger threshold for proactive verification is dynamically set based on the overall confidence of each sensor modality; if it falls below the threshold, a request is automatically issued. If neighbors agree on the conclusions after responding, the evidentiary value of this keyframe is strengthened, carrying more weight in subsequent on-chain data collection and liability determination. If neighbors disagree, a disagreement resolution process is followed to prevent low-confidence data from being included in liability determination. This proactive mechanism essentially introduces corrective capabilities in advance during uncertain phases, significantly reducing the probability of misjudgment.
[0044] According to an embodiment of the present invention, the method further includes dynamically adjusting the verification strategy based on the event type, specifically including: Identify the event type of the abnormal event, including collision events, environmental hazard events, and operational abnormal events; For collision events, priority is given to validating joint torque data and collision signal data; for environmental hazard events, priority is given to validating multi-sensor fusion data; for operational anomaly events, priority is given to validating decision intent data. Configure different keyframe extraction parameters and the number of verification robots according to the event type.
[0045] It should be noted that in this embodiment, collision events, environmental hazard events, and operational anomaly events each have different focuses in terms of data requirements and verification priorities. When a collision occurs, the torque usually increases sharply, and the collision sensor will also react. Therefore, the verification strategy prioritizes ensuring the accuracy and depth of torque and collision signal acquisition and analysis. Environmental hazard events such as fires and leaks involve information from multiple sensors, including smoke, temperature, and gas concentration. The verification focus is on the consistency of multimodal data. Operational anomalies, such as a robot deviating from its planned route, require examining the decision intent data to determine whether the deviation was due to autonomous choice or external interference. Depending on the event type, parameters such as the sampling frequency of keyframes, data bit depth, and time window length will be adjusted accordingly, and the number of robots participating in the verification can also be flexibly increased or decreased. The advantage of this approach is that, under limited computing and communication conditions, verification resources can be concentrated on where they are most needed.
[0046] According to an embodiment of the present invention, the step of matching the extracted data according to the multidimensional responsibility determination rule base to determine the responsibility level specifically includes: A weighted comprehensive score is calculated based on joint torque data, collision signal data, and environmental feature vectors. When the overall score falls within the full liability range, the robot is deemed to bear full responsibility; when the overall score falls within the partial liability range, the liability sharing coefficient is determined based on the score ratio; when the overall score falls within the force majeure range, it is deemed to be exempt from liability due to force majeure. Different levels of liability correspond to different insurance payout ratios. Full liability corresponds to full compensation, partial liability is compensated according to the liability sharing coefficient ratio, and force majeure activates a special claims channel.
[0047] It should be noted that in this embodiment, the determination of multi-dimensional liability adopts a weighted comprehensive scoring method to achieve fine-grained classification. The scoring incorporates data from three dimensions: joint torque reflects the degree of active force applied by the robot, collision signal reflects the severity and directional attributes of the impact, and environmental characteristics reflect external factors such as the distance of obstacles and ground conditions. Each dimension has its own weight coefficient, which is pre-calibrated based on historical accident databases and expert experience. Full liability generally occurs when the robot actively performs an operation that causes the accident in a normal environment, with torque exceeding the range and no external collision signal to support it; partial liability usually involves a causal relationship between internal and external factors; force majeure refers to encountering unforeseen and uncontrollable external factors such as sudden natural disasters. Different liability levels correspond to different compensation strategies and proportions, avoiding the previous one-size-fits-all approach, and ensuring that the compensation results are better matched with the actual degree of fault.
[0048] Please see Figure 4The present invention also provides a multimodal perception trajectory uplinking and responsibility delineation system 4 for physical robots. The system includes a memory 401 and a processor 402. The memory includes a method program for multimodal perception trajectory uplinking and responsibility delineation of physical robots. When the processor executes the method program for multimodal perception trajectory uplinking and responsibility delineation of physical robots, it performs the following steps: A lightweight multimodal artificial intelligence model is deployed on the robot's onboard chip to monitor the multimodal data collected by sensors in real time. The multimodal data includes visual images, sound signals, joint torques, and distance sensing data. When the lightweight multimodal AI model detects an abnormal event, it triggers intelligent keyframe cropping to extract semantic keyframes containing timestamps, environmental feature vectors, and decision intent. Cross-validation of perception among multiple robots is achieved through a local peer-to-peer network. Each validating robot independently evaluates the semantic keyframes and returns the validation results. The semantic keyframe data is constructed into a Merkle tree structure, and the Merkle root hash is anchored to the blockchain to obtain a trusted timestamp; Smart contracts deployed on the blockchain automatically determine liability and pay insurance claims based on preset judgment rules, combined with joint torque, safety thresholds, and collision signals.
[0049] It should be noted that in this embodiment, the system consists of a memory and a processor. The memory stores the complete program code for the multimodal perception trajectory on-chaining and liability determination method of the physical robot. The processor loads and executes the program in the following general process: First, the aforementioned lightweight multimodal model is deployed on the onboard chip side, allowing it to continuously monitor the data streams sent by the visual, sound, torque, and distance sensors. Once the model detects an abnormal signal, it immediately initiates keyframe pruning to extract semantic keyframes from the continuous stream. Next, the keyframes are sent to neighboring devices via a peer-to-peer network for cross-validation. After the validation results are collected, they are summarized and analyzed, and divergent events are marked and prioritized. Subsequently, the keyframe data is organized into a Merkle tree structure, and only the root hash is on-chained to obtain a timestamp certificate. Finally, the smart contract deployed on the blockchain completes the liability determination according to the built-in rules and initiates the compensation through the insurance interface. The above steps are linked together to form a complete automated closed loop from perception and monitoring to evidence solidification and liability determination.
[0050] According to an embodiment of the present invention, the step of triggering intelligent keyframe cropping when the lightweight multimodal artificial intelligence model detects an abnormal event specifically includes: Set multi-level anomaly detection thresholds, including a first-level early warning threshold and a second-level emergency threshold; When the sensor data exceeds the first-level warning threshold, the pre-recording buffer is activated, and multimodal data segments within the preset time window before and after are extracted as primary keyframes. When sensor data exceeds the secondary emergency threshold, it is marked as a high-priority event, which improves the extraction frequency and accuracy of key frames, and records the peak joint torque and collision signal intensity at the moment the event is triggered.
[0051] It should be noted that in this embodiment, the keyframe extraction management of the system relies on a two-level anomaly detection threshold. When the first-level warning threshold is triggered, it enters the primary extraction mode, and the buffer begins recording data within the preceding and following windows as preliminary evidence. When the second-level emergency threshold is triggered, it switches to a high-priority extraction mode, which not only increases the sampling frequency and accuracy but also additionally records two key quantitative indicators—peak torque and collision intensity—for subsequent accountability. The two mechanisms work together to ensure that commensurate evidence recording accuracy is obtained regardless of the severity of the anomaly. Of course, the first-level warning threshold and the second-level emergency threshold in this application can be set by those skilled in the art according to actual needs, and this application will not elaborate further.
[0052] According to an embodiment of the present invention, the step of realizing cross-verification of perception among multiple robots through a local point-to-point network specifically includes: Once keyframe intelligent cropping is triggered, the semantic keyframe is broadcast to neighboring robots within a preset range via a local peer-to-peer network. Each neighboring robot performs independent anomaly evaluation on the semantic keyframes based on its own deployed lightweight multimodal artificial intelligence model, and generates verification conclusions. The verification conclusions of all neighboring robots are summarized, and the verification consistency ratio is calculated. If the consistency ratio is lower than the preset consensus threshold, it is determined to be a perceived divergence event, which is automatically marked and prioritized for on-chain processing.
[0053] It should be noted that in this embodiment, the system utilizes a peer-to-peer network to achieve cross-validation of perception among multiple robots. After each neighboring robot evaluates using its independent model, it returns its own conclusion. The system collects all conclusions and calculates the consistency ratio. Events whose ratio does not reach the consensus threshold are automatically labeled with a perception divergence tag and scheduled for on-chain with the highest priority. The purpose of this step is to intercept false evidence generated by single-machine failures or attacks as much as possible outside the responsibility determination process.
[0054] The present invention also provides a computer-readable storage medium storing a computer program thereon, the computer-readable storage medium including a method program for multimodal perception trajectory uplinking and responsibility delineation of an entity robot, wherein when the multimodal perception trajectory uplinking and responsibility delineation method program of the entity robot is executed by a processor, the steps of the multimodal perception trajectory uplinking and responsibility delineation method of the entity robot as described above are implemented.
[0055] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for multimodal perception trajectory uplinking and responsibility delineation for a physical robot, characterized in that, The method includes: S102, Deploy a lightweight multimodal artificial intelligence model on the robot's onboard chip to monitor the multimodal data collected by the sensors in real time. The multimodal data includes visual images, sound signals, joint torques, and distance sensing data. S104, when the lightweight multimodal artificial intelligence model detects an abnormal event, it triggers keyframe intelligent cropping to extract semantic keyframes containing timestamps, environmental feature vectors and decision intentions; S106, realize the perception cross-verification among multiple robots through the local point-to-point network, and each verification robot independently evaluates the semantic keyframe and returns the verification result; S108, construct the semantic keyframe data into a Merkle tree structure, and anchor the Merkle root hash to the blockchain to obtain a trusted timestamp; S110, a smart contract deployed on the blockchain, automatically executes liability determination and insurance compensation based on preset judgment rules, combined with joint torque, safety threshold and collision signal.
2. The method for multimodal perception trajectory uplinking and responsibility delineation of a physical robot according to claim 1, characterized in that, When the lightweight multimodal AI model detects an abnormal event, it triggers intelligent keyframe cropping, specifically including: Set multi-level anomaly detection thresholds, including a first-level early warning threshold and a second-level emergency threshold; When the sensor data exceeds the first-level warning threshold, the pre-recording buffer is activated, and multimodal data segments within the preset time window before and after are extracted as primary keyframes. When sensor data exceeds the secondary emergency threshold, it is marked as a high-priority event, which improves the extraction frequency and accuracy of key frames, and records the peak joint torque and collision signal intensity at the moment the event is triggered.
3. The method for multimodal perception trajectory uplinking and responsibility delineation of a physical robot according to claim 1, characterized in that, The method of achieving cross-verification of perception among multiple robots through a local point-to-point network specifically includes: Once keyframe intelligent cropping is triggered, the semantic keyframe is broadcast to neighboring robots within a preset range via a local peer-to-peer network. Each neighboring robot performs independent anomaly evaluation on the semantic keyframes based on its own deployed lightweight multimodal artificial intelligence model, and generates verification conclusions. The verification conclusions of all neighboring robots are summarized, and the verification consistency ratio is calculated. If the consistency ratio is lower than the preset consensus threshold, it is determined to be a perceived divergence event, which is automatically marked and prioritized for on-chain processing.
4. The method for multimodal perception trajectory uplinking and responsibility delineation of a physical robot according to claim 3, characterized in that, If the consensus ratio is lower than a preset consensus threshold, it is determined to be a perceived divergence event, automatically marked, and prioritized for on-chain processing. Specifically, this includes: When a perception divergence event is determined, a divergence evidence package is generated, which includes the original semantic keyframe, the verification conclusions of each neighboring robot, and the quantification value of the degree of divergence. The divergence evidence package is marked as the highest priority on-chain data, skipping the regular on-chain queuing sequence, and directly performing Merkle tree construction and hash anchoring operations; The discrepancy evidence package records the sensor channel identifier that caused the discrepancy and the corresponding environmental feature vector differences.
5. The method for multimodal perception trajectory uplinking and responsibility delineation of a physical robot according to claim 1, characterized in that, The step of constructing the semantic keyframe data into a Merkle tree structure and anchoring the Merkle root hash to the blockchain specifically includes: Based on the data importance of semantic keyframes and storage resource constraints, keyframe data is divided into on-chain anchored data and off-chain stored data. Calculate hash digests for off-chain stored data, construct a Merkle tree using all hash digests as leaf nodes, and calculate the Merkle root hash. Only the Merkle root hash and metadata are anchored to the blockchain, while off-chain data is stored in a distributed storage network; Establish an index mapping relationship between on-chain hashes and off-chain data to ensure that the integrity of off-chain data can be located and verified through on-chain credentials during subsequent evidence collection.
6. The method for multimodal perception trajectory uplinking and responsibility delineation of a physical robot according to claim 1, characterized in that, The smart contract deployed on the blockchain automatically performs liability determination and insurance compensation based on preset judgment rules, combined with joint torque, safety threshold, and collision signals, specifically including: The smart contract has a pre-built multi-dimensional liability determination rule base, which includes joint torque threshold range, safety distance threshold and collision signal level determination conditions; The Merkle root hash and metadata corresponding to the semantic keyframes anchored on the receiving chain are used to extract joint torque data, collision signal data and timestamp data. The extracted data is matched according to the multidimensional liability determination rule base to determine the liability level, which includes full liability, partial liability, and force majeure. Based on the determined level of liability, a claim request is automatically generated and sent to the associated insurance platform interface, triggering the insurance payout process.
7. A multimodal perception trajectory uplink and responsibility delineation system for a physical robot, characterized in that, The system includes a memory and a processor. The memory includes a program for multimodal perception trajectory uploading and responsibility delimitation of a physical robot. When the processor executes the program for multimodal perception trajectory uploading and responsibility delimitation of a physical robot, it performs the following steps: A lightweight multimodal artificial intelligence model is deployed on the robot's onboard chip to monitor the multimodal data collected by sensors in real time. The multimodal data includes visual images, sound signals, joint torques, and distance sensing data. When the lightweight multimodal AI model detects an abnormal event, it triggers intelligent keyframe cropping to extract semantic keyframes containing timestamps, environmental feature vectors, and decision intent. Cross-validation of perception among multiple robots is achieved through a local peer-to-peer network. Each validating robot independently evaluates the semantic keyframes and returns the validation results. The semantic keyframe data is constructed into a Merkle tree structure, and the Merkle root hash is anchored to the blockchain to obtain a trusted timestamp; Smart contracts deployed on the blockchain automatically determine liability and pay insurance claims based on preset judgment rules, combined with joint torque, safety thresholds, and collision signals.
8. A multimodal perception trajectory uplink and responsibility delineation system for a physical robot according to claim 7, characterized in that, When the lightweight multimodal AI model detects an abnormal event, it triggers intelligent keyframe cropping, specifically including: Set multi-level anomaly detection thresholds, including a first-level early warning threshold and a second-level emergency threshold; When the sensor data exceeds the first-level warning threshold, the pre-recording buffer is activated, and multimodal data segments within the preset time window before and after are extracted as primary keyframes. When sensor data exceeds the secondary emergency threshold, it is marked as a high-priority event, which improves the extraction frequency and accuracy of key frames, and records the peak joint torque and collision signal intensity at the moment the event is triggered.
9. A multimodal perception trajectory uplink and responsibility delineation system for a physical robot according to claim 7, characterized in that, The method of achieving cross-verification of perception among multiple robots through a local point-to-point network specifically includes: Once keyframe intelligent cropping is triggered, the semantic keyframe is broadcast to neighboring robots within a preset range via a local peer-to-peer network. Each neighboring robot performs independent anomaly evaluation on the semantic keyframes based on its own deployed lightweight multimodal artificial intelligence model, and generates verification conclusions. The verification conclusions of all neighboring robots are summarized, and the verification consistency ratio is calculated. If the consistency ratio is lower than the preset consensus threshold, it is determined to be a perceived divergence event, which is automatically marked and prioritized for on-chain processing.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer-readable storage medium includes a method program for multimodal sensing trajectory uplinking and responsibility delineation of an entity robot. When the method program is executed by a processor, it implements the steps of a method for multimodal sensing trajectory uplinking and responsibility delineation of an entity robot as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Decentralization swarm robot system framework based on block chain
CN115686778A
5alpha, 8alpha-peroxy ergosterol-3-thiazolidinedione derivative as well as preparation method and application thereof
CN118108785A