A Method for Parsing Embodied Intelligence Protocols and Verifying Behaviors in Multi-Agent Collaboration

By employing an embodied intelligence protocol parsing and behavior verification method based on multi-agent collaboration, the problems of low protocol parsing efficiency and insufficient robot behavior analysis in multi-protocol environments are solved, achieving efficient and intelligent protocol recognition and action execution, and improving the system's adaptability and security.

CN122137895APending Publication Date: 2026-06-02SHENYANG INSTITUTE OF CHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENYANG INSTITUTE OF CHEMICAL TECHNOLOGY
Filing Date
2026-01-19
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing embodied intelligence systems with multi-protocol and multi-device collaboration, protocol parsing is inefficient and inaccurate, lacks intelligent collaboration, and robot behavior analysis lacks knowledge support, making it difficult to meet the reasoning needs of complex tasks.

Method used

This paper proposes a method for embodied intelligent protocol parsing and behavior verification in multi-agent collaboration. By defining intelligent agents for protocol identification, protocol scheduling, and behavior reasoning, and utilizing protocol fingerprints and dual verification rules, the method achieves automatic protocol identification, dynamic scheduling and parsing, and intelligent analysis of robot behavior.

Benefits of technology

It improves the system's intelligence and adaptability, enhances the accuracy and efficiency of protocol recognition and parsing, ensures the compliance and security of robot action execution, and strengthens the robustness and autonomous evolution capabilities of the embodied intelligent system in complex multi-protocol environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122137895A_ABST
    Figure CN122137895A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of industrial communication and robot control technology, and discloses a method for embodied intelligence protocol parsing and behavior verification in multi-agent collaboration. The method includes: capturing communication messages between the embodied intelligence system and external devices; extracting statistical and semantic features and fusing them to construct a protocol fingerprint; obtaining structured semantic data by having the agent identify the protocol type and scheduling an adaptation parsing tool; generating action execution schemes by combining an external knowledge base; verifying and deducing through dual-verification rules; and triggering regeneration and dynamic optimization in case of anomalies. This invention achieves accurate multi-protocol identification and parsing, and safe and compliant action execution through multi-agent collaboration, feature fusion modeling, dual security verification, and a closed-loop feedback mechanism. It enhances the system's robustness and autonomous evolution capabilities in complex environments and is applicable to various scenarios such as intelligent manufacturing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of industrial communication and robot control technology, and more specifically, to a method for parsing embodied intelligence protocols and verifying behaviors in multi-agent collaboration. Background Technology

[0002] With the development of artificial intelligence and automation technologies, embodied intelligent systems are widely used in various scenarios such as robotics, intelligent manufacturing, and warehousing and logistics. Embodied intelligent systems typically consist of multiple functional modules collaborating to complete tasks. For robots to achieve efficient and accurate action execution, they rely on the parsing and understanding of various types of data, including task protocols, communication protocols, and behavioral instructions. In complex systems, protocol types are diverse, and their formats, content, and semantics differ. Accurately identifying protocols and scheduling the corresponding parsing modules is a critical problem that urgently needs to be solved in embodied intelligent systems.

[0003] Existing protocol parsing systems typically rely on manually preset rules or parsing fixed protocol structures. While these methods work well when the number of protocols is limited and the structure is fixed, they lack flexibility and cannot effectively adapt to the complexity of multi-protocol environments when faced with ever-increasing device types, protocol formats, and changing application scenarios. This results in low parsing efficiency and insufficient accuracy. Especially in embodied intelligence systems with multi-protocol collaboration, achieving dynamic protocol recognition and parsing while ensuring the intelligence and real-time performance of robot behavior reasoning has become a current technological bottleneck.

[0004] In embodied intelligent systems that collaborate across multiple protocols and devices, communication between agents and robot behavior execution instructions often exist in different protocol formats, including status reporting protocols, control instruction protocols, and task allocation protocols. Due to the complexity and sheer number of protocol formats, single-point parsing is insufficient to meet application requirements. Therefore, there is an urgent need for an efficient mechanism that can automatically perform protocol feature extraction, protocol modeling, protocol identification, and parser scheduling to improve the system's real-time performance and scalability.

[0005] To address the challenges of diverse protocol types and significant content variations, data-driven approaches are emerging as a new technological trend. Accurate identification can be achieved by collecting communication data and combining feature extraction with protocol modeling to generate "protocol fingerprints." However, current technologies rely on manual feature selection or fixed patterns, lacking the generalization ability of protocol models and exhibiting a lack of intelligent collaboration between protocol identification and parsing, making it difficult to handle system scaling. Furthermore, embodied intelligent systems require the transformation of parsing results into task understanding, behavioral reasoning, and action generation, while traditional decision-making methods often rely on fixed rules or limited knowledge bases, failing to effectively address the reasoning needs of complex tasks.

[0006] Designing an integrated system that combines data acquisition, feature extraction, protocol modeling, protocol recognition agent, protocol scheduling agent, and behavior analysis capabilities of a retrieval-enhanced generative architecture agent to achieve automatic protocol recognition, dynamic scheduling and parsing, and intelligent analysis of robot behavior is a key problem that current technology urgently needs to solve.

[0007] Therefore, it is necessary to design a multi-agent collaborative embodied intelligent protocol parsing and behavior verification method to solve the problems of insufficient protocol processing capabilities, lack of intelligent collaboration in the parsing process, and lack of knowledge support for robot behavior analysis in the existing technology. Summary of the Invention

[0008] In view of this, this invention proposes a multi-agent collaborative embodied intelligence protocol parsing and behavior verification method, which can achieve efficient and intelligent protocol identification and parsing in a multi-protocol environment, and support intelligent reasoning of robot behavior, thereby improving the system's intelligence and adaptability. It aims to solve problems in existing technologies such as insufficient protocol processing capabilities, lack of intelligent collaboration in the parsing process, and lack of knowledge support for robot behavior analysis.

[0009] This invention proposes a method for parsing and verifying embodied intelligence protocols in multi-agent collaboration, applicable to embodied intelligence systems. The method includes: Define intelligent agents corresponding to protocol identification, protocol scheduling, and behavior reasoning; the protocol identification intelligent agent is used to compare and determine protocol types, the protocol scheduling intelligent agent is used to select an appropriate protocol parsing tool and execute scheduling, and the behavior reasoning intelligent agent is used to associate external knowledge and generate action execution plans; and define a protocol fingerprint for protocol feature representation, which is constructed based on the fusion of statistical features and semantic features, and is used to describe the differences in structure and field patterns of different protocol types; define dual verification rules for behavior verification, which include semantic logic consistency verification rules and physical environment constraint inference rules, the semantic logic consistency verification rules are constructed based on the temporal and mutual exclusion relationships of action instructions, and the physical environment constraint inference rules are constructed based on the spatial relationship between the robot's motion state and environmental obstacles; Based on the aforementioned intelligent agents, protocol fingerprints, and dual verification rules, a multi-agent collaborative parsing and behavior verification model is constructed. By utilizing historical communication data, we complete the construction and storage of protocol fingerprints, as well as the adaptation and debugging of multi-agent collaborative parsing and behavior verification models, to obtain a target collaborative model that can adapt to multi-protocol environments and output compliant action execution schemes. When the embodied intelligent system is running online, it captures real-time communication messages between the system and external devices, extracts statistical and semantic features of the communication messages, and fuses them to form a comprehensive feature vector. The comprehensive feature vector is compared with a stored protocol fingerprint, and the protocol type is determined by a protocol identification agent and transmitted to a protocol scheduling agent. The protocol scheduling agent selects an appropriate protocol parsing tool to perform structured parsing of the communication messages, obtaining structured semantic data. The structured semantic data is input into a behavior reasoning agent, which correlates with an external knowledge base to generate a robot action execution plan. The action execution plan is verified and deduced using a dual verification rule, triggering regeneration when an anomaly is detected. Updated system communication data and action execution feedback information are collected, and protocol parsing and behavior verification are performed cyclically to achieve dynamically optimized protocol processing and behavior control.

[0010] Furthermore, the intelligent agents defined corresponding to protocol identification, protocol scheduling, and behavior reasoning include: The protocol recognition agent takes feature comparison as its core function. Its input includes a comprehensive feature vector formed by fusing statistical feature vectors and semantic feature vectors of communication messages. By calculating the similarity with the stored protocol fingerprint, it outputs the protocol type label and corresponding confidence information. The protocol scheduling agent takes parsing tool matching as its core function. Its inputs include the protocol type label output by the protocol identification agent, the version adaptation information of each protocol parsing tool, the semantic parsing completeness record, and the current running load data. Through multi-dimensional evaluation and screening, it outputs the corresponding protocol parsing tool identifier and scheduling instructions. The behavioral reasoning agent takes action plan generation as its core function. Its inputs include structured semantic data, robot motion parameters from an external knowledge base, and knowledge fragments related to safety regulations and environmental constraints. Through associative reasoning and logical combination, it outputs robot action execution plans and corresponding parameter sets.

[0011] Furthermore, the construction of the protocol fingerprint includes: Extract statistical features from communication messages, including field occurrence frequency, field length distribution, byte entropy value, overall message length pattern, and field position statistics. Extract semantic features from communication messages, including inter-byte relationships, field semantic block patterns, and instruction logical association information; The statistical features and semantic features are concatenated or weighted to form a feature vector, which is then standardized and mapped to a fixed-dimensional feature representation through feature embedding. Based on the aforementioned feature representation, the differences in structure and field patterns of different protocol types are identified, a unique corresponding protocol fingerprint is generated, and the fingerprint is stored in the fingerprint database according to the protocol type. When the similarity between the comprehensive feature vector corresponding to a new communication message and all protocol fingerprints in the fingerprint database is lower than a set threshold, it is marked as an unknown protocol. A new protocol fingerprint candidate is generated based on the feature vector of the message and added to the fingerprint database after verification.

[0012] Furthermore, the construction of the semantic logic consistency verification rules includes: Obtain the temporal dependencies of various robot action commands, clarify the correspondence between actions executed first and actions executed later, and form a list of temporal constraints; Identify mutually exclusive action instruction combinations, clarify action instruction pairs that cannot be executed simultaneously or consecutively, and form a list of mutual exclusion constraints. A logical association matrix is ​​constructed based on the time constraint list and the mutual exclusion constraint list. The rows and columns of the logical association matrix correspond to various action instructions, and the matrix elements identify the compatibility status of the corresponding action instruction combinations. The instruction sequence in the action execution plan is compared with the logical association matrix. If there is an instruction combination that violates the timing constraint or mutual exclusion constraint, it is judged as a semantic logic anomaly.

[0013] Furthermore, the construction of the physical environment constraint deduction rules includes: Obtain the robot's kinematic parameters, including joint range of motion, end effector trajectory limits, and motion speed limits; Establish an environmental obstacle model to clarify the spatial location, size, and distribution information of the obstacles; The motion parameters in the motion execution scheme are mapped to the robot end effector pose sequence, which includes the spatial coordinates and attitude angles of the end effector at each moment. Calculate the Euclidean distance between each coordinate point in the pose sequence and the surface of the environmental obstacle model, and preset a safe distance threshold. If the Euclidean distance corresponding to any coordinate point is less than the preset safe distance threshold, it is determined to be a physical constraint violation.

[0014] Furthermore, the adaptation and debugging of the multi-agent collaborative parsing and behavior verification model using historical communication data includes: A simulation test environment is constructed based on historical communication data. The environment simulates the transmission of communication messages of multiple protocol types, the operating status of different protocol parsing tools, and the physical scenario of robot action execution. Historical communication data is categorized by protocol type to construct a test dataset, which includes communication messages in different protocol formats, corresponding standard parsing results, and compliance action execution plans. The test dataset is input into the multi-agent collaborative parsing and behavior verification model, driving the protocol recognition agent, protocol scheduling agent and behavior reasoning agent to run in sequence, and collecting the output results, parsing accuracy and action plan compliance rate data of each agent. Based on the collected data, the model parameters are adjusted, including the feature weights of the protocol fingerprint, the evaluation index thresholds of the agent, and the inference logic parameters, and the testing and adjustment are carried out iteratively. When the model's protocol recognition accuracy, parsing tool matching accuracy, and action plan compliance rate all reach the preset standards, and there are no significant fluctuations in the number of consecutive preset tests, the model adaptation and debugging are deemed complete, and the current parameters are solidified to form the target collaborative model.

[0015] Furthermore, the step of extracting the statistical and semantic features of the communication message and fusing them to form a comprehensive feature vector includes: The communication messages are preprocessed, including field splitting, invalid information removal and format standardization, to obtain standardized message data; Statistical features are extracted from standardized message data, and feature parameters such as field frequency and length distribution are obtained through statistical calculations to form a statistical feature vector. By analyzing the relationships between bytes, the semantic relationships between fields, and the instruction logic in the normalized message data, semantic feature parameters are extracted to form a semantic feature vector; The statistical feature vector and the semantic feature vector are aligned according to a preset dimension and fused using a weighted summation or concatenation method to generate a comprehensive feature vector with a fixed dimension. The weight coefficients of the weighted summation are set based on the contribution of different features to protocol recognition.

[0016] Furthermore, the triggering of regeneration upon detecting an anomaly includes: When the semantic logic consistency check determines that there is an anomaly, extract the instruction combination, timing conflict point and mutual exclusion constraint violation information associated with the anomaly to form a logical anomaly feature; When physical environment constraint deduction determines that there is an anomaly, the pose coordinates of the out-of-bounds position, the corresponding motion parameters, and the distance data to the obstacle are extracted to form physical anomaly features; The logical or physical anomaly features are used as constraint information and fed back to the behavioral reasoning agent, while also being transmitted to the protocol scheduling agent to confirm whether there are any deviations in the parsing process. The behavioral reasoning agent adjusts the reasoning boundary based on constraint information, re-retrieves relevant knowledge fragments in conjunction with the external knowledge base, and regenerates the action execution plan within the scope of excluding abnormal constraints. The double verification rule is applied again to the regenerated action execution plan until the plan passes the verification or the preset maximum number of regeneration attempts is reached.

[0017] Furthermore, it also includes vectorizing and storing the verified action execution plan and abnormal amendment examples, specifically including: The verified action execution schemes are structured, and the core action parameters, corresponding protocol types, environmental constraints and execution result feedback information are extracted and converted into standardized data formats. The abnormal amendment cases are broken down to extract the original abnormal action plan, abnormal type, constraint information, correction strategy and final compliance plan, forming structured case data; The standardized data format action execution plan and case structured data are transformed into fixed-dimensional vector data by using feature embedding. The vector data contains the core features and related information of the plan. Vector data is categorized by data type and stored in a dedicated database, with indexes established. These indexes include key fields such as protocol type, action scenario, and exception type, which are used for rapid retrieval and retrieval during subsequent reasoning processes.

[0018] Furthermore, it also includes dynamically optimizing and adjusting the target collaboration model, specifically including: Real-time monitoring of the operational metrics of the target collaboration model, including protocol recognition accuracy, parsing time, action plan compliance rate, and anomaly correction success rate; Set a baseline threshold range for the operating metrics. When any operating metric is detected to exceed the baseline threshold range and continues for a preset duration, the model will be dynamically optimized. If the accuracy of protocol recognition decreases due to the addition of new protocol types, features are extracted based on the communication data of the new protocol to generate a new protocol fingerprint and update it to the fingerprint database. At the same time, the comparison parameters of the protocol recognition agent are adjusted. If the anomaly rate of physical constraint deduction increases due to changes in environmental constraints, update the environmental obstacle model and robot kinematic parameters, and adjust the safety distance threshold in the physical environmental constraint deduction rules. If the parsing time is prolonged due to changes in the performance of the parsing tool, optimize the parsing tool evaluation metrics and selection logic of the protocol scheduling agent to ensure that scheduling efficiency matches parsing results.

[0019] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention constructs protocol fingerprints by capturing communication messages and fusing statistical and semantic features, achieving accurate identification of multiple protocol types and dynamic adaptation of unknown protocols. It leverages intelligent agent collaboration to select and adapt parsing tools for structured parsing, generates action execution schemes by combining external knowledge bases, and forms a closed-loop anomaly correction mechanism through semantic logic consistency verification and physical environment constraint deduction. This effectively solves the problems of poor multi-protocol adaptability, insufficient parsing scheduling coordination, and low action execution security in traditional methods, significantly improving protocol identification accuracy and parsing efficiency, ensuring the compliance and security of robot action execution. Simultaneously, it supports dynamic system optimization through data feedback, enhancing the robustness, adaptability, and autonomous evolution capabilities of embodied intelligent systems in complex multi-protocol environments. It is suitable for efficient and stable operation in various scenarios such as intelligent manufacturing and robot control. Attached Figure Description

[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart illustrating the embodied intelligence protocol parsing and behavior verification method for multi-agent collaboration provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the embodied intelligence protocol parsing and behavior verification method for multi-agent collaboration provided in an embodiment of the present invention. Detailed Implementation

[0021] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0022] See Figure 1-2 As shown in some embodiments of this application, this embodiment provides a method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration, applied to embodied intelligence systems, including the following steps: S100: Define the intelligent agents corresponding to protocol identification, protocol scheduling, and behavior reasoning; the protocol identification intelligent agent is used to compare and determine the protocol type, the protocol scheduling intelligent agent is used to select the appropriate protocol parsing tool and execute the scheduling, and the behavior reasoning intelligent agent is used to associate external knowledge and generate action execution plans; and define the protocol fingerprint for protocol feature representation, which is constructed based on the fusion of statistical features and semantic features, and is used to describe the differences in structure and field patterns of different protocol types; define the dual verification rules for behavior verification, which include semantic logic consistency verification rules and physical environment constraint inference rules, the semantic logic consistency verification rules are constructed based on the temporal and mutual exclusion relationships of action instructions, and the physical environment constraint inference rules are constructed based on the spatial relationship between the robot's motion state and environmental obstacles.

[0023] It is worth noting that the communication data of the target system, including communication messages such as interaction messages and data exchange content between devices, are obtained as the raw dataset for the protocol identification agent. This ensures the integrity and representativeness of the data and provides a foundation for subsequent feature extraction. Statistical and semantic features of the raw dataset are extracted to form a feature set. Based on the feature set, a protocol fingerprint representing the protocol structure features is constructed. The protocol identification agent compares and infers the protocol fingerprint with the data to be identified, identifies the protocol type to which the data belongs, and passes the identification result as input to the protocol scheduling agent. The protocol scheduling agent selects the corresponding parser from multiple protocol parsers according to the identified protocol type and schedules the corresponding parser to perform protocol content parsing on the data to be identified, generating structured semantic data with clear field meanings and a unified format. The parsed structured semantic data is input into the retrieval-enhanced generation framework agent, i.e., the behavior reasoning agent. The behavior reasoning agent combines external knowledge retrieval to analyze the task semantics, action intentions, and behavior planning of the embodied intelligent system robot, generating a corresponding action execution plan. The dual verification rule is used to perform closed-loop auditing and security verification of the behavior logic after protocol parsing. When an anomaly is detected, it triggers regeneration and stores the verified scheme and amendment in a vectorized manner. Protocol parsing and behavior verification are performed cyclically to achieve dynamically optimized protocol processing and behavior control.

[0024] It's worth noting that communication data can be acquired non-destructively using the switch's port mirroring function or by connecting a logic analyzer probe in parallel on the controller's local area network. This communication data includes robot control commands, sensor data, and status feedback. The acquisition process can be achieved through network sniffing, protocol analysis tools, or hardware interfaces. Interaction messages include key information such as request packets, status update messages, and task feedback from the communication protocol. For each message packet, predefined filtering rules are used to identify control commands and robot status, discarding irrelevant data.

[0025] It is worth noting that data is continuously acquired until the number of samples meets the requirements for protocol recognition training.

[0026] Specifically, the intelligent agents corresponding to protocol identification, protocol scheduling, and behavior reasoning are defined, including: The protocol recognition agent takes feature comparison as its core function. Its input includes a comprehensive feature vector formed by fusing statistical feature vectors and semantic feature vectors of communication messages. By calculating the similarity with the stored protocol fingerprint, it outputs the protocol type label and corresponding confidence information. The protocol scheduling agent takes parsing tool matching as its core function. Its inputs include the protocol type label output by the protocol identification agent, the version adaptation information of each protocol parsing tool, the semantic parsing completeness record, and the current running load data. Through multi-dimensional evaluation and screening, it outputs the corresponding protocol parsing tool identifier and scheduling instructions. The behavioral reasoning agent takes action plan generation as its core function. Its inputs include structured semantic data, robot motion parameters from an external knowledge base, and knowledge fragments related to safety regulations and environmental constraints. Through associative reasoning and logical combination, it outputs robot action execution plans and corresponding parameter sets.

[0027] Understandably, in the protocol recognition agent, the statistical feature vector of the communication message covers low-dimensional quantitative indicators such as message length distribution, field frequency, and time interval sequence, while the semantic feature vector is generated through deep encoding of message field names, keywords, numerical ranges, and contextual relationships. The comprehensive feature vector formed by the weighted fusion of the two can comprehensively characterize the essential attributes of the message, ensuring high accuracy and robustness in similarity calculation with protocol fingerprints. When performing multi-dimensional evaluation, the protocol scheduling agent not only considers the basic compatibility of the parsing tool with the protocol type, but also uses historical parsing data to score the semantic parsing completeness of the tool for the same type of protocol (such as field extraction accuracy and nested structure parsing success rate). Combined with the current CPU utilization, memory consumption, and task queue length of each tool, an evaluation matrix is ​​constructed using the analytic hierarchy process (AHP) to ultimately select the parsing tool with the best overall performance. In the process of generating action plans, the behavioral reasoning agent first links and aligns the structured semantic data with knowledge fragments in the external knowledge base. For example, it establishes a mapping relationship between the parsed "target position coordinates" and the "joint space parameters" in the robot's kinematic model. At the same time, it strictly follows the hard constraints in the safety specifications such as "mechanical arm working radius limit" and "end effector maximum speed threshold". Through a combination of rule-based reasoning and case-based reasoning, it dynamically combines basic action units to form a complete set of action parameters, including motion trajectory, execution timing, force feedback parameters, etc., to ensure the accuracy and safety of robot actions.

[0028] Specifically, the construction of the protocol fingerprint includes: Extract statistical features from communication messages, including field occurrence frequency, field length distribution, byte entropy value, overall message length pattern, and field position statistics. Extract semantic features from communication messages. Semantic features include inter-byte relationships, field semantic block patterns, and instruction logical relationships. Statistical features and semantic features are combined or weighted to form a feature vector, which is then standardized and mapped to a fixed-dimensional feature representation through feature embedding. Based on feature representation, the differences in structure and field patterns of different protocol types are identified, and a unique corresponding protocol fingerprint is generated and stored in the fingerprint database according to protocol type. When the similarity between the comprehensive feature vector corresponding to a new communication message and all protocol fingerprints in the fingerprint database is lower than a set threshold, it is marked as an unknown protocol. A new protocol fingerprint candidate is generated based on the feature vector of the message and added to the fingerprint database after verification.

[0029] It is worth noting that the format of the fused feature vector data is standardized so that message data inputs of different lengths and structures meet the unified format requirements for protocol fingerprint construction, including length alignment, padding strategy, field position encoding and other processing methods.

[0030] Understandably, statistical features capture the surface-level structural patterns of protocols. For example, the distribution of fields in a message can reflect the protocol's format specifications, and byte entropy values ​​can reflect the randomness or regularity of the data payload. Semantic features delve deeper into the inherent logical relationships between protocol fields. For instance, a "start" command is usually accompanied by a parameter configuration field; this semantic binding between commands and parameters is key to distinguishing different control protocols. Feature embedding mapping, through dimensionality reduction or enhancement, transforms high-dimensional sparse original features into low-dimensional dense vector representations. This preserves core feature information while reducing the complexity of subsequent similarity calculations, allowing fingerprints of different protocol types to form distinguishable clusters in the vector space. When a new message arrives, the initial protocol type can be quickly determined by calculating the cosine similarity or Euclidean distance between its feature vector and the center vectors of each category in the fingerprint database. The dynamic discovery mechanism for unknown protocols ensures the self-updating capability of the fingerprint database, adapting to emerging new protocols or variants, and providing a continuous and reliable feature benchmark for protocol parsing in complex communication environments for multi-agent systems.

[0031] Specifically, the construction of semantic logic consistency verification rules includes: Obtain the temporal dependencies of various robot action commands, clarify the correspondence between actions executed first and actions executed later, and form a list of temporal constraints; Identify mutually exclusive action instruction combinations, clarify action instruction pairs that cannot be executed simultaneously or consecutively, and form a list of mutual exclusion constraints. A logical association matrix is ​​constructed based on the timing constraint list and the mutual exclusion constraint list. The rows and columns of the logical association matrix correspond to various action instructions, and the matrix elements identify the compatibility status of the corresponding action instruction combinations. The instruction sequence in the action execution plan is compared with the logical association matrix. If there is an instruction combination that violates the timing constraint or mutual exclusion constraint, it is judged as a semantic logic anomaly.

[0032] It is worth noting that temporal dependency refers to the necessary logical relationship of "first execution - last execution" between action instructions when a robot completes a specific task. That is, the initiation of the later action must be a prerequisite for the completion of the earlier action. The temporal constraint list is a standardized list that records the temporal dependencies of all action instructions, including the name of the action to be executed first, the name of the action to be executed later, and the dependent conditions (such as no additional conditions or the requirement to meet specific states).

[0033] It is understandable that a robot's task execution is essentially a combination of actions linked together in a logical sequence. Violating the temporal dependencies will lead to task failure or logical paradoxes (such as performing a placement action without grasping an object). By clearly defining the temporal dependencies and establishing the logical order boundaries of action execution, we can ensure that the action sequence conforms to the inherent logic of task completion.

[0034] It is worth noting that a mutually exclusive action instruction combination refers to a combination of two or more types of action instructions that cannot be executed simultaneously (parallel mutual exclusion) or sequentially (serial mutual exclusion) due to mechanical structure limitations, logical contradictions, or other reasons. A mutual exclusion constraint list is a standardized list recording all mutually exclusive action instruction combinations, including information such as the mutual exclusion action group, the mutual exclusion type (parallel / serial), and the mutual exclusion reason.

[0035] Understandably, robots have inherent limitations in their mechanical structure (such as joint range of motion and end effector trajectory) and motion logic: parallel mutual exclusion stems from the fact that the mechanical structure cannot simultaneously complete two conflicting actions (such as the left arm swinging to the left and right); sequential mutual exclusion stems from the fact that the state of a previous action can prevent the execution of a subsequent action (such as being unable to "release material" again without re-grabbing after executing "release material"). Identifying mutually exclusive combinations can prevent mechanical failures or logical confusion caused by conflicting motion commands.

[0036] It is worth noting that the logical association matrix is ​​a two-dimensional matrix constructed with all robot action commands as rows and columns. The matrix elements are specific identifiers (such as 0, 1, 2) used to intuitively represent the compatibility state of any two action command combinations.

[0037] Understandably, the temporal constraint list and the mutual exclusion constraint list are linear lists, requiring a sequential search when directly comparing action sequences, which is inefficient. The logical association matrix, through "coordinate positioning," transforms the constraint relationship between any two actions into matrix elements, enabling rapid querying of the constraint state of action combinations, improving verification efficiency, and reducing the complexity of the verification logic.

[0038] The construction of semantic logic consistency verification rules involves the following steps: First, identify the core task types of the embodied intelligent system, such as "grabbing-handling-assembly" in intelligent manufacturing and "movement-interaction-execution" in service robots. For each task type, decompose the complete action execution chain. For example, the action chain for "material assembly task" is "move to material location → grab material → move to assembly location → assemble material → reset". Identify "first-last" dependency pairs from the action chain. For example, "grabbing material" must precede "moving to assembly location", and "assembling material" must precede "reset". Organize all dependency pairs in the format of "execute action first → execute action last", supplement the dependency conditions (e.g., the dependency condition for "grabbing material" is "the end effector reaches the material grabbing coordinates"), and form a temporal constraint list.

[0039] Based on the robot's mechanical design drawings and kinematic characteristics, identify actions that cannot be performed simultaneously due to structural interference, such as "clockwise rotation of the waist" and "counterclockwise rotation of the waist" (parallel and mutually exclusive), and "end-effector clamping" and "end-effector releasing" (parallel and mutually exclusive). Combined with task logic, identify action combinations that do not have structural conflicts but are logically contradictory, such as "rapid movement" and "precise positioning" (parallel and mutually exclusive, rapid movement will affect positioning accuracy), and "placing material" and "grabbing the same material" (serial and mutually exclusive, the material position changes after placement, and repositioning is required before grasping). Verify the initially identified mutually exclusive combinations through simulation testing or actual operation, and eliminate false mutual exclusions (such as actions that seem to conflict but can actually be achieved through trajectory optimization). Organize the verified mutually exclusive combinations in the format of "Action A - Action B (mutual exclusion type: XX, mutual exclusion reason: XX)" to form a mutual exclusion constraint list.

[0040] Summarize all robot action commands (such as grasping, releasing, moving, rotating, and positioning). Assuming the total number of actions is N, construct an N×N logical association matrix, where both row and column indices are action command names. Define the meaning of matrix elements, for example: 1 represents "compatible" (no timing constraints and not mutually exclusive; actions can be executed in any order or simultaneously), 0 represents "timing violation" (column actions must be executed before other actions; if an action is executed first, the constraint is violated), and 2 represents "mutually exclusive" (two actions cannot be executed simultaneously or consecutively). Fill the matrix one by one, referring to the timing constraint list and the mutual exclusion constraint list. If an action is the action that precedes a column action, and the column action is the action that follows an action, then the matrix element (row, column) = 1, and the matrix element (column, row) = 0. If the action and column actions are mutually exclusive combinations, then the matrix element (row, column) = matrix element (column, row) = 2; If two actions have no temporal dependency and are not mutually exclusive, then the matrix element (row, column) = the matrix element (column, row) = 1; Traverse the matrix, verify the consistency of element filling (such as the bidirectionality of temporal dependencies and the symmetry of mutual exclusion relationships), correct filling errors, and form the final logical association matrix.

[0041] The action execution plan is broken down into an ordered sequence of action instructions along the time axis (e.g., [A→B→C→D]). If there are parallel actions (e.g., A and E are executed simultaneously), the parallel action group (A, E) is extracted separately. For consecutive action pairs in a sequence (e.g., A→B, B→C, C→D), query the corresponding element in the logical association matrix (row = previous action, column = next action): If the element is 0, it indicates a violation of the timing constraint (the subsequent action should be executed before the previous action), and is marked as a semantic logic exception; If the element is 2, it indicates a violation of the mutual exclusion constraint (two actions cannot be executed consecutively), and is marked as a semantic logic exception; If the element is 1, it means the logic is compatible, and proceed to the next comparison step; Parallel action comparison: For a group of parallel actions (e.g., A, E), query the matrix elements (A, E) and (E, A): If any element is 2, it means that the two actions are mutually exclusive and is marked as a semantic logic exception; If all elements are 1, it means that parallel execution is possible and the comparison is successful; Exception summary: Record all marked abnormal action combinations and exception types (timing violation / mutual exclusion violation) to form a semantic logic exception report.

[0042] Understandably, the instruction sequence in the action execution plan is a combination of actions arranged in chronological order. By comparing adjacent and parallel actions in the sequence one by one with the logical association matrix, it is possible to quickly identify whether there are combinations that violate timing constraints or mutual exclusion constraints. "Using the matrix as the standard and matching with the sequence," and replacing linear retrieval with coordinate queries, the efficiency and accuracy of anomaly detection are improved.

[0043] Specifically, the construction of physical environment constraint deduction rules includes: Obtain the robot's kinematic parameters, including joint range of motion, end effector trajectory limits, and motion speed limits; Establish an environmental obstacle model to clarify the spatial location, size, and distribution information of the obstacles; The motion parameters in the motion execution plan are mapped to the robot end effector pose sequence, which includes the spatial coordinates and attitude angles of the end effector at each moment. Calculate the Euclidean distance between each coordinate point in the pose sequence and the surface of the environmental obstacle model, and preset a safe distance threshold. If the Euclidean distance corresponding to any coordinate point is less than the preset safe distance threshold, it is determined to be a physical constraint violation.

[0044] It is worth noting that kinematic parameters are key parameters describing the motion characteristics of a robot's mechanical structure. They are the physical boundaries of the robot's motion execution. The range of motion of a joint is the angular range (e.g., -90° to 90°) that each movable joint of the robot (such as the shoulder, elbow, and wrist joints) can rotate. The end effector trajectory limit is the spatial range boundary that the robot's end effector (such as a manipulator or gripper) can reach (e.g., 0 to 1000 mm for the X-axis, 0 to 800 mm for the Y-axis, and 0 to 500 mm for the Z-axis). The motion speed limit is the maximum value of the robot's joint rotation speed and end effector movement speed (e.g., maximum joint rotation speed of 60° / s and maximum end effector movement speed of 500 mm / s).

[0045] Understandably, a robot's movement is limited by the physical characteristics of its mechanical structure. Actions exceeding the range of kinematic parameters can lead to joint damage, motor overload, or end effector trajectory deviation. Obtaining kinematic parameters is the foundation for constructing physical constraints, clarifying the "physically feasible domain" of robot movements, and preventing equipment failures caused by exceeding motion parameter limits.

[0046] It is worth noting that the environmental obstacle model is a three-dimensional digital model constructed based on the spatial information of obstacles in the physical environment. It is used to accurately represent the position, size, and distribution of obstacles and is the core basis for collision detection. Spatial position refers to the three-dimensional coordinates (X, Y, Z) of the obstacle in a preset coordinate system (such as the robot's base coordinate system), usually based on the geometric center or bottom center point of the obstacle; size information refers to the length, width, and height of the obstacle (for regular obstacles) or the coordinates of key points on its surface contour (for irregular obstacles); distribution information refers to the relative positional relationships of multiple obstacles in the environment.

[0047] It is understandable that the physical environment in which robots perform actions contains various obstacles (such as machine tools and shelves in industrial settings, and furniture and walls in service settings). If the robot's trajectory intersects with an obstacle, it can lead to collision damage. The core of building an environmental obstacle model is to "digitize" the physical environment, enabling a quantitative comparison between the spatial position of the robot's actions and the spatial position of the obstacles, thereby predicting the risk of collision.

[0048] It is worth noting that motion parameters are the explicitly defined robot motion control parameters in the motion execution plan, such as joint rotation angles, motion speed, motion time, and target position coordinates; the end effector pose sequence is a set of spatial states of the robot end effector at various moments, arranged in chronological order. Each state includes spatial coordinates (Xt, Yt, Zt) and attitude angles (pitch angle αt, yaw angle βt, roll angle γt), where t is a time node (e.g., t=0.1s, t=0.2s). Attitude angles describe the orientation of the end effector in space; pitch angle is the rotation angle around the X-axis, yaw angle is the rotation angle around the Y-axis, and roll angle is the rotation angle around the Z-axis.

[0049] It is understandable that the essence of robot motion execution is joint movement driving the movement of the end effector. There is a clear mathematical mapping relationship (i.e., forward kinematics relationship) between motion parameters (such as joint angles) and the spatial state (pose) of the end effector. Through this mapping relationship, abstract motion parameters can be transformed into concrete end effector poses, thereby enabling visualization and quantitative analysis of motion trajectories and providing specific comparison objects for collision detection.

[0050] It is worth noting that the safety distance threshold is a preset minimum safety distance value (such as 0.05m or 0.1m) to avoid collisions between the end effector and obstacles. It needs to be set according to factors such as the size of the robot end effector, the surface characteristics of the obstacle, and the movement speed. Physical constraint violation occurs when the Euclidean distance corresponding to the pose of the end effector is less than the safety distance threshold, that is, there is a risk of collision or the kinematic parameter limit is exceeded.

[0051] Understandably, the Euclidean distance between the end effector and the obstacle directly reflects their spatial distance: the greater the distance, the lower the risk of collision; when the distance is 0, a direct collision occurs; when the distance is less than the safe distance threshold, although there is no direct collision, a collision may occur due to factors such as motion errors or protrusions on the obstacle surface. By calculating the Euclidean distance and comparing it with the safe threshold, it is possible to accurately predict whether physical constraints have been exceeded and to avoid collision risks in advance.

[0052] The specific implementation process for constructing the physical environment constraint deduction rules is as follows: Extract initial values ​​of the kinematic parameters from the product manuals and technical specifications provided by the robot manufacturer; issue calibration commands through the robot control system to drive each joint to rotate at a preset step size and the end effector to move along a preset path, and collect actual motion data using devices such as encoders and laser positioning instruments: record the minimum and maximum angles at which each joint can rotate stably, and correct the theoretical values ​​in the factory parameters (e.g., factory specification -90°~90°, actual calibration is -85°~88°); drive the end effector to move along the X, Y, and Z axes to a position where it can no longer move forward, record the coordinate values, and determine the actual trajectory limit range; gradually increase the joint rotation speed and end effector movement speed, and record the maximum speed when the equipment has no vibration and no overload alarm, as the motion speed limit value; classify and organize the calibrated kinematic parameters into a standardized kinematic parameter table according to the format of "joint name-range of motion", "end effector-trajectory limit" and "motion speed-limit value".

[0053] Using devices such as LiDAR, visual sensors (e.g., depth cameras), and ultrasonic sensors, the robot's operating environment is scanned from all angles to collect spatial data on obstacles. Collect the coordinates (X0, Y0, Z0) and length (L), width (W), and height (H) of the center point of the bottom surface of regular obstacles (such as cubic shelves); Collect the key point coordinates (such as several (Xi,Yi,Zi)) of the surface contour of irregular obstacles (such as irregularly shaped workpieces and equipment protrusions); The collected obstacle data is converted to a coordinate system consistent with the robot's kinematic parameters (such as the robot's base coordinate system) to ensure the comparability of position data; using CAD software, point cloud processing tools, etc., a three-dimensional digital model of the obstacle is constructed based on the collected data. For regular obstacles, a standard geometric model (such as a cuboid or cylinder) is generated by "center coordinates + dimensions"; for irregular obstacles, a surface contour model is generated by fitting key point cloud data. The 3D models of all obstacles are integrated into the same environmental coordinate system according to their actual distribution locations, forming a complete environmental obstacle model that supports spatial location query and distance calculation.

[0054] Extract specific parameters for each action from the action execution plan, such as the target coordinates (X target, Y target, Z target), movement speed v, and movement time T for the "movement action"; and the joint rotation angle Δθ and rotation speed ω for the "rotation action". Based on the motion time T and accuracy requirements, several time nodes are divided (e.g., T=1s, divided into 10 time nodes t0~t9 at 0.1s intervals). Based on the robot's mechanical structure (such as link length and joint type), establish the forward kinematic equations for the motion parameters to the pose, for example: For the movement motion, the coordinates of the end effector change with time as Xt = Xinitial + (Xtarget - Xinitial) × (t / T), and Yt and Zt are similar; For rotational motion, the mapping equation between joint angles and end-effector attitude angles is established using the DH parameter method, and αt, βt, and γt are calculated at each time point. Generate pose sequence: Substitute the time t of each time node into the forward kinematics equation to calculate the corresponding spatial coordinates (Xt, Yt, Zt) and attitude angles (αt, βt, γt), and arrange them in time order to form the end pose sequence.

[0055] Extract the end effector spatial coordinates (Xt, Yt, Zt) for each time point from the end effector pose sequence. Extract the surface key point coordinates (e.g., surface vertices of regular obstacles, contour points of irregular obstacles) from the environmental obstacle model. For each end effector coordinate (Xt, Yt, Zt), calculate its Euclidean distance to all obstacle surface key points, and take the minimum value as the actual distance Dt between the end effector and the obstacle at that time point. Combined with the robot end effector size (e.g., gripper diameter 0.03m), movement speed (the faster the speed, the larger the safety distance needs to be), and obstacle surface hardness (hard surfaces require a larger safety distance), set a safety distance threshold D0 (e.g., 0.1m). Compare the actual distance Dt at each time point with the safety threshold D0. If Dt≥D0, it means that the end effector maintains a safe distance from the obstacle, and the physical constraint is satisfied; If Dt < D0, it indicates a collision risk and is judged as a physical constraint violation. If the end-effector coordinates exceed the limits of the end-effector's motion trajectory, it is directly determined as a physical constraint violation. Record the time of the boundary violation, the coordinates of the end point, the corresponding obstacle, the difference between the actual distance and the safety threshold, and generate a physical constraint boundary violation report.

[0056] S200: Based on intelligent agents, protocol fingerprints, and dual verification rules, a multi-agent collaborative parsing and behavior verification model is constructed.

[0057] Understandably, the output of the protocol recognition agent is connected to the input of the protocol scheduling agent, the output of the protocol scheduling agent is connected to the protocol parsing tool, the output of the protocol parsing tool is connected to the input of the behavior reasoning agent, the output of the behavior reasoning agent is connected to the input of the dual verification rule, and the feedback of the dual verification rule is connected to the protocol scheduling agent and the behavior reasoning agent respectively, forming an interactive link of "recognition-scheduling-parsing-reasoning-verification-feedback". Data within the model flows along the path of "raw message → feature vector → protocol type → structured semantic data → action execution plan → verification result → feedback data", ensuring that data at each stage can be accurately transmitted to the next stage, providing a basis for subsequent processing.

[0058] S300: Utilizes historical communication data to construct and store protocol fingerprints, as well as adapt and debug multi-agent collaborative parsing and behavior verification models, to obtain a target collaborative model that can adapt to multi-protocol environments and output compliant action execution schemes.

[0059] Specifically, historical communication data is used to adapt and debug the multi-agent collaborative analysis and behavior verification model, including: A simulation test environment is built based on historical communication data. The environment simulates the transmission of communication messages of multiple protocol types, the running status of different protocol parsing tools, and the physical scenario of robot action execution. Historical communication data is categorized by protocol type to construct a test dataset. The dataset contains communication messages in different protocol formats, corresponding standard parsing results, and compliance action execution plans. The test dataset is input into the multi-agent collaborative parsing and behavior verification model, driving the protocol recognition agent, protocol scheduling agent and behavior reasoning agent to run in sequence, and collecting the output results, parsing accuracy and action plan compliance rate data of each agent. Based on the collected data, the model parameters are adjusted, including the feature weights of the protocol fingerprint, the evaluation index thresholds of the agent, and the inference logic parameters, and the testing and adjustment are carried out iteratively. When the model's protocol recognition accuracy, parsing tool matching accuracy, and action plan compliance rate all reach the preset standards, and there are no significant fluctuations in the number of consecutive preset tests, the model adaptation and debugging are deemed complete, and the current parameters are solidified to form the target collaborative model.

[0060] It is worth noting that key scenario features are extracted from historical communication data, including: protocol type and distribution ratio (e.g., control command protocol accounts for 60%, status reporting protocol accounts for 30%, and unknown protocol accounts for 10%), message format (field structure, length range) of various protocols, transmission frequency (e.g., 100 messages per second); historical operation data of parsing tools (e.g., parsing tool version, parsing success rate, and load change curve for different protocols); and physical environment parameters of robot operation (e.g., motion space dimensions, obstacle position coordinates, and joint range of motion).

[0061] It is worth noting that the construction of the simulation test environment includes: building a virtual message generator based on the extracted protocol features, which can generate communication messages of different protocol types according to the historical distribution ratio, simulating the real transmission timing and data characteristics; building a virtual instance of the parsing tool, configuring version attributes, parsing logic, and load response model consistent with the real tool, which can simulate various states such as normal operation of the parsing tool, version incompatibility, and high load; and using a 3D modeling tool to construct a robot motion space and obstacle model, importing historical physical environment parameters, restoring robot kinematic constraints (such as joint range of motion and end-effector trajectory limits), and supporting physical simulation and deduction of the action execution process.

[0062] Understandably, by comparing the consistency between the message data generated in the virtual environment, the running status data of the parsing tool, and the physical scene parameters with historical real data (such as message format matching degree ≥99%, physical scene size error ≤1%), the effectiveness of the simulation test environment is verified, and the credibility of the test results is ensured.

[0063] It is worth noting that historical communication data is filtered to cover all known protocol types, ensure data integrity, and eliminate transmission errors. Duplicate data, incomplete data (such as messages with missing fields), and abnormal data (such as messages with malformed formats) are removed to ensure data quality. The filtered messages are categorized according to protocol type (e.g., control command protocol, sensor data reporting protocol, task allocation protocol), with a separate data subset established for each protocol type. A certain proportion of messages from unknown protocols are also retained (for testing the model's ability to handle unknown protocols). For each protocol type, messages are parsed manually (in conjunction with protocol documentation) or using authoritative parsing tools to obtain structured data such as field meanings, task instructions, and behavioral parameters. This data serves as the standard parsing result and is associated with the corresponding message. Based on the task requirements, robot kinematic parameters, and environmental constraints in the standard parsing result, a motion execution scheme conforming to semantic logic and physical safety is constructed, specifying the motion sequence and execution parameters (e.g., motion speed, joint angles). This serves as a compliance benchmark and is stored in association with the standard parsing result. The dataset for each protocol type is divided into a debug set (for parameter adjustment) and a validation set (for final performance verification) at a preset ratio (e.g., 7:3) to ensure data independence between the debugging and validation processes and avoid overfitting.

[0064] The communication messages of the debug set are input one by one into the multi-agent collaborative parsing and behavior verification model according to the transmission simulation rules of the simulation test environment to ensure that the input timing is consistent with the real environment.

[0065] The protocol identification agent receives messages, extracts feature vectors, compares them with the protocol fingerprint database, and outputs the protocol type and confidence level. The protocol scheduling agent receives protocol type information, queries the parsing tool set and selects an adaptation tool. The scheduling tool parses the messages and outputs structured semantic data. The behavior reasoning agent receives the structured semantic data, retrieves external knowledge bases, and infers and generates action execution plans, outputting the complete plan content. The output results of each agent are recorded in real time, including protocol type identifiers, parsing tool identifiers, and action plan details, and are associated and stored with the corresponding standard results in the test dataset. Record auxiliary data such as response time, agent decision-making time, and parsing tool load during model operation for comprehensive performance evaluation.

[0066] If the protocol identification accuracy is low, analyze erroneous cases: if the statistical features do not fully reflect the differences between protocols, increase the weight of the statistical features; if the similarity threshold is set too high, resulting in missed judgments, decrease the similarity threshold; if the semantic features are not captured sufficiently, increase the weight of the semantic features.

[0067] If the parsing tool has a low matching accuracy, optimize the evaluation metrics of the protocol scheduling agent: if the selection of an incompatible parsing tool is due to an excessively low version compatibility weight, increase the version compatibility weight; if the parsing tool is overloaded due to not considering the load, optimize the load evaluation coefficient.

[0068] If the parsing accuracy is low, check the matching degree between the parsing tool's scheduling logic and the protocol type, and adjust the scoring rules for parsing tool selection (such as increasing the weight of semantic parsing completeness).

[0069] If the compliance rate of the action plan is low, adjust the reasoning logic parameters of the behavioral reasoning agent: if the knowledge retrieval is inaccurate, lower the retrieval matching threshold; if the action parameters exceed the physical constraints, increase the weight coefficient of the physical constraint parameters.

[0070] Based on the analysis results, adjust the model parameters accordingly, modifying only 1-2 types of parameters each time (e.g., first adjust the protocol fingerprint feature weights, then adjust the evaluation threshold) to avoid difficulties in attributing performance fluctuations caused by simultaneous modification of multiple parameters. After parameter adjustment, re-input the debug set into the model and repeat the execution and data collection process in step 3 to calculate new performance indicators. Compare the changes in indicators before and after adjustment to verify the adjustment effect. Continuously repeat the iterative process of "analysis-adjustment-test" until each performance indicator gradually approaches the preset standard.

[0071] Set performance preset standards according to application scenario requirements, such as: protocol recognition accuracy ≥ 98%, parsing tool matching accuracy ≥ 97%, and action plan compliance rate ≥ 99%; set a significant fluctuation range of ±0.5% (i.e., in continuous testing, the difference between the maximum and minimum values ​​of each indicator ≤ 0.5%). Input the validation set into the model and continuously execute the test a preset number of times (e.g., 5 times). Each test is conducted independently, and the performance indicators of each test are recorded.

[0072] If the accuracy of protocol recognition, parsing tool matching, and action plan compliance rate in 5 tests are all greater than or equal to the preset standards, and the fluctuation range of each indicator is less than or equal to 0.5%, the model performance is judged to be stable and meets the standards. If the requirements are not met, analyze the reasons for the failure (such as the validation set containing uncovered protocol features), continue to adjust the parameters and repeat the iterative test until the requirements are met.

[0073] Once the model meets the requirements, record all current model parameters (protocol fingerprint feature weights, agent evaluation thresholds, inference logic parameters, etc.), solidify the parameters into the model configuration file, form the target collaborative model, prohibit subsequent unverified parameter modifications, and ensure the stability of the model during online operation.

[0074] S400: When the embodied intelligent system is running online, it captures real-time communication messages between the system and external devices, extracts statistical and semantic features of the communication messages, and fuses them to form a comprehensive feature vector. The comprehensive feature vector is compared with the stored protocol fingerprint, and the protocol identification agent determines the protocol type and transmits it to the protocol scheduling agent. The protocol scheduling agent selects an appropriate protocol parsing tool to perform structured parsing of the communication messages, obtaining structured semantic data. The structured semantic data is input into the behavior reasoning agent, which correlates with an external knowledge base to generate a robot action execution plan. The action execution plan is verified and deduced using dual verification rules, triggering regeneration when an anomaly is detected. Updated system communication data and action execution feedback information are collected, and protocol parsing and behavior verification are performed cyclically to achieve dynamically optimized protocol processing and behavior control.

[0075] Specifically, the statistical and semantic features of communication messages are extracted and fused to form a comprehensive feature vector, including: The communication messages are preprocessed, including field splitting, invalid information removal and format standardization, to obtain standardized message data; Statistical features are extracted from standardized message data, and feature parameters such as field frequency and length distribution are obtained through statistical calculations to form a statistical feature vector. By analyzing the relationships between bytes, the semantic relationships between fields, and the instruction logic in the normalized message data, semantic feature parameters are extracted to form a semantic feature vector; The statistical feature vector and the semantic feature vector are aligned according to a preset dimension, and then fused using a weighted summation or concatenation method to generate a comprehensive feature vector with a fixed dimension. The weight coefficients of the weighted summation are set based on the contribution of different features to protocol recognition.

[0076] Understandably, during the generation of the comprehensive feature vector, it is necessary to ensure the consistency of the dimensions of the statistical features and the semantic features. For example, both the statistical feature vector and the semantic feature vector are normalized to 128 dimensions, and then the comprehensive feature vector is obtained by element-wise weighted summation (e.g., statistical feature weight 0.4, semantic feature weight 0.6). For the concatenation method, the two vectors are directly connected end to end to form a 256-dimensional comprehensive feature vector. The specific fusion method can be dynamically selected based on the feature importance analysis results of historical data.

[0077] Specifically, regeneration is triggered when an anomaly is detected, including: When the semantic logic consistency check determines that there is an anomaly, extract the instruction combination, timing conflict point and mutual exclusion constraint violation information associated with the anomaly to form a logical anomaly feature; When physical environment constraint deduction determines that there is an anomaly, the pose coordinates of the out-of-bounds position, the corresponding motion parameters, and the distance data to the obstacle are extracted to form physical anomaly features; Logical or physical anomalies are used as constraint information and fed back to the behavioral reasoning agent, while also being transmitted to the protocol scheduling agent to confirm whether there are any deviations in the parsing process. The behavioral reasoning agent adjusts the reasoning boundary based on constraint information, re-retrieves relevant knowledge fragments in conjunction with the external knowledge base, and regenerates the action execution plan within the scope of excluding abnormal constraints. The double verification rule is applied again to the regenerated action execution plan until the plan passes the verification or the preset maximum number of regeneration attempts is reached.

[0078] It's worth noting that setting the preset upper limit for the number of regeneration attempts needs to comprehensively consider both the system's real-time requirements and the success rate of anomaly repair. Typically, based on the complexity of the multi-agent collaborative task and the frequency of dynamic environmental changes, the upper limit is set to 3-5 times. If no effective solution is generated after reaching the upper limit, the protocol scheduling agent will automatically initiate a degradation processing mechanism, prioritizing the stable operation of basic collaborative functions. It will also record the anomaly details and failure process in the system log and trigger an alarm signal to notify maintenance personnel for manual intervention and troubleshooting, thus avoiding excessive consumption of system resources or task execution delays caused by continuous invalid attempts.

[0079] Understandably, the aforementioned anomaly detection and solution regeneration mechanism forms a closed-loop dynamic optimization process. This process is not simply error correction, but rather, through the precise extraction and feedback of abnormal features, it prompts the behavioral reasoning agent to perform knowledge retrieval and solution construction under more reasonable boundary conditions. For example, when semantic logic consistency verification detects a timing conflict such as "robotic arm A has not completed gripping the workpiece, while conveyor belt B has already started conveying," the extracted logical anomaly features, such as the "incomplete gripping instruction state," the "timing point of the conveyor belt starting prematurely," and the "mutual exclusion constraint between workpiece gripping and conveyor belt starting," will cause the behavioral reasoning agent to prioritize the retrieval of knowledge fragments related to "action timing coordination" and "task priority ranking" in subsequent reasoning, and strictly exclude instruction combination patterns that cause the conflict, thereby generating a corrective solution such as "after robotic arm A sends the 'grinding complete' signal, conveyor belt B starts with a 0.5-second delay." Similarly, if the physical environment constraint deduction shows that "in the planned path of mobile robot C, the distance between a certain point (X=5.2m, Y=3.8m) and a static obstacle is only 0.1m, which is less than the safety threshold of 0.3m," then the extracted physical anomaly features, such as the out-of-bounds pose coordinates, the current moving speed parameters, and the actual distance to the obstacle, will guide the behavioral reasoning agent to appropriately expand the influence range of the obstacle when replanning the path, or prioritize the selection of a scheme that includes knowledge of "obstacle avoidance algorithms" and "smooth path transitions," to ensure that the distance between all pose points and obstacles in the newly generated path meets the safety requirements. This mechanism can effectively improve the robustness and reliability of multi-agent cooperative systems in complex dynamic environments. Even if there are certain deviations in the initial protocol parsing or behavioral reasoning, it can gradually approach the optimal and feasible cooperative behavior scheme through continuous verification, feedback, and adjustment, avoiding the failure of the entire cooperative task due to a single anomaly.

[0080] In some embodiments of this application, the method for parsing and verifying embodied intelligence protocols for multi-agent collaboration further includes vectorizing and storing the verified action execution schemes and abnormal amendment examples, specifically including: The verified action execution schemes are structured, and the core action parameters, corresponding protocol types, environmental constraints and execution result feedback information are extracted and converted into standardized data formats. The abnormal amendment cases are broken down to extract the original abnormal action plan, abnormal type, constraint information, correction strategy and final compliance plan, forming structured case data; The standardized data format action execution plan and case structured data are transformed into fixed-dimensional vector data by using feature embedding. The vector data contains the core features and related information of the plan. Vector data is categorized and stored in a dedicated database according to data type, and indexes are established to link them. The indexes include key fields such as protocol type, action scenario, and exception type, which are used for fast retrieval and retrieval during subsequent inference processes.

[0081] Understandably, validated action execution plans contain effective decision-making logic adapted to specific protocol types and environmental constraints, while abnormal amendments contain mature handling strategies for specific anomalies. This historical data is a core source of experience for system optimization. Unstructured data is difficult for machines to quickly identify and reuse. Structured processing can extract core information and unify data formats; feature embedding can transform structured data into a vector form that machines can process efficiently, preserving semantic relationships and feature differences between data; classifying and storing data by data type and establishing indexes can achieve rapid mapping from "key fields to target data," solving the problems of low retrieval efficiency and high reuse difficulty in traditional storage methods. This allows historical experience to quickly empower subsequent reasoning processes and provides a structured data foundation for dynamic system optimization.

[0082] In some embodiments of this application, the method for parsing embodied intelligence protocols and verifying behaviors in multi-agent collaboration further includes dynamically optimizing and adjusting the target collaboration model, specifically including: Real-time monitoring of the target collaboration model's operational metrics, including protocol recognition accuracy, parsing time, action plan compliance rate, and anomaly correction success rate; Set a baseline threshold range for the operating metrics. When any operating metric is detected to exceed the baseline threshold range and continues for a preset duration, the model will be dynamically optimized. If the accuracy of protocol recognition decreases due to the addition of new protocol types, features are extracted based on the communication data of the new protocol to generate a new protocol fingerprint and update it to the fingerprint database. At the same time, the comparison parameters of the protocol recognition agent are adjusted. If the anomaly rate of physical constraint deduction increases due to changes in environmental constraints, update the environmental obstacle model and robot kinematic parameters, and adjust the safety distance threshold in the physical environmental constraint deduction rules. If the parsing time is prolonged due to changes in the performance of the parsing tool, optimize the parsing tool evaluation metrics and selection logic of the protocol scheduling agent to ensure that scheduling efficiency matches parsing results.

[0083] It is understandable that the dynamic optimization and adjustment mechanism in this embodiment is an organically linked closed-loop system. When the model is dynamically optimized, the system will automatically match the corresponding optimization strategy according to different reasons for anomalies. The execution of these strategies is not a one-time static adjustment, but will continuously track the changes in the optimized operating indicators. For example, after supplementing and extracting new protocol features and updating the fingerprint database, the protocol recognition agent will use new comparison parameters to identify subsequent protocol data, and the system will monitor in real time whether its accuracy has rebounded to the baseline threshold range. Similarly, after updating the environmental obstacle model and adjusting the safety distance threshold, the physical environment constraint inference module will re-verify the compliance of the robot's action plan and observe whether the anomaly rate has decreased. Through this continuous monitoring-evaluation-optimization cycle, the target collaboration model can continuously adapt to new protocol types, environmental conditions, and tool performance in complex and ever-changing real-world application scenarios, thereby maintaining a long-term efficient and stable collaborative operation state, ensuring that the embodied intelligent behavior of the multi-agent system is highly consistent with the expected goal, and effectively improving the robustness and adaptability of the entire system.

[0084] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program goods. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0085] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program goods according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for use in the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0086] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0087] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for parsing embodied intelligence protocols and verifying behaviors in multi-agent collaboration, characterized in that, Applied to embodied intelligence systems, the method includes: Define intelligent agents corresponding to protocol identification, protocol scheduling, and behavior reasoning; the protocol identification intelligent agent is used to compare and determine protocol types, the protocol scheduling intelligent agent is used to select an appropriate protocol parsing tool and execute scheduling, and the behavior reasoning intelligent agent is used to associate external knowledge and generate action execution plans; and define a protocol fingerprint for protocol feature representation, which is constructed based on the fusion of statistical features and semantic features, and is used to describe the differences in structure and field patterns of different protocol types; define dual verification rules for behavior verification, which include semantic logic consistency verification rules and physical environment constraint inference rules, the semantic logic consistency verification rules are constructed based on the temporal and mutual exclusion relationships of action instructions, and the physical environment constraint inference rules are constructed based on the spatial relationship between the robot's motion state and environmental obstacles; Based on the aforementioned intelligent agents, protocol fingerprints, and dual verification rules, a multi-agent collaborative parsing and behavior verification model is constructed. By utilizing historical communication data, we complete the construction and storage of protocol fingerprints, as well as the adaptation and debugging of multi-agent collaborative parsing and behavior verification models, to obtain a target collaborative model that can adapt to multi-protocol environments and output compliant action execution schemes. When the embodied intelligent system is running online, it captures real-time communication messages between the system and external devices, extracts statistical and semantic features of the communication messages, and fuses them to form a comprehensive feature vector. The comprehensive feature vector is compared with a stored protocol fingerprint, and the protocol type is determined by a protocol identification agent and transmitted to a protocol scheduling agent. The protocol scheduling agent selects an appropriate protocol parsing tool to perform structured parsing of the communication messages, obtaining structured semantic data. The structured semantic data is input into a behavior reasoning agent, which correlates with an external knowledge base to generate a robot action execution plan. The action execution plan is verified and deduced using a dual verification rule, triggering regeneration when an anomaly is detected. Updated system communication data and action execution feedback information are collected, and protocol parsing and behavior verification are performed cyclically to achieve dynamically optimized protocol processing and behavior control.

2. The method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration according to claim 1, characterized in that, The defined intelligent agents corresponding to protocol identification, protocol scheduling, and behavior reasoning include: The protocol recognition agent takes feature comparison as its core function. Its input includes a comprehensive feature vector formed by fusing statistical feature vectors and semantic feature vectors of communication messages. By calculating the similarity with the stored protocol fingerprint, it outputs the protocol type label and corresponding confidence information. The protocol scheduling agent takes parsing tool matching as its core function. Its inputs include the protocol type label output by the protocol identification agent, the version adaptation information of each protocol parsing tool, the semantic parsing completeness record, and the current running load data. Through multi-dimensional evaluation and screening, it outputs the corresponding protocol parsing tool identifier and scheduling instructions. The behavioral reasoning agent takes action plan generation as its core function. Its inputs include structured semantic data, robot motion parameters from an external knowledge base, and knowledge fragments related to safety regulations and environmental constraints. Through associative reasoning and logical combination, it outputs robot action execution plans and corresponding parameter sets.

3. The method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration according to claim 2, characterized in that, The construction of the protocol fingerprint includes: Extract statistical features from communication messages, including field occurrence frequency, field length distribution, byte entropy value, overall message length pattern, and field position statistics. Extract semantic features from communication messages, including inter-byte relationships, field semantic block patterns, and instruction logical association information; The statistical features and semantic features are concatenated or weighted to form a feature vector, which is then standardized and mapped to a fixed-dimensional feature representation through feature embedding. Based on the aforementioned feature representation, the differences in structure and field patterns of different protocol types are identified, a unique corresponding protocol fingerprint is generated, and the fingerprint is stored in the fingerprint database according to the protocol type. When the similarity between the comprehensive feature vector corresponding to a new communication message and all protocol fingerprints in the fingerprint database is lower than a set threshold, it is marked as an unknown protocol. A new protocol fingerprint candidate is generated based on the feature vector of the message and added to the fingerprint database after verification.

4. The method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration according to claim 3, characterized in that, The construction of the semantic logic consistency verification rules includes: Obtain the temporal dependencies of various robot action commands, clarify the correspondence between actions executed first and actions executed later, and form a list of temporal constraints; Identify mutually exclusive action instruction combinations, clarify action instruction pairs that cannot be executed simultaneously or consecutively, and form a list of mutual exclusion constraints. A logical association matrix is ​​constructed based on the time constraint list and the mutual exclusion constraint list. The rows and columns of the logical association matrix correspond to various action instructions, and the matrix elements identify the compatibility status of the corresponding action instruction combinations. The instruction sequence in the action execution plan is compared with the logical association matrix. If there is an instruction combination that violates the timing constraint or mutual exclusion constraint, it is judged as a semantic logic anomaly.

5. The method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration according to claim 4, characterized in that, The construction of the physical environment constraint deduction rules includes: Obtain the robot's kinematic parameters, including joint range of motion, end effector trajectory limits, and motion speed limits; Establish an environmental obstacle model to clarify the spatial location, size, and distribution information of the obstacles; The motion parameters in the motion execution scheme are mapped to the robot end effector pose sequence, which includes the spatial coordinates and attitude angles of the end effector at each moment. Calculate the Euclidean distance between each coordinate point in the pose sequence and the surface of the environmental obstacle model, and preset a safe distance threshold. If the Euclidean distance corresponding to any coordinate point is less than the preset safe distance threshold, it is determined to be a physical constraint violation.

6. The method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration according to claim 5, characterized in that, The adaptation and debugging of the multi-agent collaborative parsing and behavior verification model using historical communication data includes: A simulation test environment is constructed based on historical communication data. The environment simulates the transmission of communication messages of multiple protocol types, the operating status of different protocol parsing tools, and the physical scenario of robot action execution. Historical communication data is categorized by protocol type to construct a test dataset, which includes communication messages in different protocol formats, corresponding standard parsing results, and compliance action execution plans. The test dataset is input into the multi-agent collaborative parsing and behavior verification model, driving the protocol recognition agent, protocol scheduling agent and behavior reasoning agent to run in sequence, and collecting the output results, parsing accuracy and action plan compliance rate data of each agent. Based on the collected data, the model parameters are adjusted, including the feature weights of the protocol fingerprint, the evaluation index thresholds of the agent, and the inference logic parameters, and the testing and adjustment are carried out iteratively. When the model's protocol recognition accuracy, parsing tool matching accuracy, and action plan compliance rate all reach the preset standards, and there are no significant fluctuations in the number of consecutive preset tests, the model adaptation and debugging are deemed complete, and the current parameters are solidified to form the target collaborative model.

7. The method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration according to claim 6, characterized in that, The step of extracting statistical and semantic features of the communication message and fusing them to form a comprehensive feature vector includes: The communication messages are preprocessed, including field splitting, invalid information removal and format standardization, to obtain standardized message data; Statistical features are extracted from standardized message data, and feature parameters such as field frequency and length distribution are obtained through statistical calculations to form a statistical feature vector. By analyzing the relationships between bytes, the semantic relationships between fields, and the instruction logic in the normalized message data, semantic feature parameters are extracted to form a semantic feature vector; The statistical feature vector and the semantic feature vector are aligned according to a preset dimension and fused using a weighted summation or concatenation method to generate a comprehensive feature vector with a fixed dimension. The weight coefficients of the weighted summation are set based on the contribution of different features to protocol recognition.

8. The method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration according to claim 7, characterized in that, The triggering of regeneration when an anomaly is detected includes: When the semantic logic consistency check determines that there is an anomaly, extract the instruction combination, timing conflict point and mutual exclusion constraint violation information associated with the anomaly to form a logical anomaly feature; When physical environment constraint deduction determines that there is an anomaly, the pose coordinates of the out-of-bounds position, the corresponding motion parameters, and the distance data to the obstacle are extracted to form physical anomaly features; The logical or physical anomaly features are used as constraint information and fed back to the behavioral reasoning agent, while also being transmitted to the protocol scheduling agent to confirm whether there are any deviations in the parsing process. The behavioral reasoning agent adjusts the reasoning boundary based on constraint information, re-retrieves relevant knowledge fragments in conjunction with the external knowledge base, and regenerates the action execution plan within the scope of excluding abnormal constraints. The double verification rule is applied again to the regenerated action execution plan until the plan passes the verification or the preset maximum number of regeneration attempts is reached.

9. The method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration according to claim 8, characterized in that, This also includes vectorizing and storing the validated action execution plans and exception amendment examples, specifically including: The verified action execution schemes are structured, and the core action parameters, corresponding protocol types, environmental constraints and execution result feedback information are extracted and converted into standardized data formats. The abnormal amendment cases are broken down to extract the original abnormal action plan, abnormal type, constraint information, correction strategy and final compliance plan, forming structured case data; The standardized data format action execution plan and case structured data are transformed into fixed-dimensional vector data by using feature embedding. The vector data contains the core features and related information of the plan. Vector data is categorized by data type and stored in a dedicated database, with indexes established. These indexes include key fields such as protocol type, action scenario, and exception type, which are used for rapid retrieval and retrieval during subsequent reasoning processes.

10. The method for parsing and verifying embodied intelligence protocols and behaviors in multi-agent collaboration according to claim 9, characterized in that, It also includes dynamically optimizing and adjusting the target collaboration model, specifically including: Real-time monitoring of the operational metrics of the target collaboration model, including protocol recognition accuracy, parsing time, action plan compliance rate, and anomaly correction success rate; Set a baseline threshold range for the operating metrics. When any operating metric is detected to exceed the baseline threshold range and continues for a preset duration, the model will be dynamically optimized. If the accuracy of protocol recognition decreases due to the addition of new protocol types, features are extracted based on the communication data of the new protocol to generate a new protocol fingerprint and update it to the fingerprint database. At the same time, the comparison parameters of the protocol recognition agent are adjusted. If the anomaly rate of physical constraint deduction increases due to changes in environmental constraints, update the environmental obstacle model and robot kinematic parameters, and adjust the safety distance threshold in the physical environmental constraint deduction rules. If the parsing time is prolonged due to changes in the performance of the parsing tool, optimize the parsing tool evaluation metrics and selection logic of the protocol scheduling agent to ensure that scheduling efficiency matches parsing results.