Multi-robot control method and device based on multi-modal causal reasoning, electronic equipment and computer program product

By acquiring and calculating the causal influence strength between the operating status data of a multi-robot system and the global objective, and by screening and planning local action sequences, the problem of insufficient causal quantitative correlation in multi-robot systems is solved, and the accurate adaptation of local actions to the global objective is achieved.

CN122008232APending Publication Date: 2026-05-12广东省工业边缘智能创新中心有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
广东省工业边缘智能创新中心有限公司
Filing Date
2026-03-24
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, the causal quantification data source is of a single modality, and there is a lack of causal quantification correlation between the global goal of multiple robots and the local actions of a single robot, making it difficult for the local actions of a single robot to accurately adapt to the global goal.

Method used

By acquiring operational status data and global target quantification data from multiple business dimensions of various robots, the causal influence intensity is calculated, target operational status data that meet the set causal influence intensity are selected, and action planning is performed based on this data to generate action execution instructions for the robots.

Benefits of technology

It achieves precise matching between local actions and global goals in multi-robot systems, improves the pertinence and adaptability of multi-robot collaborative control, and ensures that local action sequences match global goals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122008232A_ABST
    Figure CN122008232A_ABST
Patent Text Reader

Abstract

The invention is suitable for the field of robot control, and provides a multi-mode causal reasoning-based multi-robot control method and device, electronic equipment and a computer program product, and the method comprises the steps: obtaining the operation state data of a plurality of business dimensions of a plurality of robots and the global target quantification data corresponding to a global target; calculating the causal influence intensity between the operation state data of each business dimension of each robot and the global target quantitative data; screening out target operation state data of which the causal influence intensity meets set causal influence intensity from the operation state data of the multiple business dimensions of each robot; performing action planning based on the target operation state data and the global target quantitative data of the robots to obtain a local action sequence of each robot; and based on the local action sequence, generating an action execution instruction of the robot, and sending the action execution instruction to the robot. According to the scheme, the adaptability of the local action and the global target of the single robot can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of robot control, and in particular relates to a multi-robot control method, device, electronic device and computer program product based on multimodal causal reasoning. Background Technology

[0002] With the development of intelligent manufacturing, multiple robots in the industrial field often collaborate to efficiently handle complex tasks, such as material handling, parts assembly, and order fulfillment. To achieve orderly cooperation among multiple robots, existing technologies have introduced control methods such as hierarchical architecture scheduling and preset rule-driven approaches. However, these methods generally suffer from the following problems: the causal quantification data source modality is singular, and there is a lack of causal quantification correlation between the global goal of multiple robots and the local actions of a single robot, making it difficult for the local actions of a single robot to accurately adapt to the global goal. Summary of the Invention

[0003] This application provides a multi-robot control method, device, electronic device, and computer program product based on multimodal causal reasoning to solve the problem that in the prior art, the causal quantification data source is single-modal, and there is a lack of causal quantification correlation between the global goal of multiple robots and the local actions of a single robot, which makes it difficult for the local actions of a single robot to accurately adapt to the global goal.

[0004] The first aspect of this application provides a multi-robot control method based on multimodal causal reasoning, including: The system acquires operational status data for multiple robots across multiple business dimensions, as well as global target quantification data corresponding to the global target; the operational status data is the planning impact data of local actions. Calculate the causal influence strength between the operational status data of each robot in each business dimension and the global target quantification data; the causal influence strength characterizes the degree of influence of the operational status data on the global target; From the operational status data of each robot across multiple business dimensions, target operational status data whose causal influence strength meets the set causal influence strength is selected. Motion planning is performed based on the target operating state data of multiple robots and the global target quantization data to obtain local motion sequences for each robot; each local motion sequence includes at least one local motion that fits the global target. Based on the local action sequence of each robot, an action execution instruction for the robot is generated and sent to the robot.

[0005] A second aspect of this application provides a multi-robot control device based on multimodal causal reasoning, comprising: The acquisition module is used to acquire the operational status data of multiple robots in multiple business dimensions and the global target quantization data corresponding to the global target; the operational status data is the planning influence data of local actions; The calculation module is used to calculate the causal influence strength between the operating status data of each business dimension of each robot and the global target quantification data; the causal influence strength characterizes the degree of influence of the operating status data on the global target; The filtering module is used to filter out target operating state data whose causal influence strength meets the set causal influence strength from the operating state data of multiple business dimensions of each robot. The planning module is used to perform motion planning based on the target running state data of multiple robots and the global target quantization data to obtain a local motion sequence for each robot; the local motion sequence includes at least one local motion that fits the global target; A generation and sending module is used to generate motion execution instructions for each robot based on the local motion sequence of each robot, and send the motion execution instructions to the robot.

[0006] A third aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described in the first aspect.

[0007] A fourth aspect of this application provides a computer program product comprising a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0008] A fifth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0009] As can be seen from the above, this application calculates the causal influence strength between the operational status data of each robot in each business dimension and the global target quantification data. It establishes a direct and quantifiable causal relationship between operational status data and the global target using multimodal data. Based on this causal quantification relationship, it filters out target operational status data with a high degree of influence on the global target from the operational status data. The operational status data represents the planning influence data for local actions, while the filtered target operational status data is data related to local actions and has an impact on the global target. Action planning is performed using this target operational status data and the global target quantification data to obtain the local action sequence of each robot. Correspondingly, action execution instructions are generated and issued to each robot, achieving control of multiple robots. The local actions in the local action sequence are determined based on the target operational status data that has a causal relationship with the global target. Relying on the causal quantification relationship between the target operational status data and the global target, the local actions also form a causal quantification relationship with the global target. The generated local actions accurately fit the global target, thus effectively solving the problem that the lack of a causal quantification relationship between the global target of multiple robots and the local actions of a single robot makes it difficult for the local actions of a single robot to accurately adapt to the global target. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart of a multi-robot control method based on multimodal causal reasoning provided in an embodiment of this application; Figure 2 This is a schematic diagram of a target cause-effect graph provided in an embodiment of this application; Figure 3 This is an architecture diagram of a multi-robot control system based on multimodal causal reasoning provided in an embodiment of this application; Figure 4 This is a structural diagram of a multi-robot control device based on multimodal causal reasoning provided in an embodiment of this application; Figure 5 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0013] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0014] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0015] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0016] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [the described condition or event] is detected," or "in response to detection of [the described condition or event]."

[0017] In specific implementations, the terminals described in the embodiments of this application include, but are not limited to, other portable devices such as mobile phones, laptop computers, or tablet computers with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0018] The following discussion describes terminals that include displays and touch-sensitive surfaces. However, it should be understood that terminals may include one or more other physical user interface devices such as physical keyboards, mice, and / or joysticks.

[0019] The terminal supports a variety of applications, such as one or more of the following: drawing applications, presentation applications, word processing applications, website creation applications, disc burning applications, spreadsheet applications, game applications, telephone applications, video conferencing applications, email applications, instant messaging applications, exercise support applications, photo management applications, digital camera applications, digital camcorder applications, web browsing applications, digital music player applications, and / or digital video player applications.

[0020] Various applications that can run on the terminal can use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the terminal can be adjusted and / or changed between and / or within applications. In this way, the terminal's common physical architecture (e.g., the touch-sensitive surface) can support various applications with user interfaces that are intuitive and transparent to the user.

[0021] It should be understood that the sequence number of each step in this embodiment does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application embodiment.

[0022] To illustrate the technical solution described in this application, specific embodiments are provided below.

[0023] See Figure 1 , Figure 1 This is a flowchart illustrating a multi-robot control method based on multimodal causal reasoning, provided in an embodiment of this application. Figure 1 As shown, a multi-robot control method based on multimodal causal reasoning is proposed, which includes the following steps: Step 101: Obtain the operational status data of multiple business dimensions for each of the multiple robots and the global target quantification data corresponding to the global target; the operational status data is the planning impact data of local actions.

[0024] In this application, the business dimension is a business classification dimension used to divide robot operating state data. It is also called the state attribute dimension or functional dimension. It is an independent attribute category of robot operating state and is completely different from the feature dimension (data dimension) of the feature vector. Each business dimension corresponds to a set of parameters that describe a certain category of robot operating state.

[0025] In some embodiments, common business dimensions of robot operating status include, but are not limited to: position dimension, such as robot current coordinates, target point distance, path offset, etc.; motion dimension, such as speed, acceleration, angular velocity, motion trajectory smoothness, etc.; load dimension, such as load weight, load balance status, grasping posture, etc.; energy consumption dimension, such as instantaneous power consumption, remaining power, energy efficiency, etc.; and sensing dimension, such as sensor detection accuracy, obstacle perception distance, signal strength, etc.

[0026] Each business dimension is an independent business attribute, and the causal influence of its corresponding operational status data on the global objective can be calculated.

[0027] Operational status data refers to the data generated by the robot during operation that can affect local motion planning.

[0028] Operational status data across multiple business dimensions includes robot body data (such as pose data, joint torque, end effector pressure, battery power, etc.), environmental interaction data (such as RGB images / video streams, LiDAR 3D point cloud data, etc.), and task context data (such as order task information, urgency level labels, deadlines, etc.).

[0029] Operational status data can influence the planning of a robot's local actions (such as grasping, obstacle avoidance, path selection, etc.), providing a basis for decision-making in local action planning.

[0030] Global target quantification data refers to the set of target thresholds after quantifying the key performance indicators (KPIs) presented by multiple robots when performing tasks. It is a specific manifestation of the global target and can be set manually or automatically sensed and updated by the system. Examples include order timeliness ≤30min, energy consumption ≤500W, and warehouse turnover rate ≥90%.

[0031] In some embodiments, data acquisition and processing are accomplished through multi-source sensors and data interaction interfaces.

[0032] In some embodiments, robot body data such as pose data, joint torque, end effector pressure, battery power, and speed are collected in real time by sensors such as encoders, six-dimensional force sensors, ultra-wideband positioning modules, and inertial measurement units (IMUs) mounted on the robot. The robot body data is usually text modal data.

[0033] In some embodiments, RGB images / video streams are acquired by a vision sensor (camera) mounted on the robot, and 3D point cloud data is acquired by a LiDAR sensor. Real-time environmental perception data such as dynamic obstacles (e.g., manual forklifts, temporary material stacks) and channel occupancy (e.g., forklift density) are captured. The environmental perception data is usually image modal data and point cloud modal data.

[0034] In some embodiments, a communication connection is established with the upper-level Warehouse Management System (WMS) or task scheduling platform to receive task list, urgency level tags for each task (such as urgent task, normal task), task deadline, and other task context data. The task context data is usually text modal data.

[0035] In some embodiments, based on application scenario requirements, technicians can pre-set global target quantification data, or automatically sense and update global target quantification data by real-time monitoring of global operating status (such as overall energy consumption and order completion progress) to ensure that it matches the actual operating scenario.

[0036] By acquiring operational status data and global target quantification data for multiple robots across various business dimensions, this provides complete and accurate basic data support for the quantitative analysis of the causal relationship between the robot's local actions and global targets. This avoids biases in causal analysis and subsequent decision-making errors caused by incomplete or missing data, ensuring the relevance and adaptability of multi-robot collaborative control.

[0037] Step 102: Calculate the causal influence strength between the operating status data of each business dimension of each robot and the global target quantification data; the causal influence strength characterizes the degree of influence of the operating status data on the global target.

[0038] In this application, the causal influence strength refers to the degree of direct causal influence of operational status data of a certain business dimension on the achievement of the global goal. Its value ranges from [-1, +1], encompassing both positive and negative influences of operational status data on the achievement of the global goal. The causal influence strength includes signs ("-" and "+", with "+" optional), representing the direction of influence: "-" indicates a negative influence, and "+" indicates a positive influence. The absolute value of the causal influence strength represents the magnitude of the influence. The larger the absolute value, the stronger the causal relationship between the operational status data and the global goal, and the more significant the impact on the achievement of the global goal. Therefore, this operational status data is more likely to be included in local action planning as action planning data. 0 indicates no impact on the global goal.

[0039] In some embodiments, the operational status data of each robot across multiple business dimensions are preprocessed. If the data is missing or abnormal, interpolation or outlier removal algorithms are used to correct it, ensuring data validity.

[0040] In some embodiments, a causal influence operator is used to quantitatively calculate the causal influence strength between the operational status data of each business dimension and the global target quantification data.

[0041] In some embodiments, the mathematical expression for the causal influence operator is: .in, This refers to running status data or variants of running status data (such as fused status data as described below). Quantify data for the global objective. To quantify the causal influence between operational status data and global target data, For the gradient of causal influence, Among all the operational state data and variant data of multiple robots, the data that maximizes the absolute value of the causal influence gradient is selected. This is the absolute value of the global maximum causal influence gradient. If the value is positive, it means that the corresponding operating status data has a positive impact on the achievement of the global goal; if the value is negative, it means that the corresponding operating status data has a negative impact on the achievement of the global goal.

[0042] By employing causal impact operators to calculate the causal impact strength between operational state data and global target quantification data, partial derivatives can accurately characterize the actual impact of changes in operational state data or its variants on global target quantification data. This transforms the abstract causal relationship between local states and global targets into quantifiable and comparable values, enabling an objective measurement of the contribution of operational state data to the global target. Simultaneously, the absolute value of the maximum causal impact gradient corresponding to the operational state data and its variants of all robots is used as a global normalization benchmark. This normalizes the causal impact values ​​of different robots and different business dimensions to the same comparable range, eliminating differences in dimensions and numerical magnitudes among various data types. This ensures the horizontal comparability of the causal impact strength of local state data, providing an accurate and reliable basis for subsequent selection of high-contribution target operational state data based on unified quantification standards. Furthermore, when global target quantification data is updated, a standardized causal impact strength adapted to the current target can be quickly recalculated, ensuring the accuracy and consistency of local-global causal relationship analysis in multi-robot collaborative scenarios.

[0043] In some embodiments, a backdoor adjustment mechanism may be introduced to block confusion factors (such as visual misjudgment caused by illumination interference, the influence of ambient temperature on sensor accuracy, etc.) to ensure the accuracy of causal influence strength calculation.

[0044] In some embodiments, the mathematical expression for the backdoor adjustment mechanism is: .in, This indicates the causal gradient after backdoor adjustment to eliminate interference from confounding factors; that is, the operational status data or its variants after the intervention operation. Quantify global target data The true, unbiased causal gradient is used to replace the original causal influence gradient in the aforementioned causal influence operator calculation formula. Similarly, the absolute value of the global maximum causal influence gradient in the denominator of the aforementioned causal influence operator formula can also be optimized through the aforementioned backdoor adjustment mechanism; the unbiased causal influence strength is then calculated. ; Represents the set of confounding factors, referring to those that simultaneously affect and Furthermore, a set of interfering factors that may create a false causal relationship between the two, including but not limited to ambient light, ambient temperature, sensor noise, and communication delay; Indicates a given set of confusion factors Under the conditions, right The conditional causal gradient; This represents the marginal probability distribution of the confusion factor set Z, which is used to perform weighted integration on the conditional causal gradients corresponding to confusion factors with different values, thereby achieving normalization adjustment of the confusion factors.

[0045] By performing weighted integration on the probability distribution of confusion factors, adaptive smoothing of various interference factors is achieved, thereby improving the robustness and stability of the causal influence intensity calculation results in complex and ever-changing multi-robot operation scenarios.

[0046] Introducing a backdoor adjustment mechanism to correct the causal influence gradient can effectively identify and block backdoor paths between confounding factors and operational status data and global target quantification data. It eliminates computational biases caused by non-causal confounding factors such as environmental interference and sensor errors, and corrects the correlation gradient at the original observation level to an intervention gradient at the pure causal level. This ensures that the calculated causal influence gradient only reflects the true direct causal effect of operational status data on the global target, further guaranteeing the reliability and accuracy of subsequent screening of target operational status data and execution of action planning based on the intensity of causal influence. This ensures that multi-robot collaborative control decisions are not misled by interfering factors and always conform to the true causal relationship logic.

[0047] By quantifying the causal relationship between local operational status data and global objectives, key data that have a substantial impact on global objectives can be accurately identified, providing a scientific basis for subsequent data screening. At the same time, the reliability of causal calculation is improved through a backdoor adjustment mechanism.

[0048] In some embodiments, the operational status data for each business dimension includes operational status sub-data of at least one modality. Calculating the causal influence strength between the operational status data for each business dimension of each robot and the global target quantization data includes: when the operational status data includes operational status sub-data of one modality, calculating the causal influence strength between the operational status sub-data and the global target quantization data, and using this as the causal influence strength between the operational status data and the global target quantization data; or, when the operational status data includes operational status sub-data of multiple modalities, performing multimodal feature fusion processing on the multiple operational status sub-data to obtain fused status data; calculating the causal influence strength between the fused status data and the global target quantization data, and using this as the causal influence strength between the operational status data and the global target quantization data.

[0049] Operational status sub-data refers to the subdivided data based on modality within the operational status data of the same business dimension. For example, operational status data in the location positioning dimension may include operational status sub-data in the text modality (global coordinate data collected by the positioning system) and operational status sub-data in the point cloud modality (relative position offset data collected by the LiDAR sensor); operational status data in the motion state dimension may include operational status sub-data in the image modality (motion velocity data collected by the visual sensor) and operational status sub-data in the text modality (acceleration and angular velocity data collected by the IMU).

[0050] Multimodal feature fusion processing refers to the process of semantically aligning, weighting, and integrating the operational state sub-data of different modalities (text modality, image modality, point cloud modality) to obtain fused state data with unified feature dimensions and complete semantic information.

[0051] Fusion status data refers to the comprehensive data obtained by fusing multimodal operational status sub-data of a robot in the same business dimension. It can eliminate intermodal bias and comprehensively reflect the overall operational status of that business dimension.

[0052] In this application, different calculation processes for the causal influence intensity exist for different modalities of the operational status data.

[0053] Scenario 1: Single-modal operational state sub-data. If the operational state data of a certain business dimension of the robot only contains operational state sub-data of one modality (such as the battery state dimension only containing text modality power data), then this operational state sub-data is directly used as the calculation object. The causal influence strength between it and the global target quantized data is calculated through causal influence operators and other methods. This strength is the causal influence strength between the operational state data of the corresponding business dimension and the global target quantized data.

[0054] Scenario 2: Multimodal operational state sub-data. If the operational state data of a robot in a certain business dimension contains multiple modalities (e.g., the task execution environment dimension contains environmental images in the image modality and environmental point clouds in the point cloud modality), then the multimodal feature fusion processing of the operational state sub-data of multiple modalities under that business dimension needs to be performed first to obtain fused state data. Then, the causal influence strength between the fused state data and the global target quantization data is calculated through causal influence operators and other methods. This strength is the causal influence strength between the operational state data of the corresponding business dimension and the global target quantization data.

[0055] In this application, a simplified calculation process is developed for single-modal data to improve computational efficiency; to address the heterogeneity problem of multimodal data, multimodal feature fusion processing is used to eliminate intermodal biases, achieve effective fusion of multimodal data, and ensure the comprehensiveness and accuracy of causal influence intensity calculation.

[0056] In some embodiments, the step of performing multimodal feature fusion processing on multiple operational state sub-data to obtain fused state data includes: performing cross-modal alignment processing on multiple operational state sub-data to obtain multiple cross-modal aligned data; calculating the modal influence weight corresponding to each cross-modal aligned data based on the global target quantization data through a cross-modal attention mechanism; and performing weighted calculation based on the multiple cross-modal aligned data and the modal influence weight corresponding to each cross-modal aligned data to obtain the fused state data.

[0057] Cross-modal alignment refers to the process of eliminating semantic differences between different modalities of data through processing methods such as query-key-value (QKV) attention mechanisms, so that the data of each modality is in the same semantic space and has comparability and fusion.

[0058] Cross-modal attention mechanisms refer to calculating the correlation between cross-modal aligned data corresponding to different modalities and the global objective, assigning different modal influence weights, and highlighting the contribution of key modalities. An example of a cross-modal attention mechanism is the KPI-aware attention mechanism.

[0059] Modal influence weight refers to the quantitative value of the degree of influence of the alignment data of a certain modality on the achievement of the global goal. The value range is [0, 1]. The sum of the modal influence weights under a certain business dimension of a certain robot is 1, which is used to adjust the contribution ratio of the running state sub-data in the fusion process.

[0060] In some embodiments, feature extraction is performed on the runtime state sub-data of each modality. For the runtime state sub-data of the text modality, feature extraction is performed using a Transformer-based Bidirectional Encoder Representations from Transformers (BERT) model to obtain text feature vectors. For the runtime state sub-data of the image modality, feature extraction is performed using models such as Residual Network (ResNet) to obtain image feature vectors. For the operational state sub-data of point cloud modes, feature extraction is performed using LiDAR encoding models to obtain point cloud feature vectors. ; In some embodiments, the feature vectors of the extracted multimodal features are concatenated pairwise to obtain composite features. This achieves initial alignment of feature dimensions between different modalities. Among them, , The feature dimension of the first modality feature vector participating in the splicing. This refers to the feature dimension of the second modality feature vector involved in the concatenation. For example, the text feature vector... With point cloud feature vectors By concatenating along the feature dimensions, composite features are obtained. .

[0061] In some embodiments, for the operational state sub-data of the two modalities, the spliced ​​composite features are linearly mapped and normalized to obtain normalized composite aligned features. The features are then split according to the proportion of the original modal feature dimensions to obtain two cross-modal aligned data that correspond one-to-one with the two modalities. The two cross-modal aligned data have the same feature dimension and no modal scale / distribution bias.

[0062] In some embodiments, cross-modal alignment of runtime state sub-data for the three modalities can be achieved through the QKV attention mechanism, specifically including the following operations: After feature concatenation, three composite features are obtained. For each composite feature, a learnable projection matrix is ​​used. and learnable projection matrix Mapped to the Key space and Value space respectively, we get and , , , , , , , For composite features projected by the Key matrix The composite key vector obtained after mapping to the key feature space serves as the key baseline feature for cross-modal alignment. For composite features projected by the Value matrix The composite Value vector obtained after mapping to the Value feature space serves as the baseline Value feature for cross-modal alignment. For each composite feature, the Key obtained by projection mapping ( ) and Value ( Using this as a unified benchmark, the feature vector of the third modality, which was not involved in the concatenation, is used as Query(Q) and input into the QKV attention formula. We obtain the cross-modal alignment features corresponding to the running state sub-data of the third modality that did not participate in feature concatenation, and finally obtain the cross-modal alignment data corresponding to the running state sub-data of each of the three modalities.

[0063] Semantic alignment of multimodal data is achieved through cross-modal alignment processing, eliminating intermodal biases.

[0064] In some embodiments, the global target quantization data is feature-encoded to obtain the KPI feature vector. Based on cross-modal perceptual attention mechanisms, such as the KPI-perceptual attention formula Calculate the modal influence weights for each cross-modal aligned data point. To align data across modalities, The modal influence weights.

[0065] The global target quantification data is encoded into KPI feature vectors and integrated into a cross-modal perception attention mechanism, so that the calculation of modal influence weights is globally target oriented. It can adaptively learn the actual impact of each cross-modal aligned data on the achievement of the global target, accurately allocate different modal influence weights, provide a weight basis that fits the global target requirements for subsequent weighted fusion, and improve the matching degree between fused state data and global target.

[0066] In some embodiments, each cross-modal alignment data is multiplied by its corresponding modal influence weight, and then all results are summed to obtain the fused state data. ,in, The number of cross-modal aligned data. The fused state data integrates the core information of operational state sub-data from multiple modalities, and highlights the influence of modalities strongly correlated with the global objective. It effectively eliminates the heterogeneous bias between multiple modalities, making the fused state data accurately match the requirements for achieving the global objective, and providing a reliable feature basis for subsequent calculation of causal influence intensity.

[0067] Multimodal alignment processing generates cross-modal aligned data, ensuring semantic consistency. The calculation of modal influence weights realizes the quantification of influence. Weighted fusion integrates the core value of each modal operational state sub-data. The final fused state data can comprehensively and accurately reflect the comprehensive information of multimodal operational state sub-data, providing high-quality input for subsequent causal influence intensity calculation and further improving the accuracy of causal association analysis.

[0068] Step 103: From the operational status data of each robot across multiple business dimensions, select target operational status data whose causal influence strength meets the set causal influence strength.

[0069] Setting the causal influence strength refers to a pre-defined threshold used to filter target operational status data. The setting of the causal influence strength is determined based on factors such as the real-time requirements of the application scenario and the availability of computing resources.

[0070] Target operational status data refers to operational status data that has been filtered and has a significant causal relationship with the global target. It is the core input data for subsequent action planning.

[0071] Satisfying the set causal influence strength means either a first set causal influence strength that is greater than or equal to a positive value, or a second set causal influence strength that is less than or equal to a negative value.

[0072] In some embodiments, if the causal influence strength is set to 0.8 and -0.8, the operating status data with causal influence strength ≥ 0.8 and causal influence strength ≤ -0.8 are retained as target operating status data.

[0073] In some embodiments, data filtering can be achieved using a Markov blanket filter.

[0074] In some embodiments, the operational status data used for screening can be operational status data used to calculate the intensity of causal influence, or operational status data acquired in real time during screening. In this application, existing data used for calculating the intensity of causal influence can be reused to improve screening efficiency; real-time data can also be called to adapt to dynamic changes in the scenario. Flexible selection of screening data sources provides flexible data support for subsequent accurate screening of highly correlated target operational status data.

[0075] In some embodiments, data filtering is performed based on a causal graph model. Accordingly, before filtering out the target operating state data whose causal influence strength meets the set causal influence strength from the operating state data of multiple business dimensions of each robot, the method further includes: taking the operating state data of each business dimension of each robot as local data nodes, taking the global target quantization data as global target nodes, and constructing causal edges between nodes based on the causal influence strength between the operating state data and the global target quantization data, thereby obtaining a target causal graph containing the local data nodes, the global target nodes, and the causal edges.

[0076] Local data nodes refer to nodes in a cause-effect graph that represent the operational status data of a certain robot in a certain business dimension. They are the cause nodes of causal relationships, such as robot A - coordinate node, robot B - robotic arm load node.

[0077] A global target node is a node in a causal graph that represents the quantified data of the global target. It is the result node of a causal relationship, such as an energy consumption node.

[0078] A causal edge is an edge that connects a local data node to a global target node. Its weight is the corresponding causal influence strength, used to characterize the degree of causal relationship between cause and effect, and its value ranges from -1 to 1.

[0079] A target causal graph is a directed acyclic graph (DAG) containing multiple local data nodes, a global target node, and causal edges. It can intuitively represent the causal quantitative relationship between runtime state data and the global target. The target causal graph is stored in a dynamic causal graph database and supports dynamic updates.

[0080] In some embodiments, the operational status data of each business dimension of each robot is defined as a local data node, and the global target is defined as a global target node.

[0081] For operational status data containing only a single modality of operational status sub-data, this operational status sub-data is used as the node feature data of the corresponding local data node; for operational status data containing multiple modalities of operational status sub-data, the fused status data obtained by multimodal feature fusion processing is used as the node feature data of the corresponding local data node. This ensures that the node feature data can comprehensively reflect the operational status of the corresponding business dimension.

[0082] The node feature data of the global target node is the global target quantization data.

[0083] A causal edge is established between each local data node and the global target node. The weight of the causal edge is the causal influence strength corresponding to that local data node.

[0084] like Figure 2 As shown, Figure 2 This is a schematic diagram of a target causal graph provided in an embodiment of this application. The edge weight of the causal edge between the robot B-arm load node and the order timeliness node is 0.85, representing that the causal influence strength of the robot arm load data on order timeliness is 0.85; the edge weight of the causal edge between the robot B-battery power node and the order timeliness node is 0.65, representing that the causal influence strength of the battery power data on order timeliness is 0.65.

[0085] In some embodiments, the constructed target causal graph is stored in a dynamic causal graph database, which stores the node set (local data nodes, global target nodes) and edge set in real time. Simultaneously, causal discovery algorithms, such as Parallel Causal Discovery with Conditional Mutual Information Constraints Plus (PCMCI+), can be used to periodically analyze the operational status data (e.g., every 24 hours) to re-identify causal relationships between data and dynamically update causal edges between nodes to adapt to changing scenarios (e.g., the environmental differences between day and night shifts in a warehousing and logistics center).

[0086] By constructing a target causal graph, the causal relationship between local operational status data and global goals is transformed into a structured and visualized node-edge model, making the intensity of causal influence concrete. This provides rigorous and practical model support for subsequent quantitative data screening based on the intensity of causal influence. The intuitive graph structure can clearly present the relationship between local data and global goals, ensuring the causal relationship and targeting of the screening logic.

[0087] In some embodiments, after the target causal graph is constructed, the step of filtering target operating state data whose causal influence strength satisfies a set causal influence strength from the operating state data of multiple business dimensions of each robot includes: filtering target causal edges whose causal influence strength satisfies the set causal influence strength in the target causal graph; and determining the operating state data corresponding to the target local data node associated with the target causal edge as the target operating state data.

[0088] In the target causal graph, all causal edges are traversed, and causal edges whose edge weights (causal influence strength) satisfy the set causal influence strength are selected as target causal edges. The local data nodes associated with the target causal edges (target local data nodes) are then found, and the operating state data corresponding to these nodes is the target operating state data. For example, the edge weight of the causal edge between the robot B-battery power node and the energy consumption node is 0.75, which is less than the set causal influence weight of 0.8, so the corresponding battery power data is removed.

[0089] By filtering target data based on causal influence strength thresholds and using a Markov blanket filter to remove redundant data with weak causal relationships, the computational complexity of subsequent action planning is significantly reduced, while ensuring that the retained data makes a crucial contribution to the overall objective, thus supporting accurate action planning. Furthermore, the dynamic updating of the causal graph further enhances the timeliness and accuracy of causal relationship analysis.

[0090] Step 104: Based on the target running state data of the multiple robots and the global target quantization data, motion planning is performed to obtain the local motion sequence of each robot; the local motion sequence includes at least one local motion that fits the global target.

[0091] A local action sequence refers to a set of actions planned for a single robot and arranged in chronological order, such as obstacle avoidance → grasping → transporting → unloading. Each action corresponds to specific execution parameters (such as path coordinates, obstacle avoidance speed, grasping force, execution time, etc.).

[0092] Aligning with the overall goal means that the execution of local actions can contribute to the achievement of the overall goal.

[0093] By combining the target operating state data of multiple robots (local states with high causal correlation) with the requirements for achieving the global goal, a matching rule between local actions and the global goal is established. For each robot's own state characteristics, an ordered combination of actions containing at least one local action that matches the global goal is planned to form a local action sequence exclusive to each robot.

[0094] By planning robot local actions based on target operating state data that has a high causal correlation with the global objective, the robot's local actions can be precisely aligned with the global objective requirements, avoiding deviation between local actions and the global objective. At the same time, the robot's local actions can be adapted to its own operating state, improving the efficiency of multi-robot collaborative execution in achieving the global objective, and ensuring the targeted execution of local actions and the consistency of global collaboration.

[0095] In some embodiments, the step of performing motion planning based on the target operating state data of multiple robots and the global target quantization data to obtain local motion sequences for each robot includes: vectorizing and combining the target operating state data of each robot to obtain an initial feature vector of the robot; performing dimensionality reduction processing on each of the initial feature vectors to obtain low-dimensional feature vectors with a set feature dimension; inputting the low-dimensional feature vectors of multiple robots, the global target quantization data, and task context data into a target motion planning engine to obtain at least one candidate motion planning scheme output by the target motion planning engine; each candidate motion planning scheme includes candidate motion sequences corresponding to multiple robots; and determining, from the at least one candidate motion planning scheme, a target motion planning scheme whose global target achievement quantization data is closest to the global target quantization data, wherein the target motion planning scheme includes the local motion sequences corresponding to each robot.

[0096] The initial feature vector is a high-dimensional vector formed by arranging and combining all target operating state data of a single robot in a preset order, which can comprehensively characterize the key operating states of the robot.

[0097] Low-dimensional feature vectors refer to feature vectors with defined feature dimensions that are obtained after dimensionality reduction processing models such as Neural Information Compressor (NIS) and retain the core causal features. For example, setting the feature dimension to 8 dimensions reduces the complexity of planning.

[0098] The target motion planning engine refers to the motion planning engine used to generate candidate motion planning solutions. The motion planning engine includes a first motion planning engine and a second motion planning engine.

[0099] Candidate action planning schemes refer to the set of schemes generated by the target action planning engine, which contains candidate action sequences for each robot. Each scheme corresponds to a possible combination of cooperative actions. For example, robot A: path A → obstacle avoidance action 1 → grasping action 2, robot B: path B → obstacle avoidance action 3 → transporting action 4.

[0100] Global goal achievement quantification data refers to the global goal quantification results predicted based on candidate action planning schemes, such as predicting order delivery time of 28 minutes or predicting energy consumption of 480W, which are used to evaluate the merits of candidate action planning schemes.

[0101] For each robot, the selected target operating state data (such as pose data, joint torque data, path selection data, etc.) are converted into vector form and then spliced ​​and combined according to the preset feature dimension order (such as pose, path, load) to form an initial feature vector.

[0102] In some embodiments, a neural information compressor is used to reduce the dimensionality of the initial feature vector. For example, the initial feature vector is mapped to an 8-dimensional low-dimensional space through dimensionality reduction.

[0103] In some embodiments, data filtering and dimensionality reduction for multiple robots can be achieved simultaneously based on a target causal graph. ,in, for Real-time operational status data for all robots across multiple business dimensions. , The elements in , … This provides operational status data for a single robot across multiple business dimensions; the Real-valued Non-Volume Preserving (RealNVP) model is used for high-dimensional data. Dimensionality reduction is performed to achieve low-dimensional mapping of features while preserving information; It can be used as a data filter such as a Markov blanket filter; For the target cause-effect graph, Provide causal relationship constraints for data filtering. This is a feature dimension mapping / projection function used to map the filtered features to a preset dimension. Low-dimensional space enables the fixation and standardization of feature dimensions; This is a set of low-dimensional feature vectors that are strongly correlated with the global target after being filtered and dimensionality reduced, corresponding to multiple robots.

[0104] Constrained by the causal structure of the target causal graph, the high-dimensional operational state data of all robots are first compressed and reduced in dimensionality using a compressor. Then, the reduced features and the target causal graph are input into a data filter. Based on causal relationships, effective data strongly correlated with the global target are selected. Finally, the data is mapped to a set low-dimensional space through a feature dimension mapping function. This completes the causal-guided filtering and high-dimensional data reduction of multi-robot data in one go, providing high-quality and lightweight data support for subsequent multi-robot motion planning.

[0105] The low-dimensional feature vectors of all robots, global target quantization data, and task context data are input into the target action planning engine. The engine generates candidate action planning schemes based on preset planning logic. Each candidate action planning scheme contains a sequence of candidate actions for all robots, and the action types cover core actions such as grasping and obstacle avoidance from the meta-skill library.

[0106] In some embodiments, when there are multiple candidate action planning schemes, the global target achievement quantification data corresponding to each candidate action planning scheme is calculated based on the action utility evaluation function, and then the candidate action planning scheme whose data is closest to the global target quantification data is selected as the target action planning scheme, and the robot candidate action sequences contained therein are the final local action sequences.

[0107] In some embodiments, if there is only one candidate motion planning scheme, that candidate motion planning scheme is the target motion planning scheme.

[0108] By vectorization and dimensionality reduction, high-dimensional key data is transformed into low-dimensional effective features, significantly reducing the planning complexity of action planning and keeping inference latency within the threshold of industrial real-time control, thus solving the problem of excessive planning latency in existing technologies.

[0109] The target operation status data of multiple robots, global target quantification data, and task context data are used as input data for the target action planning engine to ensure the coordination and global adaptability of action planning. The candidate solution screening mechanism improves the optimization of local action sequences, ensuring that they fit the global goal and solving the static defects of the rule base.

[0110] In some embodiments, before inputting the low-dimensional feature vectors of the plurality of robots, the global target quantization data, and the task context data into the target motion planning engine to obtain at least one candidate motion planning scheme output by the target motion planning engine, the method further includes: detecting potential disturbance factors on the target running state data of the plurality of robots; if no potential disturbance factors exist, then determining the first motion planning engine as the target motion planning engine; if potential disturbance factors exist, then determining the second motion planning engine as the target motion planning engine, or determining the first motion planning engine and the second motion planning engine as the target motion planning engine; wherein the planning logic of the first motion planning engine is a baseline-free planning without degradation, the planning logic of the second motion planning engine is a degradation planning, and the second motion planning engine has an embedded degradation strategy library.

[0111] Potential disturbance factors refer to various unforeseen factors that may cause abnormal robot operation and affect the achievement of overall goals.

[0112] In some embodiments, potential disturbance factors can be extracted based on the ISO 13849 Fault Tree Analysis (FTA) model. Potential disturbance factors include robotic arm joint failures, communication anomalies, energy anomalies, environmental changes, load changes, program errors, visual perception failures, positioning drift, etc.

[0113] The first action planning engine uses a non-degradation baseline planning logic, which generates the optimal action plan without degradation based on the running scenario, without considering the impact of disturbance factors, and outputs a candidate action planning scheme, namely the first candidate action planning scheme.

[0114] The second action planning engine uses a degraded planning logic and embeds a degraded strategy library built based on historical disturbance handling experience. It can generate at least one candidate action planning scheme for potential disturbance factors, namely the second candidate action planning scheme.

[0115] The degradation strategy library is a database that stores degradation strategies corresponding to various potential disturbance factors, providing strategic support for action planning, and can be updated according to actual operation.

[0116] In some embodiments, the degradation strategy library is shown in Table 1 below, including information such as disturbance type, technical feature description, degradation strategy, and recovery time limit. Here, CRC stands for Cyclic Redundancy Check, UWB for Ultra Wideband, and GPS for Global Positioning System.

[0117] Table 1. Illustration of the Degradation Strategy Library

[0118] In some embodiments, if the detection result is that there are no potential disturbance factors, the multi-robot is in normal operation, the first motion planning engine is determined as the target motion planning engine, and only the first candidate motion planning scheme without degradation is generated as the target motion planning scheme.

[0119] In some embodiments, if at least one potential disturbance factor is detected, the multi-robot is in an abnormal operating state, and the second motion planning engine is determined as the target motion planning engine, or the first motion planning engine and the second motion planning engine are determined together as the target motion planning engine, and at least one candidate motion planning scheme is generated accordingly.

[0120] In some embodiments, the target action planning engine can be selected based on the degree of disturbance impact: If the potential disturbance factors are mild (such as positioning drift, visual perception failure, small impact range, no serious consequences), then the second action planning engine is determined as the target action planning engine, and at least one degradation candidate solution (second candidate action planning solution) is generated. If the potential disturbance factors present moderate / severe disturbances (such as robotic arm joint failure, sudden load changes, program errors, which have a wide impact and may cause production line shutdown), then both the first motion planning engine and the second motion planning engine will be determined as the target motion planning engine. At the same time, a baseline candidate solution (first candidate motion planning solution) and a degraded candidate solution (second candidate motion planning solution) will be generated. Then, the optimal solution will be selected to obtain the target motion planning solution, ensuring that the global goal is achieved.

[0121] By real-time detection and classification of potential disturbance factors, dynamic perception of the operating scenario is achieved. Based on the potential disturbance factors, the action planning engine is adaptively selected. In normal scenarios, baseline planning is used to ensure optimality, while in disturbed scenarios, degraded planning or combined planning is used to ensure robustness. This solves the problems of lag in dynamic disturbance response and reliance on manual parameter tuning in existing technologies, and improves autonomous decision-making ability and anti-interference ability.

[0122] In some embodiments, historical data on the collaborative operation of multiple robots is collected. This historical data includes operational status data, global target quantification data, action execution data, and target achievement results. A training dataset is constructed based on this data. The motion planning engine is then trained using this training dataset until the engine's output reaches a preset standard, resulting in a first motion planning engine. After the first motion planning engine is put into use, it can be adjusted based on real-time operational data to continuously improve the engine's motion planning accuracy and adaptability to real-world scenarios.

[0123] In some embodiments, historical data of multi-robot collaborative operation, including disturbance data (disturbance type, pre-disturbance state, processing action, and goal achievement result) of various scenarios, is collected. Significant and scenario-related degradation processing strategies are then selected and statistically analyzed to construct a degradation strategy library. Based on the historical data and the degradation strategy library, the motion planning engine is trained, focusing on its counterfactual reasoning ability. This allows the engine to learn the correlation weights and causal relationships between disturbances, processing actions, and the global goal achievement result, optimizing engine parameters until the engine output reaches a preset standard, resulting in a second motion planning engine. After the second motion planning engine is put into use, adjustments can be made based on real-time operational data to improve the accuracy and timeliness of the engine's disturbance response and the precision of motion planning.

[0124] In some embodiments, the first action planning engine outputs a first candidate action planning scheme, and the second action planning engine outputs at least one second candidate action planning scheme. The step of determining the target action planning scheme from the at least one candidate action planning scheme that most closely approximates the global target quantification data includes: when the candidate action planning schemes only contain the first candidate action planning scheme, determining the first candidate action planning scheme as the target action planning scheme; when the candidate action planning schemes only contain at least one second candidate action planning scheme, determining the target action planning scheme from the at least one second candidate action planning scheme that most closely approximates the global target quantification data from the at least one second candidate action planning scheme using an action utility evaluation function associated with the global target; and when the candidate action planning schemes contain the first candidate action planning scheme and at least one second candidate action planning scheme, determining the target action planning scheme from the first candidate action planning scheme and at least one second candidate action planning scheme that most closely approximates the global target quantification data from the first candidate action planning scheme and at least one second candidate action planning scheme using the action utility evaluation function.

[0125] The first candidate action planning scheme refers to the candidate action planning scheme output by the first action planning engine, which is generated based on the no-degradation baseline planning logic. It is suitable for normal operation scenarios and takes the global target optimization as the core objective.

[0126] The second candidate action planning scheme refers to the candidate action planning scheme generated by the second action planning engine based on the degradation planning logic. It is applicable to scenarios with or suspected of having disturbances, with the core objective of stabilizing and approaching the global target.

[0127] The action utility evaluation function is a function used to quantify the contribution of candidate action planning schemes to the achievement of the global goal. It can output quantitative data on the achievement of the global goal and is the core basis for selecting the optimal scheme.

[0128] In some embodiments, the action utility evaluation function is an evaluation function constructed based on the decomposition and mixing of global action values ​​using Q-Mixing Networks (QMIX). .in, For the combined actions of all robots - observed historical trajectories, Represents robots The local actions - observed historical trajectories are recorded in real time during operation and stored in the experience playback buffer; It is the combined action of all robots, including the sequence of individual robot actions. Represents robots Selected local action sequence; For robots The local action value function; The global action value function, whose magnitude can be evaluated in the current joint history. The following joint actions were taken. The long-term expected return; the causal utility constraint is embedded in the utility evaluation function of this action. Quantify local action sequences Quantify global target data The strength of the causal effect, the design of the action utility evaluation function guarantees That is, the local value of a single robot. When improving, the total global value It will definitely not decrease; it will only increase or remain unchanged. It is a monotonic hybrid network, i.e., a Q-value hybrid network, with parameters of... ; For robots of The trainable parameters are obtained by iteratively updating the loss function (such as the Temporal Difference Error (TD-error) function) through gradient descent optimization.

[0129] In some embodiments, the first action planning engine and the second action planning engine have embedded the aforementioned action utility evaluation function, which is used to evaluate action utility when generating candidate action planning schemes.

[0130] In some embodiments, the degradation decision process of the second action planning engine is as follows: When potential disturbances exist, the most recently successful strategy for handling the same disturbance is matched and loaded from the degradation strategy library. This mode must meet two conditions: first, statistical significance, meaning the causal relationship implied in the strategy is confirmed to be non-random through hypothesis testing; second, scenario relevance, meaning the type of disturbance handled by the strategy highly matches the current disturbance. The most similar successful strategy to the current scenario is retrieved by querying the degradation strategy library. Based on the action utility evaluation results, the action combination with the best utility is selected from the preset action space A (the set of executable actions in the degradation strategy library) to generate a second candidate action planning scheme.

[0131] In some embodiments, if the first candidate action planning scheme (no disturbance scenario) is the only one in the candidate scheme set, it is directly determined as the target action planning scheme because it is the globally optimal scheme under normal scenario.

[0132] In some embodiments, if there is only at least one second candidate action planning scheme (slight disturbance scenario) in the candidate scheme set, the global target achievement quantification data of each scheme is calculated through the action utility evaluation function, and the scheme whose data is closest to the preset global target quantification data is selected as the target action planning scheme. For example, if the current disturbance is positioning drift, there are two second candidate action planning schemes: Scheme 1 (multi-source sensor fusion correction) has a predicted order time of 29 minutes and energy consumption of 490W, and Scheme 2 (low-speed driving repositioning) has a predicted order time of 31 minutes and energy consumption of 510W. The preset global target is order time ≤ 30 minutes and energy consumption ≤ 500W. Scheme 1 is closer to the target and is determined as the target scheme.

[0133] In some embodiments, if the candidate solution set contains two types of solutions (medium / severe disturbance scenarios), the global goal achievement quantification data of all solutions are calculated using the action utility evaluation function, and the solution closest to the preset global goal is selected as the target action planning solution. For example, if the current disturbance is a robotic arm joint failure, the first candidate action planning solution (normal path + grasping) has a predicted order time of 28 minutes but cannot be executed (robotic arm failure), while the second candidate action planning solution (backup robot takeover + path adjustment) has a predicted order time of 30 minutes and an energy consumption of 500W. In this case, the second candidate solution is selected as the target solution. If the first candidate solution is executable and has better utility, the first candidate action planning solution is selected first.

[0134] The candidate solutions are quantitatively evaluated by using an action utility evaluation function, ensuring that the selection of the target action planning solution is based on scientific evidence.

[0135] The second action planning engine, combined with a degradation strategy library, enables rapid response and autonomous degradation in disturbed scenarios without manual intervention. It formulates differentiated screening rules for different candidate scheme combinations, taking into account both the optimality of normal scenarios and the robustness of disturbed scenarios, ensuring that local action sequences that fit the global goal can be generated in various scenarios, solving the problem of cross-level causal transmission failure, and achieving causal consistency between task instructions and executed actions.

[0136] Step 105: Based on the local action sequence of each robot, generate the action execution command for the robot and send the action execution command to the robot.

[0137] Action execution instructions refer to the specific control instructions that convert a local sequence of actions into a robot that can be recognized and executed, including information such as action type and execution parameters.

[0138] Figure 3 This is an architecture diagram of a multi-robot control system based on multimodal causal reasoning, provided in an embodiment of this application. Figure 3As shown, the multi-robot control system based on multimodal causal reasoning includes a three-layer structure: The causal-driven decision-making layer deploys a multimodal causal reasoning engine (including a first action planning engine, a second action planning engine, and a scheme decision selection engine), a dynamic causal graph database, and a cross-level causal modeling engine to achieve causal modeling, action planning, instruction generation, and causal weight updates. Input data includes running status data and global target quantification data, and output data includes action execution instructions. The cross-level communication bus adopts a shared blackboard architecture, including a main channel (industrial Ethernet) and a backup channel (local area network bus (CAN bus)). It supports the publish-subscribe mode and the Raft consensus algorithm to achieve low-latency data transmission and state synchronization, while transmitting action execution commands and sensor feedback data. The real-time control layer integrates a meta-skill library (grasping unit, obstacle avoidance unit, path planning unit, trajectory planning unit, etc.) and a sensor feedback unit. After receiving action execution instructions, it realizes low-latency action control of the robot. Through encoder data stream and sensor feedback unit, it outputs status feedback to realize dynamic update of causal influence intensity and anomaly detection output.

[0139] The three-layer structure forms a closed loop through action execution instructions and status feedback, ensuring the coordinated stability of the three-layer structure and improving the system's execution accuracy and robustness.

[0140] In some embodiments, for each robot's local action sequence, the local actions in the local action sequence are transformed into specific control instructions by combining the meta-skill library of the real-time control layer. The instruction generation logic of each unit of the meta-skill library is as follows: The gripping unit, based on a proportional-integral-derivative (PID) controller with force feedback, generates gripping instructions that include gripping force (e.g., adjustable from 0-100N), clamping time, and action timing. For example, gripping instruction: force 45N, clamping time 2 seconds, executed at t=5s. The obstacle avoidance unit calculates the collision avoidance vector based on the Velocity Obstacle model and generates an obstacle avoidance command that includes the obstacle avoidance direction (e.g., deviating 30° to the left), the speed adjustment value (e.g., decreasing from 1 m / s to 0.5 m / s), and the timing of execution. The path planning unit generates path execution instructions containing coordinate sequences, driving speed, and turning angle based on the planning results of the A* search algorithm or the Rapidly Exploring Random Tree (RRT) algorithm. For example, the path instruction is: coordinates (10, 20) → (15, 25) → (20, 20), speed 0.8 m / s, turning angle ≤ 45°. The trajectory planning unit, based on the B-spline curve optimizer, generates trajectory instructions that include trajectory control points, motion acceleration, and execution cycle, ensuring smooth execution of the action.

[0141] The standardized generation of action commands is achieved through a meta-skill library, ensuring the executableness and accuracy of the commands.

[0142] In this application, instructions are generated, transmitted and corrected in real time based on a cross-level communication bus. The design of the cross-level communication bus further solves the problem of cross-level causal transmission failure.

[0143] In some embodiments, the cross-level communication bus adopts a shared blackboard architecture and a dual-redundant channel design.

[0144] The shared blackboard architecture comprises a data sharing layer and a decision-making layer. The data sharing layer establishes a globally accessible distributed memory space to store robot states (such as coordinates and robotic arm load for robots with Automated Guided Vehicle (AGV) attributes), global target quantification data (such as order timeliness and energy consumption thresholds), and control commands, achieving global data synchronization. The decision-making layer publishes action execution commands to designated topics. Each robot subscribes to its corresponding topic based on its ID and, upon receiving a command, reports its execution status via the feedback topic, such as command received successfully, action in progress, or execution completed. The Raft consensus algorithm is used to synchronize control processing states, including leader election (heartbeat timeout set to a random value of 150-300ms to avoid simultaneous elections), log replication (the leader encapsulates commands as log entries, synchronizes them to all followers, and commits execution after a majority of nodes persist the logs), and fault recovery (rapid synchronization via log backtracking after a node crash, with recovery time ≤200ms).

[0145] Based on a shared blackboard architecture and dual redundant channel communication design, low-latency and high-reliability transmission of instructions is achieved, meeting the requirements of industrial real-time control.

[0146] The dual-redundant channel consists of a primary channel and a backup channel. The primary channel uses industrial Ethernet with a bandwidth of 1Gbps and mainly transmits real-time control commands (track planning, task allocation, etc.) to ensure high-bandwidth, low-latency transmission. The backup channel uses a CAN bus with a bandwidth of 1Mbps and mainly transmits degradation commands (such as low-speed obstacle avoidance and basic grasping), serving as a backup for the primary channel.

[0147] In some embodiments, the packet loss rate of the main channel is continuously monitored. If the packet loss rate is greater than a set value, such as >5%, the system automatically switches to the backup channel. After the main channel recovers, recovery is performed according to the set recovery time, such as switching back to the main channel 35ms to 50ms after the main channel recovers, to avoid the oscillation of frequent switching affecting the stability of command transmission.

[0148] In some embodiments, during robot command execution, encoders collect motion execution data (such as joint angles, travel distance, and execution time) in real time. This data is then input into a trajectory deviation predictor, such as a gated recurrent unit (GRU), to predict motion execution deviations (such as path offset and force deviation). Compensation commands (such as torque compensation commands) are then generated to correct motion execution parameters in real time, ensuring that the motion execution accuracy matches the global target. For example, if the predicted path offset is 0.5cm, a compensation command is generated to adjust the steering angle and correct the path deviation.

[0149] By dynamically adjusting action execution parameters through a real-time correction mechanism, the accuracy of action execution is improved, ensuring the accurate implementation of local action sequences and ultimately promoting the achievement of global goals.

[0150] In this embodiment, the causal influence strength between the operational status data of each robot in each business dimension and the global target quantification data is calculated. Multimodal data is used to establish a direct and quantifiable causal relationship between the operational status data and the global target. Based on this causal quantification relationship, target operational status data with a high degree of influence on the global target is selected from the operational status data. The operational status data represents the planning influence data of local actions, while the selected target operational status data is data related to local actions and has an impact on the global target. Action planning is performed using this target operational status data and the global target quantification data to obtain the local action sequence of each robot. Accordingly, action execution instructions for each robot are generated and issued, achieving control of multiple robots. The local actions in the local action sequence are determined based on the target operational status data that has a causal relationship with the global target. Relying on the causal quantification relationship between the target operational status data and the global target, the local actions also form a causal quantification relationship with the global target. The generated local actions accurately fit the global target, thus effectively solving the problem that the lack of a causal quantification relationship between the global target of multiple robots and the local actions of a single robot makes it difficult for the local actions of a single robot to accurately adapt to the global target.

[0151] See Figure 4 , Figure 4 This is a structural diagram of a multi-robot control device based on multimodal causal reasoning provided in an embodiment of this application. For ease of explanation, only the parts related to the embodiment of this application are shown.

[0152] The multi-robot control device 400 based on multimodal causal reasoning includes: an acquisition module 401, a calculation module 402, a filtering module 403, a planning module 404, and a generation and sending module 405.

[0153] The acquisition module 401 is used to acquire the operating status data of multiple business dimensions of multiple robots and the global target quantization data corresponding to the global target; the operating status data is the planning influence data of local actions.

[0154] The calculation module 402 is used to calculate the causal influence strength between the operating status data of each business dimension of each robot and the global target quantification data; the causal influence strength characterizes the degree of influence of the operating status data on the global target.

[0155] The filtering module 403 is used to filter out target operating state data whose causal influence intensity meets the set causal influence intensity from the operating state data of multiple business dimensions of each robot.

[0156] The planning module 404 is used to perform motion planning based on the target running state data of multiple robots and the global target quantization data to obtain a local motion sequence for each robot; the local motion sequence includes at least one local motion that fits the global target.

[0157] The generation and sending module 405 is used to generate motion execution instructions for each robot based on the local motion sequence of each robot, and send the motion execution instructions to the robot.

[0158] In some embodiments, the operational status data for each business dimension includes operational status sub-data for at least one modality, and the calculation module is specifically used for: When the operational state data includes operational state sub-data of one modality, the causal influence strength between the operational state sub-data and the global target quantization data is calculated as the causal influence strength between the operational state data and the global target quantization data; or, When the operating state data contains operating state sub-data of multiple modalities, multimodal feature fusion processing is performed on the multiple operating state sub-data to obtain fused state data; The causal influence strength between the fused state data and the global target quantization data is calculated and used as the causal influence strength between the running state data and the global target quantization data.

[0159] In some embodiments, the computing module is further configured to: Cross-modal alignment processing is performed on the various operational state sub-data to obtain multiple cross-modal aligned data; Based on the global target quantization data, the modal influence weights corresponding to each of the cross-modal alignment data are calculated using a cross-modal attention mechanism. The fused state data is obtained by weighting multiple cross-modal alignment data and the modal influence weights corresponding to each cross-modal alignment data.

[0160] In some embodiments, the system further includes a building module for: The operational status data of each business dimension of each robot is taken as local data nodes, the global target quantization data is taken as global target nodes, and causal edges between nodes are constructed based on the causal influence strength between the operational status data and the global target quantization data, so as to obtain a target causal graph containing the local data nodes, the global target nodes and the causal edges. Accordingly, the filtering module is specifically used for: In the target causal graph, target causal edges whose causal influence strength satisfies the set causal influence strength are selected. The running status data corresponding to the target local data node associated with the target causal edge is determined as the target running status data.

[0161] In some embodiments, the planning module is specifically used for: The target operating state data of each robot are vectorized and combined to obtain the initial feature vector of the robot; Each of the initial feature vectors is subjected to dimensionality reduction processing to obtain a low-dimensional feature vector with a set feature dimension; The low-dimensional feature vectors of multiple robots, the global target quantization data, and the task context data are input into the target action planning engine to obtain at least one candidate action planning scheme output by the target action planning engine; each candidate action planning scheme contains multiple candidate action sequences corresponding to the robots. From at least one of the candidate action planning schemes, a target action planning scheme is determined that has the closest global target achievement quantification data to the global target quantification data, wherein the target action planning scheme includes the local action sequence corresponding to each of the robots.

[0162] In some embodiments, the planning module is further configured to: Potential disturbance factors are detected in the target operating state data of multiple robots; If there are no potential disturbances, the first motion planning engine will be determined as the target motion planning engine. If there are potential disturbance factors, the second motion planning engine will be determined as the target motion planning engine, or the first motion planning engine and the second motion planning engine will be determined as the target motion planning engine. The first action planning engine uses a non-degradation baseline planning logic, while the second action planning engine uses a degradation planning logic, and the second action planning engine has an embedded degradation strategy library.

[0163] In some embodiments, the first motion planning engine outputs a first candidate motion planning scheme, the second motion planning engine outputs at least one second candidate motion planning scheme, and the planning module is further configured to: If the candidate action planning scheme only includes the first candidate action planning scheme, the first candidate action planning scheme shall be determined as the target action planning scheme. When the candidate action planning scheme contains only at least one second candidate action planning scheme, the target action planning scheme whose global goal achievement quantification data is closest to the global goal quantification data is determined from at least one second candidate action planning scheme by using the action utility evaluation function associated with the global goal; When the candidate action planning scheme includes the first candidate action planning scheme and at least one second candidate action planning scheme, the target action planning scheme whose global goal achievement quantification data is closest to the global goal quantification data is determined from the first candidate action planning scheme and at least one second candidate action planning scheme using the action utility evaluation function.

[0164] The multi-robot control device based on multimodal causal reasoning provided in this application can realize all the processes of the above-described embodiments of the multi-robot control method based on multimodal causal reasoning, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0165] Figure 5 This is a structural diagram of an electronic device provided in an embodiment of this application. As shown in the figure, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 5 (Only one is shown in the diagram), memory 51, and computer program 52 stored in said memory 51 and executable on said at least one processor 50, wherein said processor 50 executes said computer program 52 to implement the steps in any of the above method embodiments.

[0166] The electronic device 5 can be a desktop computer, laptop, handheld computer, or cloud server, etc. The electronic device 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will understand that... Figure 5 This is merely an example of electronic device 5 and does not constitute a limitation on electronic device 5. It may include more or fewer components than shown, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.

[0167] The processor 50 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0168] The memory 51 can be an internal storage unit of the electronic device 5, such as a hard disk or memory. The memory 51 can also be an external storage device of the electronic device 5, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 51 can include both internal and external storage units of the electronic device 5. The memory 51 is used to store the computer program and other programs and data required by the electronic device. The memory 51 can also be used to temporarily store data that has been output or will be output.

[0169] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0170] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0171] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0172] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0173] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0174] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0175] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0176] The processes in the above-described embodiments can be implemented by a computer program product. When the computer program product is run on an electronic device, the electronic device executes the steps in the above-described method embodiments.

[0177] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A multi-robot control method based on multimodal causal reasoning, characterized in that, include: Acquire operational status data for multiple robots across multiple business dimensions, as well as global target quantification data corresponding to the global target; The operational status data is the planning impact data of local actions; Calculate the causal influence strength between the operational status data of each business dimension of each robot and the global target quantification data; The causal influence intensity characterizes the degree of influence of the operational status data on the global objective; From the operational status data of each robot across multiple business dimensions, target operational status data whose causal influence strength meets the set causal influence strength is selected. Motion planning is performed based on the target operating state data of multiple robots and the global target quantization data to obtain local motion sequences for each robot; each local motion sequence includes at least one local motion that fits the global target. Based on the local action sequence of each robot, an action execution instruction for the robot is generated and sent to the robot.

2. The method according to claim 1, characterized in that, The operational status data for each business dimension includes operational status sub-data of at least one modality. Calculating the causal influence strength between the operational status data for each business dimension of each robot and the global target quantization data includes: When the operational state data includes operational state sub-data of one modality, the causal influence strength between the operational state sub-data and the global target quantization data is calculated as the causal influence strength between the operational state data and the global target quantization data; or, When the operating state data contains operating state sub-data of multiple modalities, multimodal feature fusion processing is performed on the multiple operating state sub-data to obtain fused state data; The causal influence strength between the fused state data and the global target quantization data is calculated and used as the causal influence strength between the running state data and the global target quantization data.

3. The method according to claim 2, characterized in that, The process of performing multimodal feature fusion processing on multiple sub-data of the operating states to obtain fused state data includes: Cross-modal alignment processing is performed on the various operational state sub-data to obtain multiple cross-modal aligned data; Based on the global target quantization data, the modal influence weights corresponding to each of the cross-modal alignment data are calculated using a cross-modal attention mechanism. The fused state data is obtained by weighting multiple cross-modal alignment data and the modal influence weights corresponding to each cross-modal alignment data.

4. The method according to claim 1, characterized in that, Before filtering out target operational status data whose causal influence strength meets the set causal influence strength from the operational status data of multiple business dimensions of each robot, the method further includes: The operational status data of each business dimension of each robot is taken as local data nodes, the global target quantization data is taken as global target nodes, and causal edges between nodes are constructed based on the causal influence strength between the operational status data and the global target quantization data, so as to obtain a target causal graph containing the local data nodes, the global target nodes and the causal edges. The step of filtering target operational status data whose causal influence strength meets a set causal influence strength from the operational status data of multiple business dimensions of each robot includes: In the target causal graph, target causal edges whose causal influence strength satisfies the set causal influence strength are selected. The running status data corresponding to the target local data node associated with the target causal edge is determined as the target running status data.

5. The method according to claim 1, characterized in that, The motion planning based on the target operating state data of multiple robots and the global target quantization data, to obtain the local motion sequence of each robot, includes: The target operating state data of each robot are vectorized and combined to obtain the initial feature vector of the robot; Each of the initial feature vectors is subjected to dimensionality reduction processing to obtain a low-dimensional feature vector with a set feature dimension; The low-dimensional feature vectors of multiple robots, the global target quantization data, and the task context data are input into the target action planning engine to obtain at least one candidate action planning scheme output by the target action planning engine; each candidate action planning scheme contains multiple candidate action sequences corresponding to the robots. From at least one of the candidate action planning schemes, a target action planning scheme is determined that has the closest global target achievement quantification data to the global target quantification data, wherein the target action planning scheme includes the local action sequence corresponding to each of the robots.

6. The method according to claim 5, characterized in that, Before inputting the low-dimensional feature vectors of the multiple robots, the global target quantization data, and the task context data into the target action planning engine to obtain at least one candidate action planning scheme output by the target action planning engine, the method further includes: Potential disturbance factors are detected in the target operating state data of multiple robots; If there are no potential disturbances, the first motion planning engine will be determined as the target motion planning engine. If there are potential disturbance factors, the second motion planning engine will be determined as the target motion planning engine, or the first motion planning engine and the second motion planning engine will be determined as the target motion planning engine. The first action planning engine uses a non-degradation baseline planning logic, while the second action planning engine uses a degradation planning logic, and the second action planning engine has an embedded degradation strategy library.

7. The method according to claim 6, characterized in that, The first motion planning engine outputs a first candidate motion planning scheme, and the second motion planning engine outputs at least one second candidate motion planning scheme. The step of determining, from the at least one candidate motion planning scheme, the target motion planning scheme whose global target achievement quantification data is closest to the global target quantification data includes: If the candidate action planning scheme only includes the first candidate action planning scheme, the first candidate action planning scheme shall be determined as the target action planning scheme. When the candidate action planning scheme contains only at least one second candidate action planning scheme, the target action planning scheme whose global goal achievement quantification data is closest to the global goal quantification data is determined from at least one second candidate action planning scheme by using the action utility evaluation function associated with the global goal; When the candidate action planning scheme includes the first candidate action planning scheme and at least one second candidate action planning scheme, the target action planning scheme whose global goal achievement quantification data is closest to the global goal quantification data is determined from the first candidate action planning scheme and at least one second candidate action planning scheme using the action utility evaluation function.

8. A multi-robot control device based on multimodal causal reasoning, characterized in that, include: The acquisition module is used to acquire the operational status data of multiple robots across multiple business dimensions, as well as the global target quantization data corresponding to the global target. The operational status data is the planning impact data of local actions; The calculation module is used to calculate the causal influence strength between the operating status data of each business dimension of each robot and the global target quantification data; The causal influence intensity characterizes the degree of influence of the operational status data on the global objective; The filtering module is used to filter out target operating state data whose causal influence strength meets the set causal influence strength from the operating state data of multiple business dimensions of each robot. The planning module is used to perform motion planning based on the target running state data of multiple robots and the global target quantization data to obtain a local motion sequence for each robot; the local motion sequence includes at least one local motion that fits the global target; A generation and sending module is used to generate motion execution instructions for each robot based on the local motion sequence of each robot, and send the motion execution instructions to the robot.

9. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device performs the method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, Includes a computer program, which, when run, causes the method as described in any one of claims 1 to 7 to be performed.