Robot control method and device based on multi-modal data fusion perception, equipment and medium

CN122807911APending Publication Date: 2026-09-25CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611187105.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-06
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,实际工况中上述传感器普遍存在退化失效问题:激光雷达在玻璃墙体、高反射金属或长距无特征走廊中易发生镜面反射或几何对称性退化,导致点云匹配失准;深度相机在强光直射、频闪照明或电弧焊弧光干扰下易出现过曝光或深度空洞;IMU(InertialMeasurement Unit,惯性测量单元)虽不受外扰,但其积分漂移随时间呈二次方累积

Benefits of technology

[0015]本申请中,可以通过目标高层语义引导模型对输入的自然语言任务指令和场景图像进行跨模态解析,以从所述目标高层语义引导模型前向传播的隐状态中抽取固定维度的连续稠密向量,并将所述连续稠密向量作为感知引导向量;将所述目标机器人中视觉传感器、激光雷达和惯性测量单元输出的数据流分别输入对应的前置特征网络,以将所述数据流并行映射至统一维度的隐空间,得到相应的视觉特征、几何特征和惯性特征;将所述感知引导向量与所述视觉特征、所述几何特征以及所述惯性特征输入预设动态门控注意力网络,以解算出与所述视觉传感器、所述激光雷达和所述惯性测量单元分别对应的动态门控系数,并利用所述动态门控系数对所述视觉特征、所述几何特征和所述惯性特征进行加权融合,以得到融合特征;基于所述动态门控系数生成预设状态估计器的观测噪声参数,并通过所述状态估计器基于所述观测噪声参数对所述目标机器人的状态进行估计,以得到所述目标机器人的位姿信息,然后通过所述位姿信息以及所述融合特征对所述目标机器人进行动作控制。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807911A_ABST
    Figure CN122807911A_ABST
Patent Text Reader

Abstract

The application discloses a robot control method and device based on multi-modal data fusion perception, equipment and medium, relates to the field of intelligent control, including: cross-modal analysis of natural language instructions and scene images, extracting continuous dense vectors from forward propagation hidden states as perception guide vectors. The data streams of the visual sensor, laser radar and inertial measurement unit are respectively input into the corresponding front feature network and are mapped to a unified hidden space in parallel to obtain visual, geometric and inertial features. The perception guide vector and the three types of features are input into a dynamic gated attention network to calculate the dynamic gating coefficients corresponding to each sensor and obtain the fused features by weighted fusion. At the same time, the observation noise parameters of the state estimator are generated based on the dynamic gating coefficients, the robot pose is estimated by using the state estimator, and finally the action control is realized by combining the pose information and the fused features. Therefore, the precise control of the robot can be realized based on multi-modal data fusion perception.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control, and in particular to a robot control method, device, equipment and medium based on multimodal data fusion perception. Background Technology

[0002] Stable perception in complex industrial scenarios for autonomous mobile robots relies heavily on the synergistic complementarity of multiple heterogeneous sensors, including depth cameras, LiDAR, and inertial measurement units (IMUs). Depth cameras provide rich texture and relative depth information, LiDAR constructs precise 3D geometric contours, and IMUs possess high-frequency motion estimation capabilities unaffected by external interference. Ideally, fusing these three sensors through extended Kalman filtering or graph optimization algorithms can compensate for the limitations of a single mode.

[0003] However, in actual operating conditions, the aforementioned sensors generally suffer from degradation and failure: lidar is prone to specular reflection or geometric symmetry degradation in glass walls, highly reflective metals, or long, featureless corridors, leading to inaccurate point cloud matching; depth cameras are prone to overexposure or depth holes under strong direct light, strobe lighting, or arc welding light interference; and while IMUs (Inertial Measurement Units) are not affected by external disturbances, their integral drift accumulates quadratically over time. How to leverage the advantages of each mode while actively suppressing local degradation is a pressing problem in the field of sensing.

[0004] Existing fusion schemes have several drawbacks: First, traditional methods use fixed weights or static noise covariance, lacking online adaptive adjustment capabilities, and the noise features of failed channels can contaminate global state estimation. Second, there is a severe disconnect between low-level perception and high-level task semantics; the multimodal large model only acts on top-level decision-making and cannot guide the active focusing of sensor feature resources downwards, resulting in task blindness in perception. Third, the computational complexity of standard self-attention mechanisms increases quadratically with the length of the feature sequence. When dealing with dense pixels and hundreds of thousands of point clouds, the edge latency far exceeds the hard real-time requirements of industrial control cycles, severely restricting the robot's safe response capabilities. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a robot control method, device, equipment, and medium based on multimodal data fusion perception, which can achieve precise control of the robot based on multimodal data fusion perception. The specific solution is as follows: In a first aspect, this application discloses a robot control method based on multimodal data fusion sensing, applied to a target robot equipped with multimodal sensors, comprising: The target high-level semantic guidance model performs cross-modal parsing on the input natural language task instructions and scene images to extract a fixed-dimensional continuous dense vector from the hidden state of the forward propagation of the target high-level semantic guidance model, and uses the continuous dense vector as the perception guidance vector. The data streams output from the vision sensor, lidar, and inertial measurement unit in the target robot are respectively input into the corresponding pre-feature network to map the data streams in parallel to a latent space of a unified dimension, thereby obtaining the corresponding visual features, geometric features, and inertial features. The perception guidance vector, the visual features, the geometric features, and the inertial features are input into a preset dynamic gating attention network to calculate the dynamic gating coefficients corresponding to the visual sensor, the lidar, and the inertial measurement unit, respectively. The visual features, the geometric features, and the inertial features are then weighted and fused using the dynamic gating coefficients to obtain fused features. The observation noise parameters of the preset state estimator are generated based on the dynamic gating coefficients, and the state of the target robot is estimated based on the observation noise parameters by the state estimator to obtain the pose information of the target robot. Then, the target robot is controlled by the pose information and the fused features.

[0006] Optionally, before performing cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model to extract a fixed-dimensional continuous dense vector from the hidden state of the forward propagation of the target high-level semantic guidance model, and using the continuous dense vector as the perceptual guidance vector, the method further includes: The backbone network parameters of the pre-trained multimodal model are frozen, and the fine-tuning parameters of the bypass insertion in the pre-trained multimodal model are trained by a preset parameter fine-tuning algorithm to obtain the target high-level semantic guidance model; the fine-tuning parameters include at least one of low-rank matrix, prefix vector or adapter. Accordingly, the step of performing cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model to extract a continuous dense vector of fixed dimensions from the hidden states of the forward propagation of the target high-level semantic guidance model includes: The target high-level semantic guidance model performs cross-modal parsing on the input natural language task instructions and scene images, and extracts the fixed-dimensional continuous dense vector from the hidden state through a linear pooling projection layer cascaded after the last hidden state of the target high-level semantic guidance model.

[0007] Optionally, after performing cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model to extract a fixed-dimensional continuous dense vector from the hidden state of the forward propagation of the target high-level semantic guidance model, and using the continuous dense vector as the perception guidance vector, the method further includes: Perform semantic consistency verification on the perception guidance vector; Accordingly, the semantic consistency check of the perception guidance vector includes: Obtain a pre-stored semantic prototype vector library and calculate the cosine similarity between the perception guidance vector and each semantic prototype vector in the semantic prototype vector library. If the first cosine similarity with the largest value in the cosine similarity is less than the preset consistency threshold, then the perception guidance vector will be reverted to the first historical perception guidance vector that passed the semantic consistency check last time or the preset default security guidance vector. If the perception guidance vector passes the semantic consistency check, then the perception guidance vector is smoothed by an exponential moving average, and the divergence between two adjacent perception guidance vectors is monitored. If the divergence exceeds a preset divergence threshold, the update is rejected, and the second historical perception guidance vector after the previous smoothing is used.

[0008] Optionally, the step of inputting the data streams output from the vision sensor, lidar, and inertial measurement unit in the target robot into the corresponding pre-feature network to map the data streams in parallel to a latent space of a unified dimension, thereby obtaining the corresponding visual features, geometric features, and inertial features, includes: The point cloud data output by the lidar is input into the point cloud feature extraction network to extract the corresponding curvature and edge topology features; The image data output by the vision sensor is input into a lightweight image feature extraction network to extract the corresponding texture and boundary features; The pre-integrated result of the output data of the inertial measurement unit is input into a temporal convolutional network to extract the corresponding transient kinematic features; The curvature and edge topological features, texture and boundary features, and transient kinematic features are mapped to a latent space of a unified dimension to obtain the corresponding geometric features, visual features, and inertial features.

[0009] Optionally, the step of inputting the perception guidance vector, the visual features, the geometric features, and the inertial features into a preset dynamic gating attention network to calculate the dynamic gating coefficients corresponding to the visual sensor, the lidar, and the inertial measurement unit, respectively, and then using the dynamic gating coefficients to perform weighted fusion of the visual features, the geometric features, and the inertial features to obtain fused features includes: The perception guidance vector is concatenated with the visual features, the geometric features, and the inertial features to form a joint feature vector; The joint feature vector is input into the visual branch, geometric branch and inertial branch of the preset dynamic gating attention network, respectively, and the inactive confidence score is output through the visual branch and the geometric branch, respectively. Obtain sensor status diagnostic information and generate degradation flag bits corresponding to the visual sensor and the lidar based on the confidence score; If the degradation flag indicates normal operation, the confidence score is modulated by the first temperature parameter and then activated to obtain the corresponding visual gating coefficient and geometric gating coefficient; the visual gating coefficient is the dynamic gating coefficient corresponding to the visual sensor, and the geometric gating coefficient is the dynamic gating coefficient corresponding to the lidar. If the degradation flag indicates degradation, the confidence scores output by the visual branch and the geometric branch are normalized after being modulated by the second temperature parameter to obtain the visual gating coefficient and the geometric gating coefficient; the second temperature parameter is less than the first temperature parameter. The inertial branch outputs the corresponding inertial gating coefficient using a preset activation function; the inertial gating coefficient is the dynamic gating coefficient corresponding to the inertial measurement unit. The visual features, geometric features, and inertial features are weighted and fused using the visual gating coefficient, the geometric gating coefficient, and the inertial gating coefficient to obtain fused features.

[0010] Optionally, the step of generating observation noise parameters for a preset state estimator based on the dynamic gating coefficients, estimating the state of the target robot based on the observation noise parameters using the state estimator to obtain the pose information of the target robot, and then performing motion control on the target robot using the pose information and the fused features includes: The dynamic gating coefficients are mapped to the observation noise parameters of the state estimator according to a preset inverse proportional relationship; The state estimator performs forward state estimation and Kalman gain estimation based on the observed noise parameters to obtain the pose information of the target robot. The spatial boundary of the target working surface is determined based on the fusion features, and target motion parameters are generated based on the pose information and the spatial boundary; the target motion parameters include motion trajectory and action parameters. The target robot is controlled to perform operational movements based on the target motion parameters.

[0011] Optionally, the robot control method based on multimodal data fusion perception further includes: The pre-integration residual of the inertial measurement unit is calculated by the state estimator. If the pre-integration residual continuously exceeds the preset residual threshold and no effective observation data from the visual sensor and the lidar is obtained within a preset time period, the target robot is stopped.

[0012] Secondly, this application discloses a robot control device based on multimodal data fusion perception, applied to a target robot equipped with multimodal sensors, comprising: The semantic guidance module is used to perform cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model, so as to extract a continuous dense vector of fixed dimension from the hidden state of the forward propagation of the target high-level semantic guidance model, and use the continuous dense vector as the perception guidance vector. The feature encoding module is used to input the data streams output by the vision sensor, lidar and inertial measurement unit in the target robot into the corresponding pre-feature network, so as to map the data streams in parallel to the latent space of the same dimension to obtain the corresponding visual features, geometric features and inertial features. The dynamic gating fusion module is used to input the perception guidance vector, the visual features, the geometric features, and the inertial features into a preset dynamic gating attention network to calculate the dynamic gating coefficients corresponding to the visual sensor, the lidar, and the inertial measurement unit, respectively, and to use the dynamic gating coefficients to perform weighted fusion of the visual features, the geometric features, and the inertial features to obtain fused features. The state estimation and control module is used to generate observation noise parameters of a preset state estimator based on the dynamic gating coefficients, and to estimate the state of the target robot based on the observation noise parameters through the state estimator to obtain the pose information of the target robot. Then, the target robot is controlled by the pose information and the fused features.

[0013] Thirdly, this application discloses an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the robot control method based on multimodal data fusion perception as described above.

[0014] Fourthly, this application discloses a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the robot control method based on multimodal data fusion perception as described above.

[0015] In this application, a target high-level semantic guidance model can be used to perform cross-modal parsing on the input natural language task instructions and scene images to extract a continuous dense vector of fixed dimension from the hidden state of the forward propagation of the target high-level semantic guidance model, and use the continuous dense vector as the perception guidance vector; the data streams output by the visual sensor, lidar and inertial measurement unit in the target robot are respectively input into the corresponding pre-feature network to map the data streams in parallel to a hidden space of a unified dimension to obtain the corresponding visual features, geometric features and inertial features; the perception guidance vector is then compared with the visual features, the geometric features and the inertial features. The system inputs a preset dynamic gating attention network to calculate the dynamic gating coefficients corresponding to the visual sensor, the lidar, and the inertial measurement unit, respectively. The visual features, geometric features, and inertial features are then weighted and fused using these dynamic gating coefficients to obtain fused features. Based on the dynamic gating coefficients, observation noise parameters for a preset state estimator are generated. The state estimator then estimates the state of the target robot based on these observation noise parameters to obtain the robot's pose information. Finally, the pose information and the fused features are used to control the target robot's actions.

[0016] Therefore, the method of this application can perform cross-modal parsing of natural language instructions and scene images through a target high-level semantic guidance model. It extracts a fixed-dimensional continuous dense vector from the forward propagation latent state as the perception guidance vector, and ensures its semantic quality through contrastive learning and semantic consistency verification. Then, the data streams from the visual sensor, LiDAR, and inertial measurement unit are mapped in parallel through a front-end network to a unified-dimensional latent space, generating standardized multi-dimensional features. The perception guidance vector and the three types of features are concatenated and input into a dynamic gating attention network to calculate the dynamic scalar gating coefficients for each modality. The gating coefficients are used to perform weighted fusion of features, achieving hard interception before pseudo-features contaminate the global space. Finally, the gating coefficients are inversely mapped to the observation noise covariance matrix of a preset state estimator to form adaptive observation noise. Finally, the state estimator estimates the state of the target robot based on the observation noise parameters to obtain the robot's pose information. Then, the robot's motion is controlled using the pose information and the fused features. In this way, the perception guidance vector enables the underlying perception to have task-selective focusing capabilities, eliminating the task separation between perception and decision-making; the dynamic gating coefficient collapses to near zero in milliseconds under degradation conditions, hard-intercepting the pseudo-feature flow of the failed channel at the feature space entrance, preventing it from polluting the global fusion features; at the same time, the mode adaptive strategy can achieve both normal complementary enhancement and degradation competitive exclusion dual-mode consideration; the mechanism of inversely mapping the gating coefficient to the filter observation noise covariance matrix mathematically strictly guarantees that the Kalman gain component of the degradation channel is automatically cleared to zero, ensuring that the filter does not diverge, and effectively improving the precise control of the robot. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0018] Figure 1 This is a flowchart of a robot control method based on multimodal data fusion perception disclosed in this application; Figure 2 This is a flowchart of a small-sample process instruction and cross-modal prior semantic feature extraction process disclosed in this application; Figure 3 This is a flowchart of a feature extraction and fusion process disclosed in this application; Figure 4 This is a flowchart of a pose information output method disclosed in this application; Figure 5This is a flowchart illustrating the overall workflow of a robot control method based on multimodal data fusion perception disclosed in this application. Figure 6 This is a schematic diagram of a dual-ring collaboration disclosed in this application; Figure 7 This is a schematic diagram of a specific implementation technical route and algorithm flow steps disclosed in this application; Figure 8 This is a schematic diagram of a robot control device based on multimodal data fusion perception disclosed in this application; Figure 9 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Currently, existing fusion solutions have some shortcomings. For example, they use fixed weights or static noise covariance, lack online adaptive adjustment capabilities, and the noise characteristics of failed channels can pollute the global state estimation. There is a serious disconnect between low-level perception and high-level task semantics. The multimodal large model only acts on top-level decision-making and cannot guide the active focusing of sensor feature resources downward, resulting in task blindness in perception. The computational complexity of the standard self-attention mechanism increases quadratically with the length of the feature sequence. When facing dense pixels and hundreds of thousands of point clouds, the edge latency far exceeds the hard real-time requirements of industrial control cycles, which seriously restricts the robot's safe response capability.

[0021] To overcome the aforementioned technical problems, this application discloses a robot control method, device, equipment, and medium based on multimodal data fusion perception, which can achieve precise control of the robot based on multimodal data fusion perception.

[0022] See Figure 1 As shown, this embodiment of the invention discloses a robot control method based on multimodal data fusion sensing, applied to a target robot equipped with multimodal sensors, including: Step S11: Perform cross-modal parsing on the input natural language task instructions and scene images through the target high-level semantic guidance model, so as to extract a continuous dense vector of fixed dimension from the hidden state of the forward propagation of the target high-level semantic guidance model, and use the continuous dense vector as the perception guidance vector.

[0023] In this embodiment, in actual production environments such as industrial spraying, large-area irregular casting coating, and surface treatment of curved parts, to avoid the latency caused by the large number of parameters resulting from directly deploying a full end-to-end multimodal large model, the tendency of general models to exhibit illusions on the sensing side under extreme operating conditions, and the lack of deterministic safety boundaries in fully end-to-end control output, this embodiment adopts a scheme based on lightweight process adaptation and continuous high-order semantic feature extraction with efficient parameter fine-tuning. Specifically, this scheme abandons the direct output of specific low-order motor control quantities by the large model. Instead, it allows for deep parsing of unstructured textual process instructions and synchronous forward propagation to extract dense latent space perception guidance vectors that can be used for guiding the screening of underlying sensor features. This provides high-order task and process context priors for downstream multimodal adaptive gating fusion.

[0024] like Figure 2 The diagram illustrates the processing flow corresponding to small-sample process instructions and cross-modal prior semantic feature extraction. Specifically, before processing, a pre-trained multimodal model needs to be trained to obtain the target high-level semantic guidance model. Specifically, the backbone network parameters of the pre-trained multimodal model need to be frozen, and the fine-tuning parameters inserted in the pre-trained multimodal model are trained using a preset parameter fine-tuning algorithm to obtain the target high-level semantic guidance model. The fine-tuning parameters include at least one of a low-rank matrix, a prefix vector, or an adapter. The selection and localized compression quantization of the pre-trained multimodal model are crucial. A lightweight open-source pre-trained multimodal model with fewer than 7 bytes of parameters is selected as the core of high-level semantic parsing, preferably a VLA (Vision-Language-Action) model, but a multimodal large language model can also be used. To enable the model to run smoothly on in-vehicle edge computing devices, INT4 activation quantization and operator pruning are performed on the weight parameters of the original model offline. Furthermore, the pre-trained multimodal model's built-in visual encoder is used to extract coarse-grained high-dimensional feature grids from the panoramic images captured by the visual sensors inside the paint booth, while the text input end uses a standard word segmenter to convert the input natural language task instructions into discrete text tag sequences.

[0025] Furthermore, a dedicated cross-modal process instruction dataset needs to be constructed. This involves collecting 50 to 80 typical and comprehensive real-world workshop process instructions offline at the work site, covering four types of text instructions: "spraying anti-corrosion primer areas," "uniform overall spraying of high-gloss topcoat," "multi-layer coating of rough casting surfaces," and "path adjustment for specific adhesive masking areas." Each text instruction is paired with a local scene image captured by an airborne camera under the corresponding working condition, and process engineers annotate the corresponding structured meta-action sub-task sequence, forming a paired fine-tuning training set containing unstructured instruction text, global scene images, and structured process meta-action sequences.

[0026] Next, all original backbone network parameter matrices of the pre-trained multimodal model need to be completely frozen. The fine-tuning parameters are only inserted as a bypass in key layers of the Transformer architecture, such as the projection weighting matrix in multi-head self-attention layers and the cross-attention layer with cross-modal feature alignment. When the fine-tuning parameters are low-rank matrices, i.e., Low-Rank Adaptation (LoRA), is the preferred implementation, two low-rank matrices A and B are inserted as a bypass for the original weight matrix. For a 7B-scale model, a rank value r=16 is preferred, strictly controlling the overall fine-tuning parameter update amount to less than 5% of the original model parameter count, thereby reducing the computational cost of model fine-tuning.

[0027] Finally, the AdamW optimizer with weight decay correction was adopted, with the initial learning rate configured in the range of 5e-5 to 1e-4, the batch size set to 4 to 8, and the number of offline fine-tuning training epochs controlled between 20 and 30. After training, the target high-level semantic guided model was obtained. Among them, a process environment-specific noise enhancement strategy was introduced during the fine-tuning process: in the input workpiece color image, the algorithm manually injected proportionally to simulate geometric point cloud degradation noise caused by high concentration of suspended paint fog, pixel overexposure caused by strong specular reflection of large area wet paint film, and Gaussian optical interference noise such as sudden change of illumination and flicker. This forced the large model to learn the high-dimensional feature invariance of extracting stable workpiece topological edges under extreme optical and paint fog interference during the fine-tuning process.

[0028] Furthermore, such as Figure 2 As shown, a high-order continuous sensing guidance vector is required ( Online streaming extraction is performed on the target high-level semantic guidance model. This involves cross-modal parsing of the input natural language task instructions and scene images using the target high-level semantic guidance model. A linear pooling projection layer, cascaded after the last hidden state of the target high-level semantic guidance model, extracts a continuous dense vector of fixed dimension from the hidden state. During online forward propagation inference, any natural language instruction stream input by the user is published to the functional package node via a topic interface. After fine-tuning, a dedicated linear pooling projection layer is designed after the last hidden state of the model. When the node performs forward propagation inference on the input instructions and images, this layer extracts a continuous dense vector from the forward propagation activation values ​​that is strongly correlated with the environmental degradation state and task safety constraints. The dimension is fixed at D (preferably D=128), and this vector is the high-level semantic context feature vector. This feature vector will be asynchronously cached in a concurrent, safe, double-ended queue accessible by the high-frequency control layer.

[0029] The next step requires semantic consistency verification and contrastive learning supervision of the perceptual guidance vectors to ensure that the continuous dense vectors truly encode the task semantics rather than irrelevant noise. Contrastive learning supervision signals need to be introduced during the fine-tuning training phase: an InfoNCE loss function is added to the linear pooling projection layer to ensure that similar process instructions, such as "primer spraying" and "anti-corrosion primer area overlay," are consistent. Clustering in the latent space leads to closer proximity of different types of process instructions, such as "primer spraying" and "masking area bypass". Push further away. The total training loss function is: ;in For the original task loss, To compare the learning loss, λ is used as the balancing weight, preferably λ=0.1. Furthermore, during the online inference phase, semantic consistency verification of the perception guidance vector is required. This involves acquiring a pre-stored semantic prototype vector library and calculating the cosine similarity between the perception guidance vector and each semantic prototype vector in the library. If the first cosine similarity, which has the highest value, is less than a preset consistency threshold, the perception guidance vector is reverted to the first historical perception guidance vector that passed the semantic consistency verification in the previous iteration or a preset default safe guidance vector. If the perception guidance vector passes the semantic consistency verification, it is smoothed using an exponential moving average, and the divergence between two adjacent generated perception guidance vectors is monitored. If the divergence exceeds a preset divergence threshold, the update is rejected, and the second historical perception guidance vector after the previous smoothing is used. Specifically, as shown... Figure 2 As shown, during the online inference phase, the system maintains a pre-stored prototype vector library of process semantics. Among them, each type of process instruction Centroid vector. Each time. After generation, the cosine similarity between the generated vector and all prototype vectors is calculated. If the maximum similarity is lower than a preset threshold δ (preferably δ=0.6), then the current vector is considered unsuitable. There is a risk of semantic misalignment; automatic rollback to the previous verified cycle will be implemented. Or use the default secure boot vector A balanced vector with equal gating weights for each modality is used to prevent incorrect gating coefficients due to erroneous semantics. This is to prevent misadjustment of gating coefficients in adjacent inference cycles. A sharp jump caused the gating coefficient to oscillate, and the system... Implement Exponential Moving Average (EMA) smoothing: Where α is the smoothing coefficient (preferably α=0.3). Simultaneously, a KL (Kullback-Leibler Divergence) divergence threshold is set for monitoring: if two adjacent... If the KL divergence between the two exceeds the preset threshold κ, preferably κ=2.0, an abnormal alarm is triggered and the current update is rejected, and the smoothing result of the previous cycle is used.

[0030] It should be noted that in this embodiment, as an equivalent alternative to LoRA bypass matrix fine-tuning, all Attention parameter matrices within the general large model can be kept completely frozen. Instead, manually or cascaded trainable continuous high-dimensional prefix vectors (Prefix / Prompt Tokens) are added to the very beginning of the input text token sequence. During fine-tuning, only the parameters of this set of prefix vectors are updated. Its advantages are that it does not alter the internal structure of the large model at all, has good component transferability, and is suitable for production lines with relatively fixed workpiece types that mainly perform large-volume standardized spraying.

[0031] This high-level semantic guidance model, based on efficient fine-tuning of small sample parameters, supports deep semantic parsing of unstructured text instructions and stream extracts continuous, dense perception guidance vectors that encapsulate task constraint priors from the forward propagation hidden state space. Through contrastive learning supervision and semantic consistency verification mechanisms, it is ensured that the perception guidance vectors truly encode process semantics rather than irrelevant noise, and automatic regression to safe guidance vectors occurs in case of semantic misalignment. Furthermore, by introducing a bypass lightweight fine-tuning mechanism, the amount of model parameter updates is significantly reduced, eliminating the cost of large-scale, accurate multimodal labeled datasets; convergence is achieved quickly with only 50-80 paired data points. The introduction of contrastive learning supervision further improves the semantic quality of the perception guidance vectors, effectively enhancing gating accuracy with the same amount of data. Moreover, the entire algorithm, after quantization and compression, can be deployed independently on a single automotive edge computing device, maximizing hardware throughput through chip-level heterogeneous computing resource allocation strategies, effectively controlling the overall computing power overhead and memory usage at the automotive edge.

[0032] Step S12: Input the data streams output by the vision sensor, lidar and inertial measurement unit in the target robot into the corresponding pre-feature network, so as to map the data streams in parallel to the latent space of the same dimension to obtain the corresponding visual features, geometric features and inertial features.

[0033] In this embodiment, to address the severe degradation of single-modality features under complex operating conditions, such as scattering noise from LiDAR, loss of feature points from depth cameras, and divergence in visual odometry, a unified implicit feature space mapping mechanism is employed. This involves constructing a dynamically gated attention neural network guided by a high-level perception vector. This network maps the underlying geometric, visual, and inertial motion feature streams in parallel to a unified-dimensional implicit space for tokenization alignment. Simultaneously, by combining high-level process task priors with the topic states of low-level hardware diagnostic nodes, the dynamic scalar gating coefficients of each heterogeneous sensor channel are calculated online. A millisecond-level hard interception and activation switch is established at the very front of the feature layer, fundamentally solving the sensor feature degradation problem under extreme conditions.

[0034] Specifically, such as Figure 3 As shown, the point cloud data output by the LiDAR needs to be input into a point cloud feature extraction network to extract the corresponding curvature and edge topology features. Specifically, the point cloud data output by the LiDAR is input into a lightweight PointNet network, which is dedicated to extracting the curvature and spatial edge topology features of the workpiece or irregularly shaped casting surface, and outputting a geometric local feature vector. The image data output from the vision sensor is input into a lightweight image feature extraction network to extract corresponding texture and boundary features. Specifically, the image data from the depth camera is input into the MobileNetV3 network to extract the surface texture and boundary features of the workpiece, and outputs a visual local feature vector. The pre-integrated results of the inertial measurement unit (IMU) output data are input into a temporal convolutional network to extract the corresponding transient kinematic features. Specifically, the airborne IMU data stream installed at the end of the device undergoes short-period pre-integration along the time axis, and high-frequency transient kinematic features are extracted through a one-dimensional temporal convolutional network, outputting an inertial local feature vector. .

[0035] Then, curvature and edge topological features, texture and boundary features, and transient kinematic features are mapped to a latent space of a unified dimension to obtain corresponding geometric, visual, and inertial features. Specifically, through independent fully connected projection layers, the above features are uniformly and forcibly mapped to the same latent space feature dimension D. In this embodiment, D=256 is set to form standardized visual, geometric, and inertial tokens.

[0036] Step S13: Input the perception guidance vector, the visual features, the geometric features, and the inertial features into a preset dynamic gating attention network to calculate the dynamic gating coefficients corresponding to the visual sensor, the lidar, and the inertial measurement unit, respectively. Then, use the dynamic gating coefficients to perform weighted fusion of the visual features, the geometric features, and the inertial features to obtain fused features.

[0037] In this embodiment, the perceptual guidance vector is first concatenated with visual, geometric, and inertial features to form a joint feature vector. This joint feature vector is then input into the visual, geometric, and inertial branches of a pre-defined dynamic gating attention network. The visual and geometric branches output inactive confidence scores, respectively. Sensor state diagnostic information is then obtained, and the degradation flags corresponding to the visual sensor and the LiDAR are generated based on the confidence scores. Specifically, a multilayer perceptron (MLP) is constructed as a scalar gating generator. This generator simultaneously subscribes frequently to high-order perceptual guidance vectors issued by the slow system. In addition, there are the three local feature tokens that are currently released in a fast system cycle. The four heterogeneous feature vectors are vertically concatenated to construct a joint feature space input vector: The result is then fed into the mapping branch. To ensure the nonlinear expression capability of subsequent multi-path complementary modulation, the lidar branch and the depth camera branch directly output the inactive, unnormalized confidence score (Logits scalar) in the real domain. and Since the inertial measurement unit does not participate in subsequent local space competition, its mapping branch is... The activation function independently generates a closed interval Dynamic weighting coefficients between The specific calculation formula is as follows: ; ; ; in, and These are the weight matrix and bias vector of the fully connected layer for each sensor mapping branch.

[0038] Furthermore, degradation state discrimination is required. Specifically, if the degradation flag indicates normal operation, the confidence scores are modulated by the first temperature parameter and then activated to obtain the corresponding visual gating coefficients and geometric gating coefficients; the visual gating coefficients are the dynamic gating coefficients corresponding to the visual sensor, and the geometric gating coefficients are the dynamic gating coefficients corresponding to the lidar; if the degradation flag indicates degradation, the confidence scores output by the visual branch and the geometric branch are modulated by the second temperature parameter and then normalized to obtain the visual gating coefficients and geometric gating coefficients; the second temperature parameter is less than the first temperature parameter.

[0039] It should be noted that if the depth camera diagnostics are normal and there is no overexposure causing blindness, and (Preferred) ), then the flag bit =0 (normal); otherwise =1 (Degradation). If the lidar diagnosis is normal, the echo noise rate is not excessive, and Then the flag bit =0 (normal); otherwise =1 (degenerate).

[0040] When both channels are normal ( =0 and =0), using independent Sigmoid activation allows both to maintain high confidence simultaneously, achieving complementary enhancement: ; ; Where σ is the Sigmoid function. Temperature hyperparameters for complementary modes (preferred) =1.0), used to adjust the steepness of the Sigmoid.

[0041] When any channel degrades ( =1 and =1), switch to local Softmax normalization with temperature modulation to achieve competitive exclusivity: ; ; in For the competitive mode sharpness temperature hyperparameter (preferred) =0.3). After this modulation, small differences in input confidence will be amplified nonlinearly and exponentially. Once the radar point cloud undergoes geometric distortion due to scattering, causing a slight decrease in the l_lidar value, the normalized visual gating coefficients... It will quickly gain dominance and suppress failed channels through its architectural mechanisms.

[0042] Then, the corresponding inertial gating coefficients are output through the inertial branch using a preset activation function; these inertial gating coefficients are the dynamic gating coefficients corresponding to the inertial measurement unit. Specifically, the outputs of both modes need to be uniformly recorded as the final visual gating coefficients. and geometric gating coefficient This is for use in subsequent steps.

[0043] It should be noted that the gating unit also receives sensor status diagnostic messages from the underlying hardware diagnostic node. Once the diagnostic node determines that the current depth camera is severely blinded due to overexposure caused by strong specular reflection, or that the LiDAR echo noise rate exceeds the standard, the underlying hardware interception mechanism actively intervenes, storing the normalized gating coefficients of the corresponding failed channel in memory. or Force a direct reset to zero.

[0044] Finally, the visual features, geometric features, and inertial features need to be weighted and fused using the visual gating coefficients, geometric gating coefficients, and inertial gating coefficients to obtain the fused features. That is, to eliminate the differences in the latent space mapping basis of heterogeneous feature sources, each local feature vector is first projected onto a unified cascaded semantic space through a learnable linear transformation matrix before fusion, and then weighted by scalar multiplication with the finally established dynamic gating weights to fuse them into a unified composite feature space vector. And send it to the downstream cross-attention network: ; in, , , These are the linear projection alignment matrices for the laser, visual, and inertial feature spaces, respectively.

[0045] It should be noted that, as an alternative approach, on a control platform with ample computing power, the calculation of explicit scalar gating coefficients can be eliminated. Instead, a full spatial-level Cross-Attention module can be introduced at the fusion layer. The token sequences from each sensor are uniformly unfolded and concatenated with the perceptual guidance vector. Using the perceptual guidance vector as the query matrix and the sensor features as the key matrix, a full dot product calculation is performed. This approach can capture more detailed pixel-level or local point cloud meshes of deep complementary guidance relationships.

[0046] In this way, an adaptive gating strategy is adopted to achieve complementary enhancement with independent Sigmoid under normal operating conditions, avoiding the mutual dilution of normal channels by Softmax. Under degraded operating conditions, temperature-modulated Softmax achieves competitive exclusivity, thus resolving the contradiction between single Softmax normalization and multimodal complementary enhancement.

[0047] Step S14: Generate observation noise parameters of a preset state estimator based on the dynamic gating coefficients, and estimate the state of the target robot based on the observation noise parameters through the state estimator to obtain the pose information of the target robot. Then, perform motion control on the target robot using the pose information and the fused features.

[0048] In this embodiment, addressing the sensor observation degradation problem in complex industrial scenarios, a mechanism is proposed to dynamically map the normalized gating coefficients of a multimodal gating network to the observation noise matrix of an Extended Kalman Filter (EKF), namely, the Augmented-State Extended Kalman Filter (ASEKF). This scheme reconstructs the filter's R matrix as a time-adaptive dynamic function R(t). The scalar gating coefficients output by the forward gating network are directly used as the denominator of the diagonal variance penalty for the observation noise. While maintaining the ultra-low latency advantage of the lossless single-step recursive algorithm, this approach establishes a hard interception capability for data degradation and ensures the smoothness and determinism of high-speed, high-maneuver trajectories.

[0049] like Figure 4 The diagram illustrates the specific process of outputting pose information based on high-frequency control filtering and fast closed-loop control in this embodiment. This process involves mapping the dynamic gating coefficients to the observation noise parameters of the state estimator according to a preset inverse relationship. Then, the state estimator performs forward state estimation and Kalman gain estimation based on these observation noise parameters to obtain the pose information of the target robot. Specifically, the core state vector of the state estimator is defined as the nine-degree-of-freedom pose and velocity information of the target robot's end effector in global three-dimensional space. ; in, Let be the Cartesian coordinate position of the spray gun tip in global 3D space. The roll, pitch, and yaw angles (represented by Euler angles) of the spray gun tip in three-dimensional space. Let $\frac{ ... The yaw rate of the nozzle tip around the vertical z-axis is given by ω.

[0050] Based on the kinematic constraints of the robotic arm and the physical transition relationships of the end-effector inertial measurement unit, a continuous-time state transition equation is constructed and linearized to calculate the current system state transition matrix. process noise covariance matrix .

[0051] Assume the nominal base observation noise covariance matrix of the system under standard ideal operating conditions is constant: ; Here, the two components represent the variance of pose noise in laser point cloud registration and the variance of reprojection noise in visual feature point tracking, respectively. The dynamic gating coefficients output by the dynamic gating attention network described in the preceding steps are... and The observation noise parameters, i.e., the dynamic observation noise covariance matrix, are constructed according to a preset inverse proportional relationship: ; in, To prevent small positive constants from being divided by zero, the default value is 0. The adjustable range is ~ ; The preferred non-linear penalty exponential weight is... The optimization range is .

[0052] Furthermore, semantic control interception and state correction using Kalman gain are required. The filter performs single-step recursion in each control clock step: in the prediction step, it subscribes to the angular velocity and linear acceleration data of the IMU at a high frequency of 200Hz to adjust the current prediction state. With predicted covariance Perform high-speed forward state estimation; in the update step, when the latest laser SLAM (Simultaneous Localization and Mapping) point cloud observation or visual pose is received, the above formula is called to calculate the latest dynamic state. The matrix is ​​then substituted into the solution to calculate the Kalman gain matrix. When vision is interfered with by strong reflections, the gating coefficient is reduced. When the variance approaches zero, the visual observation noise variance is amplified tens of thousands of times instantaneously. After calculating the inverse matrix, the Kalman gain... The components corresponding to the visual observation updates are instantly and automatically cleared to zero. The system automatically erases the influence of the degraded visual channel, relying entirely on high-frequency IMU pre-integration and undisturbed sensor modal extrapolation to output a smooth pose, which is then used as the pose information of the target robot. Ultimately, as... Figure 5 As shown, the spatial boundary of the target working surface is determined based on the fusion features, and target motion parameters are generated based on the pose information and the spatial boundary. The target motion parameters include motion trajectory and action parameters. Then, the target robot is controlled to perform working motion according to the target motion parameters.

[0053] To ensure that radar data, visual tracking streams, and high-frequency IMU data entering the filter are on an absolutely unified time base, the system establishes a sliding window alignment mechanism in the ROS2 (Robot Operating System 2) environment. The hard synchronization error of the timestamps of depth camera, LiDAR, and IMU data is limited to within plus or minus 5ms, and any lagging data packets exceeding this time window will be directly and forcibly discarded.

[0054] Therefore, in this embodiment, a high-level semantic guidance model can be used to perform cross-modal parsing of natural language instructions and scene images. A continuous, dense vector of fixed dimensions is extracted from the forward propagation latent state as the perception guidance vector, and its semantic quality is ensured through contrastive learning and semantic consistency verification. Then, the data streams from the visual sensor, LiDAR, and inertial measurement unit are mapped in parallel through a pre-processor network to a latent space of a unified dimension, generating standardized multi-dimensional features. The perception guidance vector and the three types of features are concatenated and input into a dynamic gating attention network to calculate the dynamic scalar gating coefficients for each modality. The gating coefficients are used to perform weighted fusion of features, achieving hard interception before pseudo-features contaminate the global space. Finally, the gating coefficients are inversely mapped to the observation noise covariance matrix of a preset state estimator to form adaptive observation noise. Finally, the state estimator estimates the state of the target robot based on the observation noise parameters to obtain the robot's pose information. Then, the robot's actions are controlled using the pose information and the fused features. In this way, the perception guidance vector enables the underlying perception to have task-selective focusing capabilities, eliminating the task separation between perception and decision-making; the dynamic gating coefficient collapses to near zero in milliseconds under degradation conditions, hard-intercepting the pseudo-feature flow of the failed channel at the feature space entrance, preventing it from polluting the global fusion features; at the same time, the mode adaptive strategy can achieve both normal complementary enhancement and degradation competitive exclusion dual-mode consideration; the mechanism of inversely mapping the gating coefficient to the filter observation noise covariance matrix mathematically strictly guarantees that the Kalman gain component of the degradation channel is automatically cleared to zero, ensuring that the filter does not diverge, and effectively improving the precise control of the robot.

[0055] As a preferred embodiment, such as Figure 5 As shown, Figure 5 The previous example illustrated the slow-thinking layer and the fast-reaction layer. Therefore, this embodiment provides a detailed explanation of what the slow-thinking layer and the fast-reaction layer are. See [link to documentation]. Figure 6 As shown, slow thinking and fast reflection is a dual-loop collaborative mechanism that is vertically layered, horizontally asynchronous, and based on temporal decoupling. It is used to resolve the contradiction between large models with high intelligence but high computational latency, and end effectors with high-speed, high-maneuverability control and low-latency hard real-time performance.

[0056] Slow thinking is a low-frequency, slow-closed-loop process cognition mechanism, operating at a relatively low frequency, such as 2Hz to 5Hz. Within this loop, the high-level semantic guidance model focuses on parsing long-cycle instructions for complex processes. Although a single inference operation requires a computational latency of 30ms to 50ms, its output perceptual guidance vector... It reflects changes in macroscopic process characteristics. Such changes are low-frequency events in the physical space, so they will not affect the immediacy of fine control.

[0057] Fast reflection is a high-frequency control and filtering fast closed-loop system. Evolving in parallel with the slow system is a hard real-time underlying sensing fast reflection loop, operating at ultra-high frequencies, such as 50Hz-100Hz. At each high-frequency clock step, the fast reflection loop does not need to wait for the slow system node to complete its next inference iteration. Instead, it directly and losslessly reuses and latches the previous cycle's cached atomic buffer in shared memory through non-blocking atomic memory read operations. Vector. The gated attention branch utilizes this semantic tone cached in memory, combined with the latest radar and high-frequency IMU raw waveforms at the current millisecond level input, to instantly fine-tune the scalar gated coefficient matrix G at the millisecond level and directly drive the dynamic update of the ASEKF R matrix, keeping the computation latency within 10ms to maintain the hard real-time baseline of high-frequency high-maneuver control.

[0058] Furthermore, the slow-thinking layer and the fast-reflection layer can perform data closed-loop feedback and self-evolution. During actual operation, the underlying hardware interface layer collects high-frequency multimodal synchronous operation logs, including historical gating coefficient fluctuation curves, ASEKF neo-vector residual distribution, online feedback, and trajectory deviations at specific edges. These logs are periodically aggregated into a difficult-condition dataset and fed back to the slow-thinking layer via topic analysis. The slow-thinking layer utilizes these failure or difficult samples from real-world scenarios to perform online incremental fine-tuning of the parameter efficiency fine-tuning layer during robot idle periods. This spontaneously optimizes the task decomposition and gating prediction accuracy for similar complex workpieces or extreme scenarios, achieving closed-loop self-evolution of the system.

[0059] It should be noted that, as a replacement for the underlying ASEKF state recursion, in the asynchronous control fast reflection loop, the EKF core can be replaced with a sliding window nonlinear least squares optimization factor graph architecture based on the iSAM2 (Incremental Smoothing and Mapping 2) operator from the GTSAM (GeorgiaTech Smoothing and Mapping) library. The normalized gating coefficients output by the gating network no longer rewrite the R matrix, but are instead used as constant scaling factors, multiplied in real-time by the respective system information matrices (i.e., the inverse of the covariance) of the laser odometry factor and the visual observation factor in the factor graph. When a sensor channel degrades, the residual term weight of the corresponding factor in the factor graph automatically returns to zero, achieving the same semantically guided heterogeneous adaptive fault-tolerant positioning at the graph optimization level.

[0060] This approach proposes a dual-loop organic collaborative architecture in the system topology, characterized by vertical layering, horizontal asynchronicity, and temporal decoupling—combining slow thinking and fast reflection. This architecture securely isolates the computationally expensive forward inference of large models within a low-frequency, slow system loop, while the underlying feature encoding, gating computation, and filter matrix recursion operate in a hard real-time, fast system loop. At each high-frequency clock step, the fast system reuses the task guidance vector cached in the previous cycle through non-blocking atomic memory read operations, keeping the single-step computation latency below 10ms to adapt to real-time, zero-latency trajectory control under high-speed, high-maneuvering conditions.

[0061] As a preferred embodiment, based on the foregoing embodiments, while the dual-loop collaborative mechanism of slow thinking and fast reflection effectively isolates the impact of slow loop delay on the real-time performance of fast loop, it introduces new security risks. For example, if slow loop inference crashes or outputs abnormally, the fast loop may continue to use expired or incorrect loops. This will cause the gating coefficient to deviate significantly from the actual operating conditions. Furthermore, the aforementioned... Semantic consistency verification needs to work in conjunction with a dual-ring temporal security mechanism. Therefore, this embodiment provides a detailed explanation of how to perform quality assurance and security protection.

[0062] Quality assurance and safety protection can be divided into four types, including: Aging detection, slow loop watchdog heartbeat detection, concurrent safety dual-end queue memory barrier protection, IMU pre-integration residual monitoring and emergency stop.

[0063] in, Aging detection is performed as follows: During fast closed-loop operation, when reading the sensing guidance vector at each high-frequency clock step, the timestamp attached to that vector is also read. If at the current moment and The difference exceeds the preset aging threshold If so, the perception guidance vector is marked as expired, and preferred. The adjustable range is 200ms to 1000ms. Expired perception guidance vectors are not used directly, but are handled according to a preset decay strategy: if the expiration time is within... to In between, the previous effective perception guidance vector is used, but its weights are multiplied by a decay factor: ; If the expiration time exceeds If so, it will completely revert to the preset default security guidance vector.

[0064] Among them, the slow-loop watchdog heartbeat detection is: the slow-closing point is at a fixed period. The optimal time is 200ms, with an adjustable range of 100-500ms. A heartbeat message is published to the ` / vla / heartbeat` topic. The fast closing loop maintains a heartbeat counter; if consecutive heartbeats occur... If no heartbeat is received within a certain cycle, the slow closed-loop system is considered to have collapsed, and it automatically switches to safe mode, which is the preferred mode. The adjustable range is 3 to 10. The preset dynamic gating attention network no longer uses the perception guidance vector, but instead relies solely on the confidence score and the sensor state diagnostic information for gating calculation, while simultaneously triggering a system alarm to notify the operator to intervene.

[0065] The IMU pre-integration residual monitoring and emergency stop mechanism includes: calculating the pre-integration residual of the inertial measurement unit using a state estimator; if the pre-integration residual continuously exceeds a preset residual threshold, and no valid observation data from the visual sensor and LiDAR is acquired within a preset time period, the target robot is stopped. In other words, the filter calculates the IMU pre-integration residual in each prediction step. ;in, These are the pre-integrated measurement values ​​of the inertial measurement unit. These are the predicted measurement values ​​calculated based on the predicted state. Continuously exceeding the preset residual threshold Preferred The adjustable range is ~ ,in The standard deviation of noise is measured for the inertial measurement unit, and within a preset time period. Preferred If no valid observation data is acquired from the vision sensor and the lidar within 1.0 to 5.0 seconds, the system is determined to have entered an irreversible perception failure state and triggers a stop operation: the target robot's robotic arm decelerates to a safe stop position, the moving chassis is locked, and an emergency alarm is sent to the control panel.

[0066] As a preferred embodiment, such as Figure 7 The diagram illustrates the specific implementation technical route and algorithm flow steps of the robot control method based on multimodal data fusion perception described in this application.

[0067] The hardware and software environment configuration for the robot control method based on multimodal data fusion perception described in this application is as follows: The vehicle-mounted edge computing platform uses an NVIDIA Jetson Orin NX (100 TOPS computing power), running an Ubuntu 22.04 real-time operating system and a ROS2 Humble build environment. The vision sensor is an Intel RealSense D435i surround-view depth camera (87°×58° field of view, depth error ≤2%). The LiDAR is a SLAMTEC RPLIDAR A3 (range range 0.1–25m, scanning frequency 10 Hz). The inertial measurement unit is an airborne IMUMPU6050 (sampling rate 100Hz). The actuator is a composite robot (collaborative robotic arm with a load capacity of 3–5 kg, repeatability ±0.02 mm, and a maximum moving chassis speed of 1.2 m / s). All sensors are connected via a hardware-level clock source aligned bus for underlying driver debugging and interface connectivity. Furthermore, an industrial painting scenario can be constructed at a moderate scale using the Gazebo simulation tool, including a sealed spray booth, explosion-proof lighting, irregularly shaped curved castings to be painted, and reserved areas for adhesive masking. The complete Unified Robot Description Format (URDF) model of the target robot is imported, and parameters for spatiotemporal synchronization, path smoothing, and safety interception logic are optimized first in the simulation scenario before being transferred to a physical prototype. In a real operating environment, approximately 50 to 80 core industrial painting process instructions are collected and defined, covering scenarios such as primer, topcoat, clear coat application, and specific boundary path adjustments. Text instructions are paired with locally captured panoramic images under the current operating conditions to construct a highly structured (industrial input instructions, action sequence description) paired multimodal fine-tuning dataset.

[0068] The perception and gating fusion module of the robot control method based on multimodal data fusion perception described in this application is developed as follows: Parallel encoding of heterogeneous local features: Lightweight PointNet network, MobileNetV3 convolutional network, and one-dimensional temporal convolutional network are used as the backbone for modal feature extraction. Surface curvature features, visual texture and boundary features, and high-frequency transient kinematic waveforms at the ends of the laser point cloud are extracted in parallel. The dimensions of the three heterogeneous features are uniformly hard-mapped to 256 dimensions through independent fully connected projection layers, forming standardized visual, geometric, and inertial tokens respectively.

[0069] Gated attention calculation and spatiotemporal synchronization: A spatiotemporal synchronization node is established in the ROS2 environment. A time sliding window is established using an approximate time alignment strategy to strictly limit the synchronization error between subscribed topics to within ±5ms. The synchronized local feature tokens are cascaded with the high-order perception guidance vector issued by the slow system and input into a scalar gating generator (MLP). The scalar gating generator automatically switches the gating mode according to the sensor degradation state: under normal operating conditions, an independent sigmoid function is used to achieve complementary enhancement; under degraded operating conditions, a local softmax function with temperature modulation is used to achieve competitive exclusivity.

[0070] The high-level semantic guidance model for process task parsing and guidance integration of the robot control method based on multimodal data fusion perception described in this application is as follows: Lightweight model fine-tuning: Offline fine-tuning is performed by inserting a low-rank LoRA matrix into the open-source large model as a bypass, keeping the parameter update amount below 5% of the original parameter count. Contrastive learning loss is also introduced. Supervised training is performed on the linear pooling projection layer to enable similar process instructions. Clustering and different categories are used to extend the range and ensure [the following]. The semantic quality is improved. After the last hidden state in the forward propagation of the model, a dedicated linear pooling projection layer is cascaded to extract a dense perception guidance vector with a fixed dimension of 128, which is used to represent the prior dependency weights of the current process task on each perception modality.

[0071] Asynchronous low-frequency distribution of guidance vectors: The fine-tuned model is encapsulated into a slow system function package node, which performs long-cycle logical inference based on the input natural language process instructions. By calling a custom service or publishing the topic / vla / perception_guidance, the perception guidance vector is sent non-blockingly to the fast system node at a low frequency of 2Hz and cached in a shared memory atomic double-ended queue.

[0072] Semantic consistency check and time smoothing: each time After generation, the cosine similarity between the generated vector and the pre-stored process semantic prototype vector library is calculated. If the maximum similarity is lower than δ=0.6, it reverts to the previous valid vector. or Before issuing, Perform EMA smoothing (α=0.3) and monitor neighboring nodes. If the KL divergence exceeds κ=2.0, then the update is rejected.

[0073] The adaptive filter and control system based on the multimodal data fusion perception robot control method described in this application are developed as follows: ASEKF Filter Implementation: An adaptive extended Kalman filter recursive logic is developed, defining a nine-DOF system state vector containing quaternions of position, velocity, and attitude. In the state prediction step, forward state estimation is performed by subscribing to the IMU data stream at an ultra-high frequency of 200Hz. In the observation update step, the gating coefficients generated at the front end are substituted as a penalty term in the denominator into the diagonal component of the dynamic observation noise covariance matrix R(t) (penalty exponent p=2). The physical consistency and dimensionality verification of the inverse proportional mapping function of R(t) are as shown in the previous embodiment and will not be repeated here. When sudden degradation occurs, the corresponding gating value collapses. By nonlinearly amplifying the observation noise variance, the corresponding Kalman update gain component is instantaneously algebraically cleared to zero, completely erasing the influence of the failed noise data. After state correction is completed, a unified multimodal fusion feature space topic / perception / fused_feature is published at high frequency.

[0074] Downstream motion generation and trajectory planning: Execution nodes in the fast system frequently subscribe to fusion feature topics. The semantic localization state estimation branch extrapolates and outputs a drift-free, real-time nine-DOF global pose topic / robot_pose; the embodied motion execution branch directly latches the precise spatial boundary of the target working surface from the 3D fusion features, and the planner generates a collision-free 3D motion trajectory in real time, which is then frequently sent to the drive box via the topic / arm_controller / joint_trajectory to control the desired motion of the motor.

[0075] The security mechanism of the robot control method based on multimodal data fusion perception described in this application is deployed as in the foregoing embodiments. Aging detection, slow loop watchdog heartbeat detection, memory barrier protection for concurrent safety double-ended queues, IMU pre-integration residual monitoring and emergency stop are not elaborated here.

[0076] See Figure 8 As shown, this embodiment of the invention discloses a robot control device based on multimodal data fusion perception, applied to a target robot equipped with multimodal sensors, comprising: The semantic guidance module 11 is used to perform cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model, so as to extract a continuous dense vector of fixed dimension from the hidden state of the forward propagation of the target high-level semantic guidance model, and use the continuous dense vector as the perception guidance vector. The feature encoding module 12 is used to input the data streams output by the vision sensor, lidar and inertial measurement unit in the target robot into the corresponding pre-feature network, so as to map the data streams in parallel to the latent space of the same dimension to obtain the corresponding visual features, geometric features and inertial features. The dynamic gating fusion module 13 is used to input the perception guidance vector, the visual features, the geometric features, and the inertial features into a preset dynamic gating attention network to calculate the dynamic gating coefficients corresponding to the visual sensor, the lidar, and the inertial measurement unit, respectively, and to use the dynamic gating coefficients to perform weighted fusion of the visual features, the geometric features, and the inertial features to obtain fused features. The state estimation and control module 14 is used to generate observation noise parameters of a preset state estimator based on the dynamic gating coefficient, and to estimate the state of the target robot based on the observation noise parameters through the state estimator to obtain the pose information of the target robot. Then, the target robot is controlled by the pose information and the fused features.

[0077] In some embodiments, the robot control device based on multimodal data fusion perception further includes: The model training unit is used to freeze the backbone network parameters of the pre-trained multimodal model and train the fine-tuning parameters of the bypass insertion in the pre-trained multimodal model through a preset parameter fine-tuning algorithm to obtain the target high-level semantic guidance model; the fine-tuning parameters include at least one of low-rank matrix, prefix vector or adapter.

[0078] In some embodiments, the semantic guidance module 11 may specifically include: The semantic guidance unit is used to perform cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model, so as to extract the fixed-dimensional continuous dense vector from the hidden state through a linear pooling projection layer cascaded after the last hidden state of the target high-level semantic guidance model.

[0079] In some embodiments, the robot control device based on multimodal data fusion perception further includes: The semantic consistency verification submodule is used to perform semantic consistency verification on the perception guidance vector; In some embodiments, the semantic consistency verification submodule includes: The cosine similarity calculation unit is used to obtain a pre-stored semantic prototype vector library and calculate the cosine similarity between the perception guidance vector and each semantic prototype vector in the semantic prototype vector library. The first vector determination unit is used to revert the perception guidance vector to the first historical perception guidance vector or the preset default security guidance vector if the first cosine similarity with the largest value in the cosine similarity is less than the preset consistency threshold. The divergence calculation unit is used to perform exponential moving average smoothing on the perception guidance vector if the perception guidance vector passes the semantic consistency check, and to monitor the divergence between two adjacent perception guidance vectors generated. The second vector determination unit is used to reject the current update and use the second historical perception guidance vector after the previous smoothing if the divergence exceeds a preset divergence threshold.

[0080] In some embodiments, the feature encoding module 12 may specifically include: The first feature extraction unit is used to input the point cloud data output by the lidar into the point cloud feature extraction network to extract the corresponding curvature and edge topology features. The second feature extraction unit is used to input the image data output by the visual sensor into a lightweight image feature extraction network to extract the corresponding texture and boundary features; The third feature extraction unit is used to input the pre-integration result of the output data of the inertial measurement unit into the temporal convolutional network to extract the corresponding transient kinematic features; The feature mapping unit is used to map the curvature and the edge topology features, the texture and the boundary features, and the transient kinematic features to a latent space of a unified dimension, so as to obtain the corresponding geometric features, visual features, and inertial features.

[0081] In some embodiments, the dynamic gating fusion module 13 may specifically include: A feature concatenation unit is used to concatenate the perception guidance vector with the visual features, the geometric features, and the inertial features into a joint feature vector; The confidence score output unit is used to input the joint feature vector into the visual branch, geometric branch and inertial branch of the preset dynamic gating attention network respectively, and output the inactive confidence score through the visual branch and the geometric branch respectively. A flag generation unit is used to acquire sensor state diagnostic information and generate degradation flags corresponding to the visual sensor and the lidar in combination with the confidence score. The first gating coefficient determination unit is used to activate the confidence score after modulation by the first temperature parameter if the degradation flag indicates normal, so as to obtain the corresponding visual gating coefficient and geometric gating coefficient; the visual gating coefficient is the dynamic gating coefficient corresponding to the visual sensor, and the geometric gating coefficient is the dynamic gating coefficient corresponding to the lidar; The second gating coefficient determination unit is used to normalize the confidence scores output by the visual branch and the geometric branch after modulating them with a second temperature parameter if the degradation flag indicates degradation, so as to obtain the visual gating coefficient and the geometric gating coefficient; the second temperature parameter is less than the first temperature parameter. The third gating coefficient determination unit is used to output the corresponding inertial gating coefficient through the inertial branch using a preset activation function; the inertial gating coefficient is the dynamic gating coefficient corresponding to the inertial measurement unit; The feature fusion unit is used to perform weighted fusion of the visual features, the geometric features, and the inertial features using the visual gating coefficient, the geometric gating coefficient, and the inertial gating coefficient to obtain fused features.

[0082] In some embodiments, the state estimation and control module 14 may specifically include: The coefficient mapping unit is used to map the dynamic gating coefficients to the observation noise parameters of the state estimator according to a preset inverse proportional relationship. The pose information determination unit is used to perform forward state estimation and Kalman gain estimation based on the observation noise parameters through the state estimator to obtain the pose information of the target robot. A motion trajectory generation unit is used to determine the spatial boundary of the target working surface based on the fused features, and to generate target motion parameters based on the pose information and the spatial boundary; the target motion parameters include motion trajectory and action parameters; A motion control unit is used to control the target robot to perform operational movements based on the target motion parameters.

[0083] In some embodiments, the robot control device based on multimodal data fusion perception may further include: The motion stop control unit is used to calculate the pre-integration residual of the inertial measurement unit through the state estimator. If the pre-integration residual continuously exceeds a preset residual threshold and no effective observation data from the vision sensor and the lidar is acquired within a preset time period, the target robot is stopped.

[0084] Furthermore, embodiments of this application also disclose an electronic device, Figure 9 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0085] Figure 9This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the robot control method based on multimodal data fusion perception disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0086] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0087] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0088] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the robot control method based on multimodal data fusion perception executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0089] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned robot control method based on multimodal data fusion perception. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0090] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0091] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0092] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0093] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0094] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A robot control method based on multimodal data fusion sensing, characterized in that, Applications include target robots equipped with multimodal sensors, including: The target high-level semantic guidance model performs cross-modal parsing on the input natural language task instructions and scene images to extract a fixed-dimensional continuous dense vector from the hidden state of the forward propagation of the target high-level semantic guidance model, and uses the continuous dense vector as the perception guidance vector. The data streams output from the vision sensor, lidar, and inertial measurement unit in the target robot are respectively input into the corresponding pre-feature network to map the data streams in parallel to a latent space of a unified dimension, thereby obtaining the corresponding visual features, geometric features, and inertial features. The perception guidance vector, the visual features, the geometric features, and the inertial features are input into a preset dynamic gating attention network to calculate the dynamic gating coefficients corresponding to the visual sensor, the lidar, and the inertial measurement unit, respectively. The visual features, the geometric features, and the inertial features are then weighted and fused using the dynamic gating coefficients to obtain fused features. The observation noise parameters of the preset state estimator are generated based on the dynamic gating coefficients, and the state of the target robot is estimated based on the observation noise parameters by the state estimator to obtain the pose information of the target robot. Then, the target robot is controlled by the pose information and the fused features.

2. The robot control method based on multimodal data fusion perception according to claim 1, characterized in that, Before performing cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model to extract a fixed-dimensional continuous dense vector from the hidden state of the forward propagation of the target high-level semantic guidance model, and using the continuous dense vector as the perception guidance vector, the method further includes: The backbone network parameters of the pre-trained multimodal model are frozen, and the fine-tuning parameters of the bypass insertion in the pre-trained multimodal model are trained by a preset parameter fine-tuning algorithm to obtain the target high-level semantic guidance model; the fine-tuning parameters include at least one of low-rank matrix, prefix vector or adapter. Accordingly, the step of performing cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model to extract a continuous dense vector of fixed dimensions from the hidden states of the forward propagation of the target high-level semantic guidance model includes: The target high-level semantic guidance model performs cross-modal parsing on the input natural language task instructions and scene images, and extracts the fixed-dimensional continuous dense vector from the hidden state through a linear pooling projection layer cascaded after the last hidden state of the target high-level semantic guidance model.

3. The robot control method based on multimodal data fusion perception according to claim 1, characterized in that, The step of performing cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model to extract a fixed-dimensional continuous dense vector from the hidden state of the forward propagation of the target high-level semantic guidance model, and using the continuous dense vector as the perception guidance vector, further includes: Perform semantic consistency verification on the perception guidance vector; Accordingly, the semantic consistency check of the perception guidance vector includes: Obtain a pre-stored semantic prototype vector library and calculate the cosine similarity between the perception guidance vector and each semantic prototype vector in the semantic prototype vector library. If the first cosine similarity with the largest value in the cosine similarity is less than the preset consistency threshold, then the perception guidance vector will be reverted to the first historical perception guidance vector that passed the semantic consistency check last time or the preset default security guidance vector. If the perception guidance vector passes the semantic consistency check, then the perception guidance vector is smoothed by an exponential moving average, and the divergence between two adjacent perception guidance vectors is monitored. If the divergence exceeds a preset divergence threshold, the update is rejected, and the second historical perception guidance vector after the previous smoothing is used.

4. The robot control method based on multimodal data fusion perception according to claim 1, characterized in that, The data streams output from the visual sensor, lidar, and inertial measurement unit in the target robot are respectively input into the corresponding pre-feature network to map the data streams in parallel to a latent space of a unified dimension, thereby obtaining the corresponding visual features, geometric features, and inertial features, including: The point cloud data output by the lidar is input into the point cloud feature extraction network to extract the corresponding curvature and edge topology features; The image data output by the vision sensor is input into a lightweight image feature extraction network to extract the corresponding texture and boundary features; The pre-integrated result of the output data of the inertial measurement unit is input into a temporal convolutional network to extract the corresponding transient kinematic features; The curvature and edge topological features, texture and boundary features, and transient kinematic features are mapped to a latent space of a unified dimension to obtain the corresponding geometric features, visual features, and inertial features.

5. The robot control method based on multimodal data fusion perception according to claim 1, characterized in that, The process involves inputting the perception guidance vector, visual features, geometric features, and inertial features into a preset dynamic gating attention network to calculate dynamic gating coefficients corresponding to the visual sensor, the lidar, and the inertial measurement unit, respectively. The dynamic gating coefficients are then used to weight and fuse the visual features, geometric features, and inertial features to obtain fused features, including: The perception guidance vector is concatenated with the visual features, the geometric features, and the inertial features to form a joint feature vector; The joint feature vector is input into the visual branch, geometric branch and inertial branch of the preset dynamic gating attention network, respectively, and the inactive confidence score is output through the visual branch and the geometric branch, respectively. Obtain sensor status diagnostic information and generate degradation flag bits corresponding to the visual sensor and the lidar based on the confidence score; If the degradation flag indicates normal operation, the confidence score is modulated by the first temperature parameter and then activated to obtain the corresponding visual gating coefficient and geometric gating coefficient; the visual gating coefficient is the dynamic gating coefficient corresponding to the visual sensor, and the geometric gating coefficient is the dynamic gating coefficient corresponding to the lidar. If the degradation flag indicates degradation, the confidence scores output by the visual branch and the geometric branch are normalized after being modulated by the second temperature parameter to obtain the visual gating coefficient and the geometric gating coefficient; the second temperature parameter is less than the first temperature parameter. The inertial branch outputs the corresponding inertial gating coefficient using a preset activation function; the inertial gating coefficient is the dynamic gating coefficient corresponding to the inertial measurement unit. The visual features, geometric features, and inertial features are weighted and fused using the visual gating coefficient, the geometric gating coefficient, and the inertial gating coefficient to obtain fused features.

6. The robot control method based on multimodal data fusion perception according to claim 1, characterized in that, The process of generating observation noise parameters for a preset state estimator based on the dynamic gating coefficients, estimating the state of the target robot based on the observation noise parameters using the state estimator to obtain the pose information of the target robot, and then performing motion control on the target robot using the pose information and the fused features includes: The dynamic gating coefficients are mapped to the observation noise parameters of the state estimator according to a preset inverse proportional relationship; The state estimator performs forward state estimation and Kalman gain estimation based on the observed noise parameters to obtain the pose information of the target robot. The spatial boundary of the target working surface is determined based on the fusion features, and target motion parameters are generated based on the pose information and the spatial boundary; the target motion parameters include motion trajectory and action parameters. The target robot is controlled to perform operational movements based on the target motion parameters.

7. The robot control method based on multimodal data fusion perception according to any one of claims 1 to 6, characterized in that, Also includes: The pre-integration residual of the inertial measurement unit is calculated by the state estimator. If the pre-integration residual continuously exceeds the preset residual threshold and no effective observation data from the visual sensor and the lidar is obtained within a preset time period, the target robot is stopped.

8. A robot control device based on multimodal data fusion perception, characterized in that, Applications include target robots equipped with multimodal sensors, including: The semantic guidance module is used to perform cross-modal parsing of the input natural language task instructions and scene images through the target high-level semantic guidance model, so as to extract a continuous dense vector of fixed dimension from the hidden state of the forward propagation of the target high-level semantic guidance model, and use the continuous dense vector as the perception guidance vector. The feature encoding module is used to input the data streams output by the vision sensor, lidar and inertial measurement unit in the target robot into the corresponding pre-feature network, so as to map the data streams in parallel to the latent space of the same dimension to obtain the corresponding visual features, geometric features and inertial features. The dynamic gating fusion module is used to input the perception guidance vector, the visual features, the geometric features, and the inertial features into a preset dynamic gating attention network to calculate the dynamic gating coefficients corresponding to the visual sensor, the lidar, and the inertial measurement unit, respectively, and to use the dynamic gating coefficients to perform weighted fusion of the visual features, the geometric features, and the inertial features to obtain fused features. The state estimation and control module is used to generate observation noise parameters of a preset state estimator based on the dynamic gating coefficients, and to estimate the state of the target robot based on the observation noise parameters through the state estimator to obtain the pose information of the target robot. Then, the target robot is controlled by the pose information and the fused features.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the robot control method based on multimodal data fusion perception as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the robot control method based on multimodal data fusion perception as described in any one of claims 1 to 7.