Robot target identification method and system based on multi-source data driving
By using a multi-source data-driven approach, LSTM networks and causal graph neural networks are employed to optimize robot target recognition, solving the problems of recognition accuracy and model complexity in dynamic environments, and achieving efficient target recognition and decision support.
Patent Information
- Application Number
- CN202511008137.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-04
AI Technical Summary
Existing robot target recognition technologies lack environmental adaptability in dynamic environments. Traditional methods cannot cope with sudden changes in lighting and sensor occlusion, and deep learning methods suffer from problems such as bloated models and high computational complexity.
A multi-source data-driven approach is adopted, which generates dynamic environment vectors through LSTM network, adjusts the fusion weights of vision, sonar and lidar in real time, constructs causal graph neural network to separate pseudo-association features, and combines knowledge base retrieval and adversarial optimization to generate recognition confidence, outputting optimized target category and bounding box.
It improves the robot's recognition accuracy and generalization ability in complex and dynamic environments, enhances the interpretability and stability of the decision-making process, reduces the dependence on a large amount of labeled data, and improves the credibility of recognition results and the flexibility of practical applications.
Smart Images

Figure CN120894657A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent robot environmental perception technology, and in particular to a robot target recognition method and system based on multi-source data. Background Technology
[0002] In recent years, with the increasing demands for autonomous navigation and environmental perception in robots, target recognition technology based on multimodal sensor fusion has become a core research direction in the field of intelligent robotics. Existing technologies mainly employ two types of methods: one is multimodal fusion methods based on static weight allocation (such as early BEV fusion and feature-level stitching), which achieve sensor data fusion through fixed rules or shallow networks, but struggles to adapt to dynamic environmental changes; the other is end-to-end recognition methods based on deep learning, which can automatically extract cross-modal features, but suffer from problems such as bloated models and high computational complexity.
[0003] Existing technologies suffer from two key drawbacks: First, the multimodal fusion process lacks environmental adaptability. Traditional methods, employing fixed weights or simple heuristic rules, cannot cope with dynamic conditions such as sudden changes in illumination or sensor occlusion. Second, the feature extraction stage suffers from causal confusion. Mainstream deep networks learn only statistical correlations rather than essential causal relationships, leading to a sharp decline in generalization performance in untrained scenarios. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a robot target recognition method based on multi-source data to solve the problems of dynamic environment fusion weight mismatch and spurious correlation feature interference.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a robot target recognition method based on multi-source data, comprising: collecting multimodal data, extracting parameters such as illumination intensity, target motion speed, and scene complexity in real time, and generating a dynamic environment vector; inputting the dynamic environment vector into an LSTM network, and outputting real-time fusion weights for vision, sonar, and lidar through an activation function; performing weighted fusion of vision, sonar, and lidar based on the real-time fusion weights to construct a causal graph neural network, separating pseudo-association features through counterfactual intervention samples, and outputting a causal feature vector; inputting the causal feature vector into a pre-trained lightweight target recognition model for target classification, triggering knowledge base retrieval and comparative distillation when a new scene is detected, and generating a recognition confidence score; triggering adversarial optimization based on the recognition confidence score, generating a defense strategy, and reconstructing the missing multimodal data through a joint framework, and outputting the optimized target category and bounding box coordinates.
[0008] As a preferred embodiment of the robot target recognition method based on multi-source data driven by the present invention, the multi-modal data includes visual data, sonar data, illumination intensity data, lidar data, and motion state data.
[0009] As a preferred embodiment of the robot target recognition method based on multi-source data driven by the present invention, the specific steps for collecting multimodal data, extracting illumination intensity, target motion speed, and scene complexity parameters in real time, and generating a dynamic environment vector are as follows.
[0010] The collected multimodal data is timestamped based on a hardware clock synchronization protocol to generate a spatiotemporally consistent multimodal data stream;
[0011] Illumination intensity data is obtained by parsing the pixel brightness values of RGB images based on multimodal data streams, the target running speed is calculated by combining optical flow method, and the scene complexity parameters are evaluated by the spatial distribution entropy of three-dimensional point cloud.
[0012] The parsed illumination intensity data, the calculated target running speed, and the evaluated scene complexity parameters are fused using a weighted index to generate a dynamic environment vector.
[0013] As a preferred embodiment of the robot target recognition method based on multi-source data driven by the present invention, the specific steps of inputting the dynamic environment vector into the LSTM network and outputting the real-time fusion weights of vision, sonar, and lidar through the activation function are as follows:
[0014] The dynamic environment vector is normalized to generate a dimension-matched input sequence;
[0015] The input sequence is fed into the time step iteration unit of the LSTM network, and the effective temporal features are filtered through the forget gate and the cell state is updated.
[0016] Based on the updated cell state, multimodal feature vectors are extracted using the output gate of the LSTM network, and weight probability distributions for vision, sonar, and lidar are generated through activation functions.
[0017] The weight probability distribution is normalized using the softmax function to generate real-time fusion weights for vision, sonar, and lidar.
[0018] As a preferred embodiment of the multi-source data-driven robot target recognition method of the present invention, the steps of weighted fusion of vision, sonar, and lidar based on real-time fusion weights to construct a causal graph neural network, separating pseudo-association features through counterfactual intervention of samples, and outputting a causal feature vector are as follows:
[0019] Based on real-time fusion weights, vision, sonar, and lidar are weighted and fused to generate a spatiotemporally aligned multimodal joint feature matrix;
[0020] Based on the multimodal joint feature matrix, a causal graph neural network is constructed to extract causal relationships between nodes;
[0021] By injecting counterfactual intervention samples into the causal graph neural network, and by comparing the differences in feature distribution before and after the intervention, pseudo-correlated features that are not related to causality are separated, and biased causal subgraphs are generated.
[0022] Graph pooling is performed on the debiased causal subgraph to generate causal feature vectors representing cross-modal causality.
[0023] As a preferred embodiment of the robot target recognition method based on multi-source data driven by the present invention, the steps of inputting causal feature vectors into a pre-trained lightweight target recognition model for target classification, triggering knowledge base retrieval and comparative distillation when a new scene is detected, and generating recognition confidence scores are as follows.
[0024] The causal feature vector is input into a pre-trained lightweight target recognition model to generate a probability distribution;
[0025] The scene deviation index is calculated based on the probability distribution. When the scene deviation index exceeds the scene anomaly judgment threshold, the current scene is determined to be a new scene that has not been trained, and the knowledge graph retrieval instruction is triggered.
[0026] Based on the causal association rules and historical feature samples of knowledge graph retrieval instructions, a recognition confidence score is generated by comparing and distilling the feature distribution of a multi-head attention mechanism and a lightweight target recognition model, thereby integrating semantic prior knowledge.
[0027] As a preferred embodiment of the robot target recognition method based on multi-source data driven by the present invention, the steps of triggering adversarial optimization based on recognition confidence, generating a defense strategy, reconstructing missing multimodal data through a joint framework, and outputting optimized target category and bounding box coordinates are as follows:
[0028] When the target recognition confidence level is detected to be lower than the confidence level warning threshold, the adversarial optimization engine is triggered to generate adversarial perturbation parameters;
[0029] The counter-perturbation parameters are injected into the multimodal data stream to perform pixel-level perturbation on the visual data and to reconstruct the topology of the radar point cloud.
[0030] By using a cross-modal joint framework, missing modal completion is performed on the perturbed visual data and the reconstructed radar point cloud, and spatiotemporal features are fused to generate reconstructed data.
[0031] Multi-scale features are extracted from the reconstructed data and the target recognition confidence is recalibrated. The optimized target category and bounding box coordinates are output by combining the non-maximum suppression algorithm.
[0032] Secondly, this invention provides a robot target recognition system driven by multi-source data, comprising an environment perception module, a weight prediction module, a causal modeling module, a scene recognition module, and an adversarial optimization module. The environment perception module collects multimodal data, uses timestamp alignment, and extracts parameters such as illumination intensity, target motion speed, and scene complexity in real time to generate a dynamic environment vector. The weight prediction module inputs the dynamic environment vector into an LSTM network and outputs real-time fusion weights for vision, sonar, and LiDAR through an activation function. The causal modeling module performs weighted fusion of multimodal data based on the real-time fusion weights, constructs a causal graph neural network, separates pseudo-association features through counterfactual intervention, and outputs a causal feature vector. The scene recognition module inputs the causal feature vector into a pre-trained lightweight target recognition model for target classification. When a new scene is detected, it triggers knowledge base retrieval and comparative distillation to generate recognition confidence. The adversarial optimization module triggers adversarial optimization based on the recognition confidence, generates a defense strategy, reconstructs missing multimodal data through a joint framework, and outputs the optimized target category and bounding box coordinates.
[0033] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the robot target recognition method based on multi-source data driven by the first aspect of the present invention.
[0034] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the robot target recognition method based on multi-source data driven by the first aspect of the present invention.
[0035] The beneficial effects of this invention are as follows: By constructing a causal graph neural network and introducing a counterfactual intervention mechanism, it effectively distinguishes between causality and correlation in multimodal sensor data, thereby improving the model's recognition accuracy and generalization ability in complex dynamic environments, and enhancing the interpretability and stability of the decision-making process. Furthermore, by triggering a knowledge base retrieval and comparative distillation mechanism when detecting new scenes, it achieves the generation of recognition confidence based on fused semantic prior knowledge, enabling the method to quickly adapt to unknown environments, reducing dependence on large amounts of labeled data, and improving the credibility of the recognition results and the deployment flexibility in practical applications. Attached Figure Description
[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a flowchart of a robot target recognition method based on multi-source data.
[0038] Figure 2 This is a flowchart illustrating the multimodal data acquisition and dynamic environment vector generation of a robot target recognition method based on multi-source data.
[0039] Figure 3 This is a flowchart of the multimodal fusion and causal feature extraction of a robot target recognition method based on multi-source data.
[0040] Figure 4 This is a flowchart illustrating the target recognition and adversarial optimization of a robot target recognition method driven by multi-source data. Detailed Implementation
[0041] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0042] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0043] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0044] Reference Figures 1-4 This is one embodiment of the present invention, which provides a robot target recognition method based on multi-source data, comprising the following steps:
[0045] S1: Collect multimodal data, extract parameters such as light intensity, target motion speed and scene complexity in real time, and generate dynamic environment vectors.
[0046] S1.1: Multimodal data includes visual data, sonar data, light intensity data, lidar data, and motion state data.
[0047] It should be noted that the multimodal data consists of visual data, sonar data, illumination intensity data, lidar data, and motion state data. The visual data is obtained by parsing the pixel brightness values of RGB images. The sonar data reflects the sound wave reflection characteristics in underwater or confined spaces. The illumination intensity data comes from the quantitative analysis of pixel brightness values in the visual data. The lidar data provides spatial distribution information of three-dimensional point clouds. The motion state data is calculated based on the optical flow method to determine the target's velocity components.
[0048] S1.2: The acquired multimodal data is timestamped based on a hardware clock synchronization protocol to generate a spatiotemporally consistent multimodal data stream.
[0049] The specific process includes: visual data, sonar data, light intensity data, lidar data, and motion state data are timestamped through a hardware clock synchronization protocol. Visual data contains the pixel brightness values of RGB images, sonar data records the sound wave reflection characteristics, light intensity data quantifies the brightness information in visual data, lidar data provides the spatial distribution of three-dimensional point clouds, and motion state data contains the target running velocity component calculated by the optical flow method. The hardware clock synchronization protocol ensures that visual data, sonar data, light intensity data, lidar data, and motion state data collected by different sensors have a unified time base. After timestamping alignment, visual data, sonar data, light intensity data, lidar data, and motion state data form a spatiotemporally consistent multimodal data stream.
[0050] S1.3: Illumination intensity data is obtained by parsing the pixel brightness values of RGB images based on multimodal data streams. The target running speed is calculated using the optical flow method, and the scene complexity parameter is evaluated by the spatial distribution entropy of the 3D point cloud. The expression is:
[0051]
[0052] Where v represents the target running speed, Ω represents the set of pixels in the effective motion region, a represents the pixel index within the effective motion region Ω, and I a Let x represent the brightness value of the a-th pixel, and let u represent the horizontal coordinate axis in the image coordinate system. a Let y represent the horizontal motion component of the a-th pixel, and v represent the vertical coordinate axis in the image coordinate system. a Let represent the optical flow component of the a-th pixel in the vertical direction, exp represent the natural exponential function, and λ represent the illumination intensity adjustment coefficient (0.1≤λ≤5.0). Let represent the gradient vector of the 3D point cloud entropy corresponding to the a-th pixel. This represents the image brightness gradient vector of the a-th pixel.
[0053] The specific process includes: parsing the RGB image pixel brightness values in the multimodal data stream to obtain illumination intensity data; processing the visual data using optical flow to calculate the target's running speed component, which consists of a horizontal motion component and a vertical motion component. The horizontal motion component corresponds to the horizontal coordinate axis in the image coordinate system, and the vertical motion component corresponds to the vertical coordinate axis in the image coordinate system. The spatial distribution entropy of the 3D point cloud is used to evaluate the scene complexity parameter. During the calculation of the scene complexity parameter, the natural exponential function is used to process the illumination intensity adjustment coefficient, which is limited to a range of 0.1 to 5.0. The gradient vector of the 3D point cloud entropy and the image brightness gradient vector participate in the calculation of the scene complexity parameter. The pixel set of the effective motion region contains multiple pixel indices, and each pixel index corresponds to the brightness value of a specific pixel. These calculation processes together constitute the basis for generating dynamic environment vectors.
[0054] S1.4: The parsed light intensity data, the calculated target running speed, and the evaluated scene complexity parameters are fused using a weighted exponent to generate a dynamic environment vector.
[0055] The specific process includes: obtaining illumination intensity data by parsing the pixel brightness values of RGB images; calculating the target running speed using the optical flow method; obtaining scene complexity parameters by evaluating the spatial distribution entropy of the 3D point cloud; integrating the illumination intensity data, target running speed, and scene complexity parameters using a weighted exponential fusion method; the illumination intensity data reflecting the ambient lighting conditions; the target running speed characterizing the target's motion state; and the scene complexity parameters describing the environmental structural features. The weighted exponential fusion method assigns weights based on the relative importance of the illumination intensity data, target running speed, and scene complexity parameters. The dynamic environment vector is generated by integrating the illumination intensity data, target running speed, and scene complexity parameters using the weighted exponential fusion method, and is used to characterize the current environmental state.
[0056] S2: Input the dynamic environment vector into the LSTM network and output the real-time fusion weights of vision, sonar and lidar through the activation function.
[0057] S2.1: Normalize the dynamic environment vector to generate a dimension-matched input sequence.
[0058] The specific process includes normalizing the dynamic environment vector using a maximum-minimum scaling method, adjusting the numerical range of the illumination intensity data, target running speed, and scene complexity parameters to between 0 and 1. The normalized illumination intensity data, target running speed, and scene complexity parameters are arranged in time step order to form an input sequence. The number of dimensions of the input sequence is exactly the same as the number of nodes in the input layer of the LSTM network. The normalization process ensures that the illumination intensity data, target running speed, and scene complexity parameters have the same numerical dimensions in the input sequence. The dimension-matched input sequence can be directly used for processing by the time step iteration unit of the LSTM network.
[0059] S2.2: Input the input sequence into the time step iteration unit of the LSTM network, filter effective temporal features through the forget gate, and update the cell state.
[0060] The specific process includes: the dimension-matched input sequence enters the time-step iteration unit of the LSTM network for processing; the forget gate analyzes the retention probability of light intensity data, target running speed, and scene complexity parameters through the sigmoid activation function; the input gate determines the degree of update of the cell state by the light intensity data, target running speed, and scene complexity parameters at the current time step; the output gate controls the feature output ratio of the light intensity data, target running speed, and scene complexity parameters; the cell state is dynamically updated according to the calculation results of the forget gate and the input gate; the updated cell state stores the long-term temporal dependencies of light intensity data, target running speed, and scene complexity parameters; after the time-step iteration unit is processed, it outputs a combination of light intensity data, target running speed, and scene complexity parameters containing effective temporal features.
[0061] Effective temporal features refer to the key change patterns of time-dependent illumination intensity data, target running speed, and scene complexity parameters that are retained after being filtered by the forget gate of the LSTM network.
[0062] S2.3: Based on the updated cell state, multimodal feature vectors are extracted using the output gate of the LSTM network, and weight probability distributions for vision, sonar, and lidar are generated through activation functions, expressed as follows:
[0063]
[0064] Among them, P t Let W represent the weighted probability distribution of vision, sonar, and lidar at time t, where t represents the timestamp of the current time, and softmax represents the nonlinear normalization function that transforms the multimodal feature vector into a probability distribution. o Let represent the trainable weight matrix of the LSTM output gate, o represent the activation vector of the LSTM output gate, tanh represent the hyperbolic tangent activation function, and C represent the trainingable weight matrix of the LSTM output gate. tThis represents the cell state at time t. M represents the tensor concatenation operation along the feature dimension. e This represents the dynamic feature vector of the environment, where e represents the environmental feature identifier, and b represents the dynamic feature vector of the environment. o G represents the bias top of the LSTM output gate, ⊙ represents the Hadamard product, and G... t This represents the adaptive gating vector generated at time t.
[0065] The specific process includes: the updated cell state is processed through the output gate of the LSTM network; the output gate uses the hyperbolic tangent activation function to process the cell state and generate an activation vector; the activation vector is multiplied by the trainable weight matrix of the LSTM output gate; the environmental dynamic feature vector and environmental feature identifier are merged into a multimodal feature vector through tensor concatenation along the feature dimension; the multimodal feature vector is multiplied by the adaptive gate vector using the Hadamard product and then input into a nonlinear normalization function; the nonlinear normalization function converts the processing result into a weighted probability distribution of visual data, sonar data, and LiDAR data; the weighted probability distribution represents the fusion weight allocation of visual data, sonar data, and LiDAR data at the current timestamp; and the bias term of the LSTM output gate participates in the calculation to adjust the output result during the generation of the weighted probability distribution.
[0066] S2.4: Normalize the weight probability distribution using the softmax function to generate real-time fusion weights for vision, sonar, and lidar. The expression is:
[0067]
[0068] Among them, w i Let exp represent the normalized real-time fusion weights of the i-th type of sensor (i∈{1,2,3} corresponding to vision, sonar, and lidar, respectively), α represent the natural exponential function, and α represent the normalized real-time fusion weights of the i-th type of sensor. i Represents the environmental dynamic modulation coefficient of the i-th type of sensor (0 < α). i <1), s i This represents the real-time confidence score of the i-th type of sensor, and β represents the accuracy compensation strength coefficient (0 < β). i <1), Let α represent the historical measurement variance of the i-th type of sensor, j represent the traversal index of the sensor, and α represent the traversal index of the sensor. j Represents the environmental dynamic modulation coefficient of the j-th type of sensor (0 < α). j <1), s j This represents the real-time confidence score of the j-th type of sensor. This represents the historical measurement variance of the j-th type of sensor.
[0069] The specific process includes: normalizing the weight probability distribution using the softmax function; performing an exponential operation on the weight probability distributions of visual, sonar, and lidar data using the softmax function; processing the weight probability distributions of visual, sonar, and lidar data using the natural exponential function; participating in the weight calculation using the environmental dynamic modulation coefficients of visual, sonar, and lidar data, with the environmental dynamic modulation coefficients' values limited to between 0 and 1; using real-time confidence scores of visual, sonar, and lidar data to reflect the measurement reliability at the current moment; adjusting the weight allocation ratio of visual, sonar, and lidar data using the accuracy compensation intensity coefficient, with the accuracy compensation intensity coefficient's values limited to between 0 and 1; characterizing the long-term measurement stability using the historical measurement variance of visual, sonar, and lidar data; and converting the normalized weight probability distributions of visual, sonar, and lidar data into real-time fusion weights, which are then used for weighted fusion calculations of multimodal data.
[0070] S3: Based on real-time fusion weights, visual, sonar, and lidar signals are weighted and fused to construct a causal graph neural network. By intervening in samples with counterfactual methods, pseudo-associative features are separated, and a causal feature vector is output.
[0071] S3.1: Based on real-time fusion weights, visual, sonar, and lidar signals are weighted and fused to generate a spatiotemporally aligned multimodal joint feature matrix.
[0072] The specific process includes multiplying visual data, sonar data, and lidar data by their respective real-time fusion weights for weighted calculation. The real-time fusion weights of the visual data adjust the contribution of RGB image features, the real-time fusion weights of the sonar data adjust the participation ratio of sound wave reflection features, and the real-time fusion weights of the lidar data control the influence intensity of 3D point cloud features. The weighted visual data, sonar data, and lidar data are then integrated through feature-level fusion. Feature-level fusion ensures that the visual data, sonar data, and lidar data remain aligned in the spatiotemporal dimension. The spatiotemporally aligned visual data, sonar data, and lidar data are combined to form a multimodal joint feature matrix. The rows of the multimodal joint feature matrix represent the time step sequence, and the columns contain the fused feature vectors of the visual data, sonar data, and lidar data. The multimodal joint feature matrix serves as the input feature of the causal graph neural network.
[0073] S3.2: Based on the multimodal joint feature matrix, construct a causal graph neural network to extract causal relationships between nodes.
[0074] The specific process includes: constructing a causal graph neural network using a multimodal joint feature matrix as input; nodes in the causal graph neural network correspond to feature elements of visual data, sonar data, and lidar data; directed edges between nodes represent causal relationships between feature elements of visual data, sonar data, and lidar data; the causal graph neural network extracts causal dependencies between visual data nodes, sonar data nodes, and lidar data nodes through graph convolution operations; causal correlation analysis is performed between the feature vectors of visual data nodes and the feature vectors of sonar data nodes; causal strength evaluation is performed between the feature vectors of lidar data nodes and the feature vectors of visual data nodes; the directed acyclic graph output by the causal graph neural network clearly represents the causal influence direction between feature elements of visual data, sonar data, and lidar data; and the extracted causal relationships are used for feature decoupling operations on subsequent counterfactual intervention samples.
[0075] S3.3: Inject counterfactual intervention samples into the causal graph neural network, and by comparing the differences in feature distribution before and after the intervention, separate pseudo-correlated features that are not related to causality and generate a debiased causal subgraph.
[0076] The specific process includes: a causal graph neural network receives counterfactual intervention samples as additional input. These counterfactual intervention samples are generated by modifying specific feature values of visual, sonar, and lidar data. The distributions of visual, sonar, and lidar data features before and after intervention are compared with their corresponding distributions after intervention. The differences in visual data feature distributions reflect the degree of influence of changes in lighting conditions on the recognition results. The differences in sonar data feature distributions show the correlation strength between sound wave reflection features and the essential attributes of the target. The differences in lidar data feature distributions characterize the causal contribution of three-dimensional structural features. By analyzing the changes in feature distributions of visual, sonar, and lidar data before and after intervention, pseudo-correlated features that only have statistical correlation but no causal relationship are identified and eliminated. Feature combinations with stable causal relationships in visual, sonar, and lidar data are retained, ultimately forming a biased causal subgraph that has removed environmental interference and measurement bias.
[0077] S3.4: Perform graph pooling on the debiased causal subgraph to generate causal feature vectors representing cross-modal causality.
[0078] The specific process includes: the bias-reduced causal subgraph is processed by graph pooling; the graph pooling operation reduces and aggregates the feature vectors of visual data nodes, sonar data nodes, and lidar data nodes; the feature vectors of visual data nodes and sonar data nodes are used to extract cross-modal key features through max pooling; the feature vectors of lidar data nodes and visual data nodes are used to integrate spatial structure information through average pooling; the pooled visual data features, sonar data features, and lidar data features are formed into a unified representation through vector concatenation; the final generated causal feature vector contains the essential causal relationship patterns of visual data, sonar data, and lidar data; and the causal feature vector is used as the input feature of the lightweight target recognition model for subsequent classification tasks.
[0079] S4: Input the causal feature vector into the pre-trained lightweight target recognition model for target classification. When a new scene is detected, trigger knowledge base retrieval and comparative distillation to generate recognition confidence.
[0080] S4.1: Input the causal feature vector into the pre-trained lightweight target recognition model to generate a probability distribution.
[0081] The specific process includes: causal feature vectors are fed into a pre-trained lightweight target recognition model as input data for processing; the lightweight target recognition model extracts spatial features of the causal feature vectors through convolutional layers; the spatial features are converted into class score vectors through fully connected layers; the class score vectors are converted into probability distributions through a softmax function; the probability distributions represent the predicted probabilities of each predefined target category corresponding to the causal feature vectors; the pre-trained lightweight target recognition model ensures the accuracy of classification of causal feature vectors while maintaining a low number of parameters; and the generated probability distributions are used for subsequent scene deviation index calculation and new scene detection and judgment.
[0082] Predefined target categories are fixed sets of classifications established based on the needs of actual application scenarios, by analyzing the distribution of target types in historical data and referring to the knowledge of domain experts.
[0083] S4.2: Calculate the scene deviation index based on probability distribution. When the scene deviation index exceeds the scene anomaly judgment threshold, determine that the current scene is a new scene that has not been trained, and trigger the knowledge graph retrieval command. The expression is:
[0084]
[0085] Where D represents the scene deviation index (0≤D≤1), K represents the total number of predefined categories in the target recognition task, k represents the enumeration index of the predefined categories in the target recognition task, μ represents the dynamic weight adjustment coefficient (0≤μ≤1), and JS represents the Jensen-Shannon divergence. This represents the probability distribution of the k-th category in the current scenario. Let Δ represent the probability distribution of the k-th class in the training set. (k) This represents the absolute value of the difference in prediction entropy for the k-th category.
[0086] The specific process includes: the scene deviation index is obtained by calculating the difference between the probability distribution of the current scene and the probability distribution of the training set; the Jensen-Shannon divergence measures the statistical distance between the probability distribution of the i-th category in the current scene and the probability distribution of the i-th category in the training set; the dynamic weight adjustment coefficient adjusts the contribution weight of the Jensen-Shannon divergence according to the absolute value of the difference in prediction entropy of the i-th category; the absolute value of the difference in prediction entropy reflects the degree of deviation between the prediction uncertainty of the i-th category in the current scene and the baseline of the training set; when the calculated result of the scene deviation index exceeds the scene anomaly judgment threshold, it indicates that there is a significant difference between the probability distribution pattern of the current scene and the training data, at which point a knowledge graph retrieval instruction is triggered to obtain prior knowledge; the knowledge graph retrieval instruction locates the range of semantic information that needs to be supplemented based on the calculated result of the scene deviation index; the dynamic weight adjustment coefficient ensures that the sensitivity of the scene deviation index to changes in the probability distribution of each category remains balanced; and the final output scene deviation index value is limited to the range of 0 to 1 to facilitate comparison with the scene anomaly judgment threshold.
[0087] The scene anomaly detection threshold is a fixed discrimination criterion determined by analyzing the fluctuation range of the probability distribution in the training dataset and optimizing it in combination with the test results of the validation set.
[0088] S4.3: Based on the causal association rules and historical feature samples of knowledge graph retrieval instructions, the feature distribution of the multi-head attention mechanism and the lightweight target recognition model are compared and distilled to generate the recognition confidence that integrates semantic prior knowledge.
[0089] The specific process includes: using causal association rules and historical feature samples obtained from knowledge graph retrieval commands as additional inputs; causal association rules describing the logical constraints between visual data, sonar data, and LiDAR data; historical feature samples providing a reference for feature distribution in similar scenarios; a multi-head attention mechanism processing the causal association rules and the feature distribution currently output by the lightweight target recognition model; the query vector in the multi-head attention mechanism coming from the feature distribution of the lightweight target recognition model; key vectors and value vectors being generated by transforming historical feature samples; attention weights reflecting the degree of matching between the causal association rules and the current features; a comparative distillation process weighted and fused the knowledge graph retrieval results and the feature distribution of the lightweight target recognition model; the weighting coefficients being dynamically determined by the output of the multi-head attention mechanism; the fused feature representation containing comprehensive information from semantic prior knowledge and real-time perception data; and the final generated recognition confidence score considering both current scene features and domain knowledge in the knowledge graph. The recognition confidence score is used for subsequent adversarial optimization decisions.
[0090] S4.4: Initialize a lightweight target recognition model with an input layer, a depthwise separable convolutional sub-network, and a global pooling layer. Preprocess the multimodal data through normalization, image flipping, and color gamut transformation.
[0091] The specific process includes initializing a lightweight target recognition model with an input layer, a depthwise separable convolutional sub-network, and a global pooling layer. Visual data, sonar data, illumination intensity data, LiDAR data, and motion state data are input to the input layer after being aligned according to timestamps. The input layer performs channel-dimensional concatenation and format conversion on the multimodal data. The depthwise separable convolutional sub-network uses a combination of channel-wise and pointwise convolutions to process the concatenated multimodal data. Channel-wise convolutions independently extract the spatial features of each modality, while pointwise convolutions achieve cross-modal feature fusion. The global pooling layer compresses the feature map output by the depthwise separable convolutional sub-network to generate feature vectors. In the preprocessing stage, visual data is normalized to distribute pixel values within a standard range, data diversity is enhanced through random horizontal flipping, and brightness and contrast are adjusted using color gamut transformation to adapt to different lighting conditions. Sonar and LiDAR data are only normalized to maintain physical dimension consistency; illumination intensity and motion state data are directly input without preprocessing. The kernel parameters of the depthwise separable convolutional subnetwork are initialized using the He method, and the ReLU activation function is selected. Batch normalization layers are inserted after each convolutional operation to accelerate training convergence. The global pooling layer uses average pooling to calculate the mean of the features of each channel as the final input features of the lightweight target recognition model.
[0092] S4.5: Train the lightweight target recognition model as a whole until the network error tends to level off, freeze the backbone network parameters of the lightweight target recognition model, and train only the remaining parameters; when the error stabilizes again, unfreeze the backbone parameters and perform the final training of the lightweight target recognition model.
[0093] The specific process includes: the overall training of the lightweight target recognition model uses the cross-entropy loss function as the optimization objective; all trainable parameters of the lightweight target recognition model are updated through the backpropagation algorithm; during training, a validation set is used to monitor changes in network error; when the fluctuation of network error is less than a preset threshold for several consecutive training cycles, it is determined to be in a flattening state; then, all parameters of the depthwise separable convolutional sub-network in the lightweight target recognition model are fixed, and only the parameters of the global pooling layer and subsequent classification layers are allowed to participate in gradient updates, and training continues until the network error reaches a stable state again; finally, the parameter freeze state of the depthwise separable convolutional sub-network is lifted, the joint optimization of all parameters of the lightweight target recognition model is restored, and a strategy of reducing the learning rate is used for final training, so that the lightweight target recognition model can optimize the overall classification performance while maintaining the backbone feature extraction capability.
[0094] S4.6: The lightweight target recognition model is pruned and trained by regularization objective function and pruning mask to gradually reduce the sparsity to the target value. At the same time, the parameter distribution is adjusted based on the weight parameter modulus normalization method to finally generate a lightweight target recognition model that meets the accuracy requirements.
[0095] The specific process includes: during the pruning training phase of the lightweight target recognition model, an L2 regularization term is added to the cross-entropy loss function to form a regularization objective function, and the parameter size of the lightweight target recognition model is constrained through an iterative optimization process; the pruning mask is applied to the weight parameters of the lightweight target recognition model in the form of a binary matrix, and the proportion of zero-value elements in the mask is gradually increased according to a preset sparsity decay strategy during training; after each sparsity adjustment, the lightweight target recognition model is fine-tuned, and the weight vector of each convolution kernel is divided by its L2 norm using the weight parameter magnitude normalization method to ensure that the weight distribution of different convolution kernels remains relatively balanced; after multiple rounds of sparsity adjustment and parameter retraining, the positions of non-zero elements in the pruning mask are fixed, the pruned weight parameters are removed, and finally a lightweight target recognition model with simplified channels and meeting accuracy requirements is obtained.
[0096] S5: Trigger adversarial optimization based on recognition confidence, generate defense strategies, reconstruct missing multimodal data through a joint framework, and output optimized target categories and bounding box coordinates.
[0097] S5.1: When the target recognition confidence level is detected to be lower than the confidence level warning threshold, the adversarial optimization engine is triggered to generate adversarial perturbation parameters.
[0098] The specific process includes comparing the target recognition confidence level with the confidence level warning threshold. When the target recognition confidence level is lower than the confidence level warning threshold, the adversarial optimization engine starts generating adversarial perturbation parameters. The adversarial perturbation parameters are calculated separately for visual data and LiDAR data. The adversarial perturbation parameters for visual data optimize the key pixel region of the RGB image through gradient backpropagation. The adversarial perturbation parameters for LiDAR data adjust the three-dimensional coordinate offset based on point cloud topology analysis. The generation process of adversarial perturbation parameters takes into account the spatiotemporal consistency constraints of visual data and LiDAR data. The adversarial perturbation parameters output by the adversarial optimization engine simultaneously meet the requirements of interpretability of visual data and geometric rationality of LiDAR data. The generated adversarial perturbation parameters will be injected into the multimodal data stream for subsequent missing mode completion processing.
[0099] The confidence warning threshold is the optimal critical value determined by analyzing the historical recognition accuracy distribution of the lightweight target recognition model on the statistical validation set, combined with ROC curve analysis.
[0100] S5.2: Inject the counter-disturbance parameters into the multimodal data stream, perform pixel-level perturbation on the visual data, and reconstruct the topology of the radar point cloud.
[0101] The specific process includes: after injecting adversarial perturbation parameters into the multimodal data stream, the visual data enhances the contrast of key feature regions through pixel-level perturbation, and non-uniform brightness adjustment is applied to the RGB channels through pixel-level perturbation. The LiDAR data optimizes the spatial distribution density of the point cloud through topology reconstruction. The topology reconstruction reconstructs the adjacency relationship of the 3D point cloud based on the Delaunay triangulation method. The pixel-level perturbation of the visual data and the topology reconstruction of the LiDAR data are synchronized in terms of timestamps. The perturbated visual data retains the semantic consistency of the original image, and the reconstructed LiDAR data maintains the geometric topological constraints of the point cloud. The sonar data in the multimodal data stream retains the original measurement values unchanged during the adversarial optimization process. The processed visual data and LiDAR data are used for subsequent cross-modal joint framework to perform missing modality completion.
[0102] S5.3: By using a cross-modal joint framework, missing modal completion is performed on the perturbed visual data and the reconstructed radar point cloud, and spatiotemporal features are fused to generate reconstructed data.
[0103] The specific process includes: a cross-modal joint framework receives perturbed visual data and reconstructed LiDAR data as input; the perturbed visual data provides enhanced 2D image features; and the reconstructed LiDAR data contains optimized 3D spatial information. The cross-modal joint framework establishes a feature correspondence between visual data and LiDAR data through spatiotemporal alignment operations. The RGB features of visual data and the point cloud features of LiDAR data interact through a cross-attention mechanism. Missing sonar data features are estimated through a feature propagation algorithm in the cross-modal joint framework. The feature propagation algorithm uses the texture information of visual data and the depth information of LiDAR data to generate alternative features for sonar data. The spatiotemporal feature fusion process integrates the appearance features of visual data, the structural features of LiDAR data, and the estimated sonar data features. The fused multimodal features are decoded to generate complete reconstructed data. The reconstructed data maintains the consistency of visual data, LiDAR data, and sonar data in the spatiotemporal dimensions.
[0104] S5.4: Extract multi-scale features from the reconstructed data and recalibrate the target recognition confidence. Combine the non-maximum suppression algorithm to output the optimized target category and bounding box coordinates.
[0105] The specific process includes: reconstructing the data through a multi-scale feature extraction network; the multi-scale feature extraction network uses convolutional kernels of different sizes to perform hierarchical feature extraction on the reconstructed data; the multi-scale features of visual data contain local texture details and global semantic information; the multi-scale features of LiDAR data reflect the surface roughness and overall shape contour of the object; and the multi-scale features of sonar data characterize the reflected signal intensity and propagation attenuation characteristics. The extracted multi-scale features are input into a lightweight target recognition model to recalculate the target recognition confidence. The recalibrated target recognition confidence comprehensively evaluates the feature reliability of visual data, LiDAR data, and sonar data. The non-maximum suppression algorithm processes the target candidate boxes generated by the multi-scale features. The non-maximum suppression algorithm selects the optimal detection result based on the recalibrated target recognition confidence. Finally, the optimized target category and bounding box coordinates output maintain geometric consistency in the image space of visual data and the three-dimensional space of LiDAR data. The bounding box coordinates simultaneously meet the pixel accuracy requirements of visual data and the millimeter-level measurement accuracy of LiDAR data.
[0106] This embodiment also provides a robot target recognition system driven by multi-source data, including: an environment perception module, a weight prediction module, a causal modeling module, a scene recognition module, and an adversarial optimization module. The environment perception module is used to collect multimodal data, align it with timestamps, and extract parameters such as illumination intensity, target motion speed, and scene complexity in real time to generate a dynamic environment vector. The weight prediction module is used to input the dynamic environment vector into an LSTM network and output real-time fusion weights for vision, sonar, and LiDAR through an activation function. The causal modeling module is used to perform weighted fusion of multimodal data based on real-time fusion weights, construct a causal graph neural network, separate pseudo-association features through counterfactual intervention samples, and output a causal feature vector. The scene recognition module is used to input the causal feature vector into a pre-trained lightweight target recognition model for target classification. When a new scene is detected, it triggers knowledge base retrieval and comparative distillation to generate recognition confidence. The adversarial optimization module is used to trigger adversarial optimization based on the recognition confidence, generate a defense strategy, reconstruct missing multimodal data through a joint framework, and output the optimized target category and bounding box coordinates.
[0107] This embodiment also provides a computer device applicable to the robot target recognition method based on multi-source data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the robot target recognition method based on multi-source data as proposed in the above embodiment.
[0108] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0109] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the robot target recognition method based on multi-source data drive as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0110] In summary, this invention, by constructing a causal graph neural network and introducing a counterfactual intervention mechanism, effectively distinguishes between causality and correlation in multimodal sensor data, thereby improving the model's recognition accuracy and generalization ability in complex dynamic environments, and enhancing the interpretability and stability of the decision-making process. Furthermore, by triggering a knowledge base retrieval and comparative distillation mechanism when detecting new scenes, it achieves the generation of recognition confidence based on semantic prior knowledge, enabling the method to quickly adapt to unknown environments, reducing dependence on large amounts of labeled data, and improving the credibility of the recognition results and the deployment flexibility in practical applications.
[0111] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A robot target recognition method based on multi-source data, characterized in that: include, Collect multimodal data, extract parameters such as light intensity, target motion speed, and scene complexity in real time, and generate dynamic environment vectors; The dynamic environment vector is input into the LSTM network, and the real-time fusion weights of vision, sonar and lidar are output through the activation function; Based on real-time fusion weights, visual, sonar and lidar are weighted and fused to construct a causal graph neural network. By using counterfactual intervention samples, pseudo-associative features are separated and causal feature vectors are output. The causal feature vector is input into a pre-trained lightweight target recognition model for target classification. When a new scene is detected, knowledge base retrieval and comparative distillation are triggered to generate recognition confidence. Based on the recognition confidence level, adversarial optimization is triggered, a defense strategy is generated, and the missing multimodal data is reconstructed through a joint framework, outputting the optimized target category and bounding box coordinates.
2. The robot target recognition method based on multi-source data as described in claim 1, characterized in that: The multimodal data includes visual data, sonar data, light intensity data, lidar data, and motion state data.
3. The robot target recognition method based on multi-source data as described in claim 2, characterized in that: The process of collecting multimodal data, extracting parameters such as illumination intensity, target motion speed, and scene complexity in real time, and generating a dynamic environment vector involves the following steps: The collected multimodal data is timestamped based on a hardware clock synchronization protocol to generate a spatiotemporally consistent multimodal data stream; Illumination intensity data is obtained by parsing the pixel brightness values of RGB images based on multimodal data streams, the target running speed is calculated by combining optical flow method, and the scene complexity parameters are evaluated by the spatial distribution entropy of three-dimensional point cloud. The parsed illumination intensity data, the calculated target running speed, and the evaluated scene complexity parameters are fused using a weighted index to generate a dynamic environment vector.
4. The robot target recognition method based on multi-source data as described in claim 3, characterized in that: The process of inputting dynamic environment vectors into an LSTM network and outputting real-time fusion weights for vision, sonar, and LiDAR through an activation function is detailed below. The dynamic environment vector is normalized to generate a dimension-matched input sequence; The input sequence is fed into the time step iteration unit of the LSTM network, and the effective temporal features are filtered through the forget gate and the cell state is updated. Based on the updated cell state, multimodal feature vectors are extracted using the output gate of the LSTM network, and weight probability distributions for vision, sonar, and lidar are generated through activation functions. The weight probability distribution is normalized using the softmax function to generate real-time fusion weights for vision, sonar, and lidar.
5. The robot target recognition method based on multi-source data as described in claim 4, characterized in that: The method involves weighted fusion of visual, sonar, and lidar signals based on real-time fusion weights to construct a causal graph neural network. It then uses counterfactual intervention to separate spurious association features from the samples and outputs a causal feature vector. The specific steps are as follows: Based on real-time fusion weights, vision, sonar, and lidar are weighted and fused to generate a spatiotemporally aligned multimodal joint feature matrix; Based on the multimodal joint feature matrix, a causal graph neural network is constructed to extract causal relationships between nodes; By injecting counterfactual intervention samples into the causal graph neural network, and by comparing the differences in feature distribution before and after the intervention, pseudo-correlated features that are not related to causality are separated, and biased causal subgraphs are generated. Graph pooling is performed on the debiased causal subgraph to generate causal feature vectors representing cross-modal causality.
6. The robot target recognition method based on multi-source data as described in claim 5, characterized in that: The process involves inputting causal feature vectors into a pre-trained lightweight target recognition model for target classification. When a new scene is detected, a knowledge base retrieval and comparative distillation are triggered to generate recognition confidence. The specific steps are as follows: The causal feature vector is input into a pre-trained lightweight target recognition model to generate a probability distribution; The scene deviation index is calculated based on the probability distribution. When the scene deviation index exceeds the scene anomaly judgment threshold, the current scene is determined to be a new scene that has not been trained, and the knowledge graph retrieval instruction is triggered. Based on the causal association rules and historical feature samples of knowledge graph retrieval instructions, a recognition confidence score is generated by comparing and distilling the feature distribution of a multi-head attention mechanism and a lightweight target recognition model, thereby integrating semantic prior knowledge.
7. The robot target recognition method based on multi-source data as described in claim 6, characterized in that: The process involves triggering adversarial optimization based on recognition confidence, generating a defense strategy, reconstructing missing multimodal data using a joint framework, and outputting optimized target categories and bounding box coordinates. The specific steps are as follows: When the target recognition confidence level is detected to be lower than the confidence level warning threshold, the adversarial optimization engine is triggered to generate adversarial perturbation parameters; The counter-perturbation parameters are injected into the multimodal data stream to perform pixel-level perturbation on the visual data and to reconstruct the topology of the radar point cloud. By using a cross-modal joint framework, missing modal completion is performed on the perturbed visual data and the reconstructed radar point cloud, and spatiotemporal features are fused to generate reconstructed data. Multi-scale features are extracted from the reconstructed data and the target recognition confidence is recalibrated. The optimized target category and bounding box coordinates are output by combining the non-maximum suppression algorithm.
8. A robot target recognition system based on multi-source data, based on the robot target recognition method based on multi-source data as described in any one of claims 1 to 7, characterized in that: It includes an environmental perception module, a weight prediction module, a causal modeling module, a scene recognition module, and an adversarial optimization module. The environment perception module is used to collect multimodal data, align it with timestamps, and extract parameters such as light intensity, target movement speed, and scene complexity in real time to generate dynamic environment vectors. The weight prediction module is used to input dynamic environment vectors into the LSTM network and output real-time fusion weights of vision, sonar, and LiDAR through activation functions; The causal modeling module is used to perform weighted fusion of multimodal data based on real-time fusion weights, construct a causal graph neural network, separate pseudo-association features through counterfactual intervention samples, and output a causal feature vector. The scene recognition module is used to input causal feature vectors into a pre-trained lightweight target recognition model for target classification. When a new scene is detected, it triggers knowledge base retrieval and comparative distillation to generate recognition confidence. The adversarial optimization module is used to trigger adversarial optimization based on the recognition confidence level, generate defense strategies, reconstruct missing multimodal data through a joint framework, and output the optimized target category and bounding box coordinates.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the robot target recognition method based on multi-source data drive as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the robot target recognition method based on multi-source data drive as described in any one of claims 1 to 7.
Citation Information
Cited By
Robot control method, computer equipment and computer readable storage medium
CN121340316A
Model training method and device, environment quality evaluation method and device and storage medium
CN121524602A
Mountain environment sensing system and method based on multi-sensor fusion
CN121600412A
Traffic robot event identification method based on multi-sensor fusion
CN122223975A
A traffic robot event identification method based on multi-sensor fusion
CN122223975B