Multi-modal information system construction method for multi-tuple unmanned equipment

By building a multi-modal information system for unmanned equipment and using graph convolutional neural networks to fuse the characteristics of different modal data, the problems of insufficient information fusion and unin-depth consideration of modal correlation in traditional technologies are solved, and more efficient information processing and stronger perceived decision-making capabilities are achieved.

CN119942292AActive Publication Date: 2025-05-06XIANGJIANG LAB

Patent Information

Application Number
CN202510448376.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-06
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Traditional multimodal information fusion technology fails to fully utilize and integrate the intrinsic connections and complementarity between different modal data, resulting in the loss of some useful information and lacks in-depth consideration of the correlation between different modalities, especially when performing cross-modal tasks.

Method used

A multimodal information system construction method is proposed. By receiving RGB images, lidar point cloud images and infrared images, features are extracted and a single tuple nodes are formed, edge weights are calculated through the feature space of the binary features, graph structure is constructed, and graph convolution neural network is used for fusion and optimization.

Benefits of technology

This method can better explore and utilize potential connections between different modal data, improve the accuracy and efficiency of information processing, and thus improve the perception and decision-making support capabilities of unmanned equipment in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942292A_ABST
    Figure CN119942292A_ABST
Patent Text Reader

Abstract

The invention relates to a method for constructing a multi-modal information system of multi-tuple unmanned equipment. The method comprises the following steps: receiving multi-modal information collected by the unmanned equipment; features of the three images are extracted respectively, features of the three images are obtained respectively, and the features of the three images serve as unitary nodes of the corresponding images respectively; the features of the three images are paired in pairs and fused to obtain three shared semantic subspaces, and the three shared semantic subspaces serve as different two-tuple features; fusing all the features of the three images to obtain a shared semantic space, and taking the shared semantic space as a triple node; all the unitary group nodes are connected with one another, all the unitary group nodes are connected with the triad nodes, the weight of all edges is calculated based on the feature space of the two-tuple features, and a graph structure is constructed; fusing the graph structure by adopting a graph convolutional neural network to obtain a fusion result; and updating and optimizing the graph structure based on the fusion result to obtain an optimized graph structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of unmanned equipment information processing, and in particular to a method for constructing a multimodal information system of a multi-group unmanned equipment. Background Art

[0002] At present, in the field of unmanned equipment information processing, information processing technology is developing towards a more intelligent and automated direction. Unmanned equipment needs to process and analyze a large amount of data to perform complex tasks. These data usually include RGB image data, LiDAR point cloud data and infrared image data, which provide different perspectives and information levels. Among them, RGB image data provides rich color and texture information, LiDAR point cloud data provides accurate spatial structure information, and infrared image data provides temperature and heat distribution information.

[0003] However, traditional multimodal information fusion technologies have some limitations. When processing data from different modalities, these technologies often fail to fully utilize and integrate the intrinsic connections and complementarities between them, resulting in the loss of some useful information. In addition, since the information from each modality is relatively independent, the correlation between them is poor, and there is a lack of in-depth consideration of the correlation between different modalities, which is particularly evident when performing cross-modal tasks. Summary of the invention

[0004] Based on this, it is necessary to provide a method for constructing a multimodal information system of multi-group unmanned equipment, which includes: S1: receiving multimodal information collected by unmanned equipment, wherein the multimodal information includes RGB images, laser radar point cloud images, and infrared images; S2: extract features from the three images respectively, obtain features of the three images respectively, and use the features of the three images as tuple nodes of the corresponding images respectively; S3: Pair the features of the three images in pairs and fuse them to obtain three shared semantic subspaces, and use the three shared semantic subspaces as different bigram features; fuse all the features of the three images to obtain a shared semantic space, and use the shared semantic space as a triplet node; S4: Connect each tuple node to each other, and connect each tuple node to the triple node, calculate the weight of each edge based on the feature space of the two-tuple feature, and build a graph structure; S5: using a graph convolutional neural network to fuse the graph structure to obtain a fusion result; S6: Based on the fusion result, the graph structure is updated and optimized to obtain an optimized graph structure.

[0005] Preferably, a convolutional neural network is used to extract global features from the RGB image to obtain a first image global feature, and the first image global feature is used as a tuple node of the RGB image; Using an encoder to perform multi-scale feature extraction on the RGB image to obtain multi-scale features of a first image; Using PointNet to extract global features from the laser radar point cloud image to obtain a second image global feature, and using the second image global feature as a tuple node of the laser radar point cloud image; Using ResNet to extract global features of the infrared image to obtain a third image global feature, and using the third image global feature as a tuple node of the infrared image; An encoder is used to extract multi-scale features from the infrared image to obtain multi-scale features of a third image.

[0006] Preferably, the process of fusing the features of the first image with the features of the second image includes: Using a bilinear interpolation method to adjust the global features of the first image and the global features of the second image to the same dimension; Mapping the aligned first image global features and the second image global features to the first feature space through a first fully connected layer; Performing a weighted summation of the aligned first image global features and the second image global features in the first feature space to obtain a first fusion feature; Performing a nonlinear transformation on the first fused feature to obtain a first activated feature; Using a feature pyramid network to perform multi-scale feature extraction and fusion on the first activation feature to obtain a first multi-scale fusion feature; The first multi-scale fusion features are mapped to a first shared semantic subspace through a second fully connected layer.

[0007] Preferably, the process of fusing the features of the first image with the features of the third image includes: Convolving the multi-scale features of the first image at each scale respectively through a 3×3 convolution operation to obtain a first convolution feature; Convolving the third image multi-scale features of each scale respectively through a 3×3 convolution operation to obtain a second convolution feature; Adding the first convolution feature and the second convolution feature to obtain a common feature; Subtract the second convolution feature from the first convolution feature, and divide the result by 2 to obtain a private feature of the RGB image; Subtract the first convolution feature from the second convolution feature, and divide the result by 2 to obtain a private feature of the infrared image; The ECA attention mechanism is used to enhance the common features, private features of RGB images, and private features of infrared images respectively; The enhanced common features, the private features of the RGB image, and the private features of the infrared image are fused at each scale to obtain the second fused features at each scale; The second fusion features of all scales are reconstructed through a decoder to obtain second multi-scale fusion features; and the second multi-scale fusion features are used as the second shared semantic subspace.

[0008] Preferably, the process of fusing the features of the second image with the features of the third image includes: Performing the same spatial transformation on the second image global feature and the third image global feature; Fusing the second image global feature after spatial transformation with the third image global feature to obtain a third fused feature; The third fusion feature is optimized by using a data fitting term, a regularization term and an argmin function to obtain an optimized feature; Use multi-scale analysis to extract optimized features of different scales and fuse them to obtain the third multi-scale fusion feature; The third multi-scale fusion feature is mapped to a third shared semantic subspace through an artificial neural network.

[0009] Preferably, the step of fusing all the features of the three images to obtain a shared semantic space includes: Adjusting the second image global feature and the third image global feature to the same spatial reference system as the first image global feature; Perform feature alignment on the first image global feature, the adjusted second image global feature, and the third image global feature; Perform weighted averaging of the first image global features, the second image global features, and the third image global features that are feature aligned to obtain a fourth fusion feature; Mapping the fourth fused feature into a second activation feature through a third fully connected layer; Use multi-scale analysis to extract second activation features of different scales and fuse them to obtain a fourth multi-scale fusion feature; The fourth multi-scale fusion feature is optimized by using a data fitting term, a regularization term, and an argmin function to obtain the shared semantic space.

[0010] Preferably, the process of constructing the graph structure includes: The one-tuple node of each image belongs to the feature space of the corresponding image; the feature spaces of each image are fused to obtain the feature space to which the three-tuple node belongs; Arrange the unigram nodes around the triplet nodes, and connect each unigram node with the triplet node to obtain a plurality of first edges; The feature space of the binary feature includes the feature space after the feature spaces of any two images are fused, and the feature space after the feature spaces of any two images are compared with the global features; the fused feature space is used as the fusion weight; Taking the average of the three obtained fusion weights to obtain a weight average, and using the weight average as the weight of the first edge connecting each unigram node and the triplet node; Use the weight adjustment function to adjust the weights of the feature spaces after the comparison of the three global features, and obtain the weights of the second edges corresponding to the connection between the two unigram nodes; The graph structure is constructed based on the triplet nodes, each tuple node, a plurality of first edges and weights of the first edges, a plurality of second edges and weights of each second edge.

[0011] Preferably, the step of fusing the graph structure using a graph convolutional neural network includes: Construct the Laplacian matrix of the graph structure based on the degree matrix and adjacency matrix of the graph structure; Through each GCN layer in the graph convolutional neural network, each node in the Laplacian matrix is ​​updated, and nonlinear activation is performed after each GCN layer to obtain the representation of each node after the corresponding GCN layer is activated; Aggregate the representations of all nodes after the activation of the last GCN layer to obtain the global feature vector of the graph structure; The global feature vector of the graph structure is mapped to the fusion result through mapping weights.

[0012] Preferably, updating and optimizing the graph structure based on the fusion result includes: Analyze the fusion result to obtain an update vector, where the update vector includes nodes that need to be updated and their update directions; Adjusting a node weight vector of each node in the graph structure based on the update vector; According to the adjusted node weight vector of each node and the weight of the first edge / second edge, the UpdateWeights function is used to update the weight of the corresponding edge; According to the adjusted node weight vectors of each node and the updated weights of each edge, the AdjustTopology function is used to adjust the topological structure of the graph structure to obtain the adjusted graph structure; The OptimizeGraph function is used to optimize the adjusted graph structure to obtain the optimized graph structure.

[0013] Preferably, the unmanned equipment includes drones and autonomous driving vehicles.

[0014] Beneficial effects: By constructing a new multi-group information system, this method can better explore and utilize the potential connections between different modal data, improve the accuracy and efficiency of information processing, and thus enhance the perception and decision-making support capabilities of unmanned equipment in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0016] Figure 1 This is a flow chart of a method for constructing a multimodal information system for multi-group unmanned equipment in an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present application, so the present application is not limited by the specific embodiments disclosed below.

[0018] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0019] like Figure 1 As shown, this embodiment provides a method for constructing a multi-modal information system of a multi-group unmanned equipment, the method comprising: S1: Receive multimodal information collected by unmanned equipment, where the multimodal information includes RGB images, lidar point cloud images, and infrared images.

[0020] In this embodiment, the unmanned equipment includes but is not limited to drones and autonomous driving vehicles.

[0021] S2: Extract features from the three images respectively, obtain features of the three images respectively, and use the features of the three images as tuple nodes of the corresponding images respectively.

[0022] Specifically, extracting features from the three images respectively includes: A convolutional neural network is used to extract global features of the RGB image to obtain a first image global feature, and the first image global feature is used as a tuple node of the RGB image.

[0023] The extraction process includes: Step 1: Assume the input RGB image is a three-dimensional tensor , whose dimensions are H×W×C, where H is the image height, W is the image width, and C=3 is the number of color channels; Step 2: The convolution layer consists of multiple convolution kernels K i Each convolution kernel slides on the RGB image to calculate local features. Assuming the size of each convolution kernel is k×k, the convolution operation formula is expressed as: ; in, Represents the value of the i-th channel of the feature map at position (j, j); represents an RGB image, Represents the pixel value of the jth channel of the RGB image at position (m+1,n+1); Represents the weight value of the i-th convolution kernel at position (m,n); Step 3: Apply the nonlinear activation function ReLU to the output of each convolutional layer for nonlinear activation. The calculation formula is: ; in, Represents the activated feature map; Step 4: Reduce the spatial dimension of the feature map through the pooling layer (usually maximum pooling or average pooling), reduce the number of parameters and the amount of calculation. The pooling operation is expressed as: ; in, Represents the value of the pooled feature map at the i-th channel and position j; Represents the value at coordinate (m, n) in the pooling window on the i-th channel of the feature map before pooling; Indicates taking the maximum value within the pooling window; Step 5: After multiple convolutional and pooling layers, the feature map will be flattened and passed to the fully connected layer. Let the weight matrix of the fully connected layer be W f , the bias vector is b f , then the full connection operation is expressed as: ; in, represents the output of the pooling layer after the activation function, and Z represents the fully connected layer; Step 6: In order to extract global features, global average pooling is used after the fully connected layer to obtain the global features of the first image. The calculation formula is: ; in, represents the first image global feature; Represents the size of the feature map after convolution and pooling operations, Represents the value of the pooling layer output at the (i, j, k) position after being processed by the activation function.

[0024] The encoder is used to extract multi-scale features from the RGB image to obtain multi-scale features of the first image, and the calculation formula is: ; in, A first image multi-scale feature representing a first scale; Representing multi-scale features of the first image at a second scale; A first image multi-scale feature representing a third scale; Represents the encoder for RGB images; Represents an RGB image.

[0025] PointNet is used to extract global features of the laser radar point cloud image to obtain a second image global feature, and the second image global feature is used as a tuple node of the laser radar point cloud image.

[0026] The extraction process includes: Step 1: Set the input point cloud P Depend on N points, each point Represents a point in three-dimensional space; ; Step 2: For each point Apply a learnable multi-layer perceptron to extract local features, calculated as: ; in, Represents the local features of the i-th point; represents a multilayer perceptron; Step 3: Use Transformation Network To learn a differentiable transformation to eliminate the effect of point arrangement in the point cloud, it is expressed as: ; in, represents a first network parameter of the transformation network; A second network parameter representing the transformation network; Step 4: Transform the features Aggregate through a symmetric function. In PointNet, the maximum pooling is used as the symmetric function. The calculation formula is: ; in, Represents the features after aggregation; Step 5: Use MLP to extract global features from the aggregated features again to obtain the global features of the second image. The calculation formula is: ; in, Represents the second image global feature.

[0027] ResNet is used to extract global features of the infrared image to obtain a third image global feature, and the third image global feature is used as a tuple node of the infrared image.

[0028] The extraction process includes: Step 1: Assume input infrared image is a three-dimensional tensor with dimensions H×W×C, where H is the image height, W is the image width, and C is the number of color channels (for infrared images, C=1); Step 2: In the first convolution layer, multiple convolution kernels K are used to perform convolution operations on the infrared image to extract local features. The calculation formula is: ; in, Represents local features; represents the bias term; Step 3: Apply a nonlinear activation function (ReLU function) to the output of the convolutional layer. The calculation formula is: ; in, Represents the local features after activation; Step 4: The residual block is the core component of ResNet. It contains two convolutional layers. The output of the residual block is the sum of the output of the second convolutional layer, the output of the first convolutional layer, and the bias term in the residual block. The calculation formula is: ; in, represents the output of the residual block; Represents the output of the first convolutional layer; represents the output of the second convolutional layer; Represents the bias term of the residual block; Step 5: Use a pooling layer (such as max pooling) to reduce the spatial dimension of the output of the residual block. The calculation formula is: ; in, Represents the value at coordinate (m, n) in the pooling window on the i-th channel of the output of the residual block before pooling, Represents the value of the output of the pooled residual block at the i-th channel and position j; Step 6: At the end of the network, use global average pooling to aggregate all features and extract global features. The calculation formula is: ; in, Represents the global features; D represents the depth of the feature map after multiple convolution and pooling layers; Represents the value at position (i, j, k) after the pooling layer output; Step 7: Pass the one-dimensional feature vector after global average pooling through the fully connected layer to obtain the global features of the third image for classification or other tasks. The calculation formula is: ; in, represents the third image global feature; Represents the one-dimensional feature vector after global average pooling; Represents the weight of the fully connected layer in ResNet; Represents the bias of the fully connected layer in ResNet.

[0029] The encoder is used to extract multi-scale features from the infrared image to obtain multi-scale features of the third image. The calculation formula is: ; in, A third image multi-scale feature representing the first scale; A third image multi-scale feature representing the second scale; a third image multi-scale feature representing a third scale; represents the encoder for infrared images; Represents an infrared image.

[0030] S3: Pair the features of the three images in pairs and fuse them to obtain three shared semantic subspaces, and use the three shared semantic subspaces as different bigram features; fuse all the features of the three images to obtain a shared semantic space, and use the shared semantic space as a triplet node.

[0031] Specifically, the fusion process of the features of the first image and the features of the second image includes: The bilinear interpolation method is used to adjust the global features of the first image and the global features of the second image to the same dimension, which can be expressed as: ; ; in, represents the adjusted global features of the first image; represents the adjusted global features of the second image; represents the bilinear interpolation function; The expansion of the bilinear interpolation function is: Given a feature map F, Represents the value of the feature map at position m,n, and it is expected to interpolate F from size M×N to size , where s is the scaling factor used to determine the target size; For each position i, j in the feature map F, the calculation formula for the bilinear interpolation is: ; in, Represents the floor function; mod represents the modulus operation; this formula estimates the value of the target position (i, j) by calculating the weighted sum of the four nearest neighboring points in the feature map F, where the weight is determined by the distance from each neighboring point to the target position, thereby achieving a smooth interpolation effect.

[0032] The aligned first image global features and the second image global features are mapped to the first feature space through the first fully connected layer, which can be expressed as: ; in, represents the first feature space, represents the first fully connected layer; In the first feature space, the aligned global features of the first image and the global features of the second image are weighted summed to obtain the first fusion feature, which is calculated as follows: ; ; in, represents the first fusion feature; represents the corresponding weight of the adjusted global feature of the first image; represents the corresponding weight of the adjusted global feature of the second image; Perform a nonlinear transformation on the first fusion feature (using the sigmoid activation function) to obtain the first activation feature ; A feature pyramid network is used to extract and fuse multi-scale features of the first activation feature to obtain a first multi-scale fusion feature. , which helps detect objects of different sizes; The first multi-scale fusion feature is mapped to the first shared semantic subspace through the second fully connected layer , which can be expressed as: ; in, represents the second fully connected layer.

[0033] At the same time, a loss function is defined to train the network, which usually includes position loss, classification loss and confidence loss. The first loss function is calculated as: ; in, represents the first loss function; represents the first balance coefficient; represents the second balance coefficient; represents the third balance coefficient; Use the first loss function to back-propagate the feature pyramid network, and use an optimization algorithm (such as SGD, Adam, and other gradient-based optimization algorithms) to update the network weights; During the training process, data augmentation techniques can also be used to improve the generalization ability of the model.

[0034] Through the above steps, the first shared semantic subspace finally output contains the fused features from the RGB image and the lidar point cloud image, which can provide richer information and enhance the performance of target detection and other related tasks.

[0035] The fusion process of the features of the first image and the features of the third image includes: The multi-scale features of the first image at each scale are convolved by a 3×3 convolution operation to obtain the first convolution feature, which is calculated as follows: ; in, Indicates i The first convolutional feature of the scale; Indicates that the 3×3 convolution operation is used to extract deep features; Indicates i First image multi-scale features of scales; The third image multi-scale features of each scale are convolved respectively through a 3×3 convolution operation to obtain a second convolution feature, and the calculation formula is: ; in, Indicates i The second convolution feature of the scale; Indicatesi Multi-scale features of the third image at scales; The first convolution feature and the second convolution feature are added to obtain a common feature, which is calculated as follows: ; in, Indicates i Common features of the scales; Subtract the second convolution feature from the first convolution feature and divide it by 2 to get the private feature of the RGB image. The calculation formula is: ; in, Indicates i Private features of RGB images at different scales; Subtract the first convolution feature from the second convolution feature and divide it by 2 to obtain the private feature of the infrared image. The calculation formula is: ; in, Indicates i Private features of infrared images at different scales; The ECA attention mechanism is used to enhance the common features, private features of RGB images, and private features of infrared images respectively. The calculation formula is: ; ; ; in, Indicates that after enhancement i Common features of the scales; Indicates that after enhancement i Private features of RGB images at different scales; Indicates that after enhancement i Private features of infrared images at different scales; Represents the ECA attention mechanism; The enhanced common features, the private features of the RGB image, and the private features of the infrared image are fused at each scale to obtain the second fused features at each scale. The calculation formula is: ; in, Indicates i The second fusion feature of the scale; The second fusion features of all scales are reconstructed through the decoder to obtain the second multi-scale fusion features , the calculation formula is: ; in, represents a decoder; , , Respectively represent the second fusion features of the 1st, 2nd and 3rd scales; the second multi-scale fusion features are used as the second shared semantic subspace.

[0036] At the same time, an information-aware loss function is constructed to guide network training. The information-aware loss function is expressed as: ; in, represents the information-aware loss function; Indicates loss of strength; represents the gradient loss; represents pixel loss based on information perception; express The weight of express The weight of express The weight of An information-aware loss function is minimized using an appropriate optimization algorithm, such as the Adam optimizer, to train the network parameters.

[0037] The fusion process of the features of the second image and the features of the third image includes: The same spatial transformation is performed on the second image global feature and the third image global feature, which can be expressed as: ; ; in, represents the transformed global features of the second image; represents the global features of the transformed third image; Represents a spatial transformation operation; In this embodiment, the spatial transformation is an affine transformation, and the transformation process is as follows: First, define the affine transformation matrix. The affine transformation matrix T is a 2×3 matrix used to represent transformation operations such as rotation, translation, and scaling. The affine transformation matrix is ​​defined as: ; in, a , b , c , d , e , f These are parameters to be determined; Taking the global features of the second image as an example, each feature point (x, y) is transformed, and the transformed coordinates are (x', y'). The transformation formula is: ; in, Represents a linear transformation matrix, which is used to implement transformations such as rotation, scaling, and shearing; Represents the translation vector, which is used to achieve image translation; Transforming each feature point in the global feature of the third image by using the above transformation formula; Then, at least three groups of corresponding points are needed to determine the above six parameters to be determined. Specifically, the feature point matching algorithm SIFT is used to find the matching feature point pair (x1, y1)-(x2, y2) in the global features of the second image and the global features of the third image. Substituting these matched feature point pairs into the above transformation formula, we get a set of equations. Solving this set of equations can determine the values ​​of the six parameters. After determining the affine transformation matrix T, each feature point in the second image global feature and the third image global feature is transformed using a transformation formula that determines a parameter value, and the second image global feature and the third image global feature are transformed into the same space.

[0038] The second image global feature after spatial transformation is fused with the third image global feature to obtain a third fused feature, and the calculation formula is: ; in, represents the third fusion feature; Represents the fusion operation, which can be simple data concatenation, weighted average or other complex fusion strategies; The third fusion feature is optimized by using data fitting term, regularization term and argmin function to obtain the optimized feature, and the calculation formula is: ; in, represents the optimization feature; represents the data fitting term; represents the regularization term; Use multi-scale analysis to extract optimized features of different scales and fuse them to obtain the third multi-scale fusion feature. The calculation formula is: ; in, represents the third multi-scale fusion feature; It is a union symbol, which is used to extract features at different scales. The features after merging are performed; Indicates i Feature extraction operations at scales; In this embodiment, a Gaussian difference pyramid is used for multi-scale analysis, and the multi-scale analysis process is as follows: Constructing Gaussian pyramid: Convolve the input optimized features with Gaussian kernels of different standard deviations to obtain a series of images with different blur levels to form a Gaussian pyramid. Let the original optimized features be I(X, Y), then the k-th layer Gaussian pyramid image G k The calculation formula for (x,y) is: ; in, represents the Gaussian kernel function; represents standard deviation; Construct a Gaussian difference pyramid: By subtracting two adjacent layers of Gaussian pyramid images, we can obtain a Gaussian difference pyramid. The calculation formula is: ; in, represents the Gaussian difference pyramid image of the kth layer; represents the k+1th Gaussian pyramid image; Extract optimized features of different scales: For each layer of the Gaussian difference pyramid, extract corresponding features (such as edge features, texture features, etc.) according to specific needs; Fusion of features at different scales: Fusion of the extracted optimized features at different scales may be performed by weighted averaging to obtain the third multi-scale fusion feature.

[0039] Mapping the third multi-scale fusion feature to the third shared semantic subspace through an artificial neural network , which can be expressed as: ; in, Represents a neural network.

[0040] At the same time, a second loss function is defined to train the neural network to ensure that the third shared semantic subspace can accurately reflect the information of the input features. The expression of the second loss function is: ; in, represents the second loss function; A semantic subspace representation of the ground truth, represents the regularization parameter; Indicates the complexity of the neural network; The parameters of the neural network are trained by minimizing the second loss function using an appropriate optimization algorithm such as gradient descent.

[0041] The method of fusing all the features of the three images to obtain a shared semantic space includes: Adjusting the second image global feature and the third image global feature to the same spatial reference system as the first image global feature can be expressed as: ; ; in, represents the adjusted global features of the second image; represents the adjusted global features of the third image; represents a transformation function for adjusting the global features of the second image to the same reference frame as the global features of the first image; The transformation function that adjusts the global features of the third image to the same reference system as the global features of the first image; here the two transformation functions , They are all affine transformations, and their specific transformation process is as shown in the above spatial transformation. The only difference is that they target different objects.

[0042] The first image global feature is aligned with the adjusted second image global feature and the third image global feature, which is expressed as: ; in, represents the global features of the first image after alignment; represents the feature alignment function; The weighted average of the first image global features, the second image global features and the third image global features of the feature alignment is performed to obtain the fourth fusion feature, which is calculated as follows: ; ; in, Indicates the fourth fusion feature; express The weight coefficient of express The weight coefficient of express The weight coefficient of The fourth fusion feature is mapped to the second activation feature through the third fully connected layer ; When the feature has multi-scale characteristics, multi-scale analysis is used to extract the second activation features of different scales and fuse them to obtain the fourth multi-scale fusion feature. The calculation formula is: ; in, represents the fourth multi-scale fusion feature; Indicates i The multi-scale analysis also uses Gaussian difference pyramid for multi-scale analysis, and the multi-scale analysis process is consistent with the multi-scale analysis process of the above-mentioned optimization features, the only difference is that the objects they target are different.

[0043] The fourth multi-scale fusion feature is optimized by using a data fitting term, a regularization term, and an argmin function to obtain the shared semantic space. The calculation formula is: ; in, Represents a shared semantic space; represents the data fitting term; represents the regularization term.

[0044] At the same time, define the third loss function to train the third fully connected layer , to ensure a shared semantic space It can accurately reflect the information of the input features. The expression of the third loss function is: ; in, represents the third loss function; A feature representation representing the ground truth, represents the regularization parameter; represents the complexity of the third fully connected layer; Use an appropriate optimization algorithm (such as Adam optimizer) to minimize the third loss function and train the parameters of the third fully connected layer.

[0045] S4: Connect each tuple node to each other, and connect each tuple node to the triple node, calculate the weight of each edge based on the feature space of the two-tuple feature, and build a graph structure.

[0046] Specifically, the one-tuple node of each image belongs to the feature space of the corresponding image, which is expressed as: , , ; , , Represent the feature spaces of RGB images, LiDAR point cloud images, and infrared images respectively; The feature spaces of each image are fused to obtain the feature space to which the triplet node belongs , expressed as: ; Arrange the unigram nodes around the triplet nodes, and connect each unigram node with the triplet node to obtain a plurality of first edges; The feature space of the binary feature includes the feature space after the feature spaces of any two images are fused, and the feature space after the feature spaces of any two images are compared with the global features; the fused feature space is used as the fusion weight; The three obtained fusion weights are averaged to obtain the weight mean, and the weight mean is used as the weight of the first edge connecting each unigram node and the triplet node; the calculation formula is: ; ; ; ; in, represents the weighted mean; Represents the fusion weight of the feature space of the RGB image and the feature space of the LiDAR point cloud image; Represents the fusion weight of the feature space of the RGB image and the feature space of the infrared image; Represents the fusion weight of the feature space of the lidar point cloud image and the feature space of the infrared image; represents the fusion function; The weight adjustment function is used to adjust the weights of the feature spaces after the comparison of the three global features, and the weights of the second edges corresponding to the connection between the two unigram nodes are obtained respectively; it is expressed as: ; ; ; ; ; ; in, Represents the feature space after the first global feature comparison; Represents the feature space after the second global feature comparison; Represents the feature space after the third global feature comparison; represents the global feature comparison function; A one-tuple node representing an RGB image A tuple node with a LiDAR point cloud image The weight of the second edge of the connection; A one-tuple node representing an RGB image A tuple node with an infrared image The weight of the second edge of the connection; A one-tuple node representing a LiDAR point cloud image A tuple node with an infrared image The weight of the second edge of the connection; represents the weight adjustment function; Based on the triple node, each tuple node, multiple first edges and the weights of the first edges, multiple second edges and the weights of each second edge, the graph structure is constructed, and the graph structure is expressed as: G ={ V , E , W}; ; ; .

[0047] S5: Use a graph convolutional neural network to fuse the graph structure to obtain a fusion result.

[0048] Specifically, the adopting of a graph convolutional neural network to fuse the graph structure includes: According to the degree matrix and adjacency matrix of the graph structure, the Laplacian matrix of the graph structure is constructed, which can be expressed as: L = DA; Among them, L represents the Laplacian matrix of the graph structure; D represents the degree matrix of the graph structure; A represents the adjacency matrix of the graph structure; Through each GCN layer in the graph convolutional neural network, each node in the Laplacian matrix is ​​updated. l The update formula of the GCN layer is: ; in, Indicates l The first in the Laplacian matrix after the GCN layer update i nodes; represents the GCN layer; Indicates l The weight of the GCN layer; Indicates l The weight from the jth neuron to the ith neuron in the GCN layer, Indicates l The set of neurons in the previous layer connected to the i-th neuron in the GCN layer, Indicates l The output of the jth neuron in the GCN layer; After each GCN layer, nonlinear activation is performed to obtain the representation of each node after the activation of the corresponding GCN layer. The calculation formula is: ; in, Indicates l -1 layer of GCN layer after updating the Laplacian matrix i nodes; ReLU activation function. Aggregate the representations of all nodes after the activation of the last GCN layer to obtain the global feature vector of the graph structure, which is calculated as: ; in, Represents the global feature vector of the graph structure; V represents the number of nodes in the graph structure; Indicates L The first in the Laplacian matrix after the GCN layer is updated and activated i nodes; The global feature vector of the graph structure is mapped to the fusion result through mapping weights, which is expressed as: ; in, represents the fusion result; Represents the mapping weight.

[0049] At the same time, the fourth loss function for training GCN is defined, which includes the loss term of the task target (such as classification loss or regression loss) and the second regularization term. The fourth loss function expression is: ; in, represents the fourth loss function; represents the classification loss; represents the regularization parameter; represents the regression loss; Represents the weight of GCN; Use an optimization algorithm (such as Adam) to minimize the fourth loss function and update the network weights: ; in, Represents the network weight of the updated GCN; Represents the network weight of GCN before updating; represents partial derivative; Represents the learning rate.

[0050] S6: Based on the fusion result, the graph structure is updated and optimized to obtain an optimized graph structure.

[0051] Specifically, updating and optimizing the graph structure based on the fusion result includes: The fusion result is analyzed to obtain an update vector, which includes the nodes to be updated and their update direction (enhancement or weakening), and the calculation formula is: ; in, Represents an analytical function; represents the update vector; The node weight vector of each node in the graph structure is adjusted based on the update vector, and the calculation formula is: ; in, Represents the node weight vector of each node after adjustment; Represents the original weight vector of each node; Represents the initial weight matrix of each node; represents the Hadamard product; According to the adjusted node weight vector of each node and the weight of the first edge / second edge, the UpdateWeights function is used to update the weight of the corresponding edge, which is expressed as: ; in, Represents the weight of each edge after update; Represents the weight of each original edge; According to the adjusted node weight vectors of each node and the updated weights of each edge, the AdjustTopology function is used to adjust the topological structure of the graph structure (such as adding or deleting nodes and weighted edges) to obtain the adjusted graph structure, which is expressed as: ; in, represents the adjusted graph structure, is the set of nodes in the adjusted graph structure, is the set of weighted edges in the adjusted graph structure; The OptimizeGraph function is used to optimize the adjusted graph structure to obtain the optimized graph structure, which is expressed as: ; in, represents the optimized graph structure, is the set of nodes in the optimized graph structure, is the set of weighted edges in the optimized graph structure.

[0052] The method for constructing a multi-modal information system for multi-tuple unmanned equipment provided in this embodiment aims to construct a new multi-tuple information system in a more sophisticated and systematic way, which can better explore and utilize the potential connections between different modal data, improve the accuracy and efficiency of information processing, and thus enhance the perception ability and decision support ability of unmanned equipment in complex environments. This method is expected to promote the advancement of information processing technology for unmanned equipment and provide support for achieving a higher level of autonomy and intelligence.

[0053] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0054] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the patent application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent application shall be subject to the attached claims.

Claims

1. A method for constructing a multimodal information system of multi-group unmanned equipment, characterized in that: include: S1: receiving multimodal information collected by unmanned equipment, wherein the multimodal information includes RGB images, laser radar point cloud images, and infrared images; S2: extract features from the three images respectively, obtain features of the three images respectively, and use the features of the three images as tuple nodes of the corresponding images respectively; S3: Pair the features of the three images in pairs and fuse them to obtain three shared semantic subspaces, and use the three shared semantic subspaces as different bigram features; fuse all the features of the three images to obtain a shared semantic space, and use the shared semantic space as a triplet node; S4: Connect each tuple node to each other, and connect each tuple node to the triple node, calculate the weight of each edge based on the feature space of the two-tuple feature, and build a graph structure; S5: using a graph convolutional neural network to fuse the graph structure to obtain a fusion result; S6: Based on the fusion result, the graph structure is updated and optimized to obtain an optimized graph structure.

2. The method for constructing a multimodal information system of multi-group unmanned equipment according to claim 1, characterized in that: Using a convolutional neural network to extract global features from the RGB image to obtain a first image global feature, and using the first image global feature as a tuple node of the RGB image; Using an encoder to perform multi-scale feature extraction on the RGB image to obtain multi-scale features of a first image; Using PointNet to extract global features from the laser radar point cloud image to obtain a second image global feature, and using the second image global feature as a tuple node of the laser radar point cloud image; Using ResNet to extract global features of the infrared image to obtain a third image global feature, and using the third image global feature as a tuple node of the infrared image; An encoder is used to extract multi-scale features from the infrared image to obtain multi-scale features of a third image.

3. The method for constructing a multi-modal information system of multi-group unmanned equipment according to claim 2, characterized in that: The process of fusing the features of the first image with the features of the second image includes: Using a bilinear interpolation method to adjust the global features of the first image and the global features of the second image to the same dimension; Mapping the aligned first image global features and the second image global features to the first feature space through a first fully connected layer; Performing a weighted summation of the aligned first image global features and the second image global features in the first feature space to obtain a first fusion feature; Performing a nonlinear transformation on the first fused feature to obtain a first activated feature; Using a feature pyramid network to perform multi-scale feature extraction and fusion on the first activation feature to obtain a first multi-scale fusion feature; The first multi-scale fusion features are mapped to a first shared semantic subspace through a second fully connected layer.

4. The method for constructing a multimodal information system of multi-group unmanned equipment according to claim 2, characterized in that: The fusion process of the features of the first image and the features of the third image includes: Convolving the multi-scale features of the first image at each scale respectively through a 3×3 convolution operation to obtain a first convolution feature; Convolving the third image multi-scale features of each scale respectively through a 3×3 convolution operation to obtain a second convolution feature; Adding the first convolution feature and the second convolution feature to obtain a common feature; Subtract the second convolution feature from the first convolution feature, and divide the result by 2 to obtain a private feature of the RGB image; Subtract the first convolution feature from the second convolution feature, and divide the result by 2 to obtain a private feature of the infrared image; The ECA attention mechanism is used to enhance the common features, private features of RGB images, and private features of infrared images respectively; The enhanced common features, the private features of the RGB image, and the private features of the infrared image are fused at each scale to obtain the second fused features at each scale; The second fusion features of all scales are reconstructed through a decoder to obtain second multi-scale fusion features; and the second multi-scale fusion features are used as the second shared semantic subspace.

5. The method for constructing a multi-modal information system of multi-group unmanned equipment according to claim 2, characterized in that: The fusion process of the features of the second image and the features of the third image includes: Performing the same spatial transformation on the second image global feature and the third image global feature; Fusing the second image global feature after spatial transformation with the third image global feature to obtain a third fused feature; The third fusion feature is optimized by using a data fitting term, a regularization term and an argmin function to obtain an optimized feature; Use multi-scale analysis to extract optimized features of different scales and fuse them to obtain the third multi-scale fusion feature; The third multi-scale fusion feature is mapped to a third shared semantic subspace through an artificial neural network.

6. The method for constructing a multi-modal information system of multi-group unmanned equipment according to claim 2, characterized in that: The method of fusing all the features of the three images to obtain a shared semantic space includes: Adjusting the second image global feature and the third image global feature to the same spatial reference system as the first image global feature; Perform feature alignment on the first image global feature, the adjusted second image global feature, and the third image global feature; Perform weighted averaging of the first image global features, the second image global features, and the third image global features that are feature aligned to obtain a fourth fusion feature; Mapping the fourth fused feature into a second activation feature through a third fully connected layer; Use multi-scale analysis to extract second activation features of different scales and fuse them to obtain a fourth multi-scale fusion feature; The fourth multi-scale fusion feature is optimized by using a data fitting term, a regularization term, and an argmin function to obtain the shared semantic space.

7. The method for constructing a multi-modal information system of a multi-group unmanned equipment according to claim 2, characterized in that: The process of building a graph structure includes: The one-tuple node of each image belongs to the feature space of the corresponding image; the feature spaces of each image are fused to obtain the feature space to which the three-tuple node belongs; Arrange the unigram nodes around the triplet nodes, and connect each unigram node with the triplet node to obtain a plurality of first edges; The feature space of the binary feature includes the feature space after the feature spaces of any two images are fused, and the feature space after the feature spaces of any two images are compared with the global features; the fused feature space is used as the fusion weight; Taking the average of the three obtained fusion weights to obtain a weight average, and using the weight average as the weight of the first edge connecting each unigram node and the triplet node; Use the weight adjustment function to adjust the weights of the feature spaces after the comparison of the three global features, and obtain the weights of the second edges corresponding to the connection between the two unigram nodes; The graph structure is constructed based on the triplet nodes, each tuple node, a plurality of first edges and weights of the first edges, a plurality of second edges and weights of each second edge.

8. The method for constructing a multimodal information system of multi-group unmanned equipment according to claim 1, characterized in that: The adopting of a graph convolutional neural network to fuse the graph structure includes: Construct the Laplacian matrix of the graph structure based on the degree matrix and adjacency matrix of the graph structure; Through each GCN layer in the graph convolutional neural network, each node in the Laplacian matrix is ​​updated, and nonlinear activation is performed after each GCN layer to obtain the representation of each node after the corresponding GCN layer is activated; Aggregate the representations of all nodes after the activation of the last GCN layer to obtain the global feature vector of the graph structure; The global feature vector of the graph structure is mapped to the fusion result through mapping weights.

9. The method for constructing a multi-modal information system of multi-group unmanned equipment according to claim 1, characterized in that: The updating and optimizing of the graph structure based on the fusion result includes: Analyze the fusion result to obtain an update vector, where the update vector includes nodes that need to be updated and their update directions; Adjusting a node weight vector of each node in the graph structure based on the update vector; According to the adjusted node weight vector of each node and the weight of the first edge / second edge, the UpdateWeights function is used to update the weight of the corresponding edge; According to the adjusted node weight vectors of each node and the updated weights of each edge, the AdjustTopology function is used to adjust the topological structure of the graph structure to obtain the adjusted graph structure; The OptimizeGraph function is used to optimize the adjusted graph structure to obtain the optimized graph structure.

10. The method for constructing a multi-modal information system of multi-group unmanned equipment according to claim 1, characterized in that: The unmanned equipment includes drones and autonomous driving vehicles.

Citation Information

Patent Citations

  • Cross-scale graph similarity guide aggregation system, method and application

    CN115880552A

  • Truck intelligent driving sensing method based on multi-sensor fusion detection under cross-modal supervised learning

    CN117237919A

  • Sensor data fusion method and device based on graph neural network, and storage medium

    CN118364432A

  • Room obstacle target detection method and system based on multi-modal information

    CN119048747A

  • Unmanned aerial vehicle multi-modal data fusion method based on conversion graph neural network

    CN119251635A

Cited By

  • Target matching method based on cross-domain multi-modal fusion coding

    CN120277626A

  • A target matching method based on cross-domain multimodal fusion coding

    CN120277626B