A Method for Constructing a Multimodal Information System of Multivariate Unmanned Equipment

By building a multi-modal information system for unmanned equipment, and using graph convolutional neural network to fuse the characteristics of RGB images, lidar point cloud images and infrared images, the problems of insufficient information fusion and mode correlation in traditional technologies are solved, and more efficient information processing and stronger perceived decision-making capabilities are achieved.

CN119942292BActive Publication Date: 2025-06-27XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510448376.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-27
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

Traditional multimodal information fusion technology fails to fully utilize and integrate the intrinsic connections and complementarity between different modal data, resulting in the loss of some useful information and lacks in-depth consideration of the correlation between different modalities, affecting the execution of cross-modal tasks.

Method used

A multimodal information system construction method is proposed. By receiving RGB images, lidar point cloud images and infrared images, features are extracted and graph structures are constructed, and graph convolutional neural network is used for information fusion and optimization to form a shared semantic space to enhance the accuracy and efficiency of information processing.

Benefits of technology

By building a new multi-group information system, it can better explore and utilize potential connections between different modal data, improve the accuracy and efficiency of information processing, and improve the perception and decision-making support capabilities of unmanned equipment in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942292B_ABST
    Figure CN119942292B_ABST
Patent Text Reader

Abstract

The present application relates to a method for constructing a multi-modal information system of a multi-agent unmanned equipment, which receives multi-modal information collected by the unmanned equipment; extracts features from three types of images respectively to obtain the features of the three types of images, and uses the features of the three types of images as the single-tuple nodes corresponding to the images respectively; pairs and fuses the features of the three types of images in pairs to obtain three shared semantic sub-spaces, and uses the three shared semantic sub-spaces as different binary-tuple features respectively; fuses all the features of the three types of images to obtain a shared semantic space, and uses the shared semantic space as a triple-tuple node; connects each single-tuple node to each other, and connects each single-tuple node to the triple-tuple node, calculates the weights of each edge based on the feature space of the binary-tuple features, and constructs a graph structure; uses a graph convolutional neural network to fuse the graph structure to obtain a fusion result; updates and optimizes the graph structure based on the fusion result to obtain an optimized graph structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of unmanned equipment information processing, and particularly to a method for constructing a multi-modal information system for multi-element unmanned equipment. Background Art

[0002] Currently, in the field of unmanned equipment information processing, information processing technology is developing towards a more intelligent and automated direction. Unmanned equipment needs to process and analyze a large amount of data to perform complex tasks. These data usually include RGB image data, lidar point cloud data, and infrared image data, which provide different perspectives and information levels respectively; among them, RGB image data provides rich color and texture information, lidar point cloud data provides accurate spatial structure information, and infrared image data provides temperature and heat distribution information.

[0003] However, traditional multi-modal information fusion technologies have some limitations. These technologies often fail to fully utilize and integrate the internal connections and complementarities between different modal data when processing them, resulting in the loss of some useful information; in addition, due to the relative independence of each modal information, the correlation between them is poor, and the in-depth consideration of the correlation between different modalities is lacking, which is particularly obvious when performing cross-modal tasks. Summary of the Invention

[0004] Based on this, it is necessary to provide a method for constructing a multi-modal information system for multi-element unmanned equipment, which includes:

[0005] S1: Receive the multi-modal information collected by the unmanned equipment, where the multi-modal information includes RGB images, lidar point cloud images, and infrared images;

[0006] S2: Extract features from the three types of images respectively to obtain the features of the three types of images, and use the features of the three types of images as the single-element nodes corresponding to the respective images;

[0007] S3: Pair and fuse the features of the three types of images in pairs to obtain three shared semantic sub-spaces, and use the three shared semantic sub-spaces as different two-element features respectively; fuse all the features of the three types of images to obtain a shared semantic space, and use the shared semantic space as the three-element node;

[0008] S4: Connect the single-element nodes to each other, and connect each single-element node to the three-element node, calculate the weights of each edge based on the feature space of the two-element features, and construct a graph structure;

[0009] S5: Use a graph convolutional neural network to fuse the graph structure to obtain a fusion result;

[0010] S6: Update and optimize the graph structure based on the fusion result to obtain an optimized graph structure.

[0011] Preferably, a convolutional neural network is used to extract global features of the RGB image to obtain a first global image feature, and the first global image feature is used as a unary node of the RGB image;

[0012] An encoder is used to extract multi-scale features of the RGB image to obtain a first multi-scale image feature;

[0013] PointNet is used to extract global features of the lidar point cloud image to obtain a second global image feature, and the second global image feature is used as a unary node of the lidar point cloud image;

[0014] ResNet is used to extract global features of the infrared image to obtain a third global image feature, and the third global image feature is used as a unary node of the infrared image;

[0015] An encoder is used to extract multi-scale features of the infrared image to obtain a third multi-scale image feature.

[0016] Preferably, the fusion process of the features of the first image and the features of the second image includes:

[0017] The first global image feature and the second global image feature are adjusted to the same dimension by using bilinear interpolation;

[0018] The aligned first global image feature and second global image feature are mapped to a first feature space through a first fully connected layer;

[0019] The aligned first global image feature and second global image feature are weighted and summed in the first feature space to obtain a first fusion feature;

[0020] The first fusion feature is non-linearly transformed to obtain a first activation feature;

[0021] A feature pyramid network is used to extract and fuse multi-scale features of the first activation feature to obtain a first multi-scale fusion feature;

[0022] The first multi-scale fusion feature is mapped to a first shared semantic subspace through a second fully connected layer.

[0023] Preferably, the fusion process of the features of the first image and the features of the third image includes:

[0024] Each scale of the first multi-scale image feature is convolved through a 3×3 convolution operation to obtain a first convolution feature;

[0025] Perform convolution on the multi-scale features of the third image at each scale respectively through a 3×3 convolution operation to obtain second convolution features;

[0026] Add the first convolution features and the second convolution features to obtain common features;

[0027] Subtract the second convolution features from the first convolution features, and then divide by 2 to obtain the private features of the RGB image;

[0028] Subtract the first convolution features from the second convolution features, and then divide by 2 to obtain the private features of the infrared image;

[0029] Adopt the ECA attention mechanism to perform feature enhancement on the common features, the private features of the RGB image, and the private features of the infrared image respectively;

[0030] Fuse the enhanced common features, the private features of the RGB image, and the private features of the infrared image at each scale respectively to obtain second fusion features at each scale;

[0031] Reconstruct all the second fusion features at all scales through a decoder to obtain second multi-scale fusion features; use the second multi-scale fusion features as the second shared semantic subspace.

[0032] Preferably, the fusion process of the features of the second image and the features of the third image includes:

[0033] Perform the same spatial transformation on the global features of the second image and the global features of the third image;

[0034] Fuse the spatially transformed global features of the second image and the global features of the third image to obtain third fusion features;

[0035] Optimize the third fusion features by using a data fitting term, a regularization term, and an argmin function to obtain optimized features;

[0036] Use multi-scale analysis to extract and fuse the optimized features at different scales to obtain third multi-scale fusion features;

[0037] Map the third multi-scale fusion features to a third shared semantic subspace through an artificial neural network.

[0038] Preferably, the fusion of the features of the three images to obtain a shared semantic space includes:

[0039] Adjust the global features of the second image and the global features of the third image to the same spatial reference system as the global features of the first image;

[0040] Align the global features of the first image with the adjusted global features of the second image and the global features of the third image;

[0041] Perform weighted averaging on the globally feature-aligned global features of the first image, the second image, and the third image to obtain a fourth fused feature;

[0042] Map the fourth fused feature to a second activation feature through a third fully connected layer;

[0043] Use multi-scale analysis to extract and fuse the second activation features at different scales to obtain a fourth multi-scale fused feature;

[0044] Optimize the fourth multi-scale fused feature using a data fitting term, a regularization term, and an argmin function to obtain the shared semantic space.

[0045] Preferably, the process of constructing the graph structure includes:

[0046] The unary nodes of each image belong to the feature space of the corresponding image; fuse the feature spaces of each image to obtain the feature space to which the ternary nodes belong;

[0047] Arrange the unary nodes around the ternary nodes and connect each unary node to the ternary node to obtain multiple first edges;

[0048] The feature space of the binary feature includes the feature space after fusing the feature spaces of any two images and the feature space after globally comparing the feature spaces of any two images; use the fused feature space as the fusion weight;

[0049] Take the average of the three obtained fusion weights to obtain a weight average value, and use the weight average value as the weight of the first edge connecting each unary node to the ternary node;

[0050] Use a weight adjustment function to respectively adjust the weights of the feature spaces after globally comparing the three global features to obtain the weights of the second edges connecting the corresponding two unary nodes to each other;

[0051] Construct the graph structure based on the ternary nodes, each unary node, multiple first edges and the weights of the first edges, multiple second edges and the weights of each second edge.

[0052] Preferably, the fusion of the graph structure using a graph convolutional neural network includes:

[0053] Construct the Laplacian matrix of the graph structure according to the degree matrix and the adjacency matrix of the graph structure;

[0054] Through each GCN layer in the graph convolutional neural network, each node in the Laplacian matrix is updated, and non-linear activation is performed after each GCN layer to obtain the representation of each node after activation of the corresponding GCN layer;

[0055] Aggregate the representations of all nodes after activation of the last GCN layer to obtain the global feature vector of the graph structure;

[0056] Map the global feature vector of the graph structure to the fusion result through the mapping weight.

[0057] Preferably, the updating and optimizing the graph structure based on the fusion result includes:

[0058] Analyze the fusion result to obtain an update vector, where the update vector includes the nodes to be updated and their update directions;

[0059] Based on the update vector, adjust the node weight vectors of each node in the graph structure;

[0060] According to the node weight vectors of each node after adjustment and the weights of the first edge / second edge, use the UpdateWeights function to update the weights of the corresponding edges;

[0061] According to the node weight vectors of each node after adjustment and the weights of each edge after update, use the AdjustTopology function to adjust the topological structure of the graph structure to obtain the adjusted graph structure;

[0062] Use the OptimizeGragh function to optimize the adjusted graph structure to obtain the optimized graph structure.

[0063] Preferably, the unmanned equipment includes unmanned aerial vehicles and autonomous vehicles.

[0064] Beneficial effects: By constructing a new multi-tuple information system, this method can better explore and utilize the potential connections between different modality data, improve the accuracy and efficiency of information processing, and thus enhance the perception ability and decision support ability of unmanned equipment in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0066] Figure 1 It is a flowchart of the method for constructing a multi-tuple unmanned equipment multi-modal information system in the embodiments of the present application. Detailed implementation manners

[0067] To make the above objects, features, and advantages of the present application more obvious and understandable, the following will describe in detail the specific implementation manners of the present application with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0068] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0069] As Figure 1 shown, this embodiment provides a method for constructing a multi-modal information system for multi-agent unmanned equipment, and the method includes:

[0070] S1: Receive multi-modal information collected by unmanned equipment, where the multi-modal information includes RGB images, lidar point cloud images, and infrared images.

[0071] In this embodiment, the unmanned equipment includes but is not limited to unmanned aerial vehicles and autonomous vehicles.

[0072] S2: Extract features from the three types of images respectively, obtain the features of the three types of images, and use the features of the three types of images as the unary tuple nodes corresponding to the respective images.

[0073] Specifically, the extracting features from the three types of images respectively includes:

[0074] Use a convolutional neural network to perform global feature extraction on the RGB image to obtain the first image global feature, and use the first image global feature as the unary tuple node of the RGB image.

[0075] The extraction process includes:

[0076] Step 1: Assume that the input RGB image is a three-dimensional tensor , whose dimension is H×W×C, where H is the image height, W is the image width, and C = 3 is the number of color channels;

[0077] Step 2: The convolutional layer consists of multiple convolutional kernels K iIt consists of each convolutional kernel sliding on the RGB image to calculate local features. Suppose the size of each convolutional kernel is k×k, and the convolution operation formula is expressed as:

[0078] ;

[0079] Among them, represents the value of the i-th channel of the feature map at position (j,j); represents the RGB image, represents the pixel value of the j-th channel of the RGB image at position (m+1,n+1); represents the weight value of the i-th convolutional kernel at position (m,n);

[0080] Step 3: Apply the non-linear activation function ReLU to the output of each convolutional layer for non-linear activation. The calculation formula is:

[0081] ;

[0082] Among them, represents the activated feature map;

[0083] Step 4: Reduce the spatial dimension of the feature map through the pooling layer (usually max pooling or average pooling) to reduce the number of parameters and the amount of calculation. The pooling operation is expressed as:

[0084] ;

[0085] Among them, represents the value of the pooled feature map at the i-th channel and position j; represents the value of the i-th channel of the feature map before pooling at the coordinate (m,n) within the pooling window; represents taking the maximum value within the pooling window;

[0086] Step 5: After multiple convolutional layers and pooling layers, the feature map will be flattened and passed to the fully connected layer. Suppose the weight matrix of the fully connected layer is W f , and the bias vector is b f , then the fully connected operation is expressed as:

[0087] ;

[0088] Among them, represents the output of the pooling layer processed by the activation function, and Z represents the fully connected layer;

[0089] Step 6: In order to extract global features, global average pooling is used after the fully connected layer to obtain the first image global feature. The calculation formula is:

[0090] ;

[0091] Wherein, represents the global feature of the first image; represents the size of the feature map after convolution and pooling operations, represents the value of the output of the pooling layer processed by the activation function at the position (i, j, k).

[0092] An encoder is used to perform multi-scale feature extraction on the RGB image to obtain the multi-scale feature of the first image. The calculation formula is:

[0093] ;

[0094] Wherein, represents the multi-scale feature of the first image at the first scale; represents the multi-scale feature of the first image at the second scale; represents the multi-scale feature of the first image at the third scale; represents the encoder for the RGB image; represents the RGB image.

[0095] PointNet is used to perform global feature extraction on the lidar point cloud image to obtain the global feature of the second image, and the global feature of the second image is used as the unary tuple node of the lidar point cloud image.

[0096] The extraction process includes:

[0097] Step 1: Assume the input point cloud P consists of N points, and each point represents a point in three-dimensional space;

[0098] ;

[0099] Step 2: Apply a learnable multi-layer perceptron to each point to extract local features. The calculation formula is:

[0100] ;

[0101] Wherein, represents the local feature of the i-th point; represents the multi-layer perceptron;

[0102] Step 3: Use the transformation network to learn a differentiable transformation to eliminate the influence of the arrangement of points in the point cloud, expressed as:

[0103] ;

[0104] Among them, represents the first network parameter of the transformation network; represents the second network parameter of the transformation network;

[0105] Step 4: Aggregate the transformed features through a symmetric function. Max pooling is used as the symmetric function in PointNet, and the calculation formula is:

[0106] ;

[0107] Among them, represents the aggregated features;

[0108] Step 5: Pass the aggregated features through the MLP again to extract global features, obtaining the second image global feature, and the calculation formula is:

[0109] ;

[0110] Among them, represents the second image global feature.

[0111] Use ResNet to extract global features from the infrared image, obtaining the third image global feature, and use the third image global feature as the unary tuple node of the infrared image.

[0112] The extraction process includes:

[0113] Step 1: Assume that the input infrared image is a three-dimensional tensor with dimensions H×W×C, where H is the image height, W is the image width, and C is the number of color channels (for infrared images, C = 1);

[0114] Step 2: In the first convolutional layer, use multiple convolutional kernels K to perform convolution operations on the infrared image to extract local features, and the calculation formula is:

[0115] ;

[0116] Among them, represents the local features; represents the bias term;

[0117] Step 3: Apply a non-linear activation function (ReLU function) to the output of the convolutional layer, and the calculation formula is:

[0118] ;

[0119] Among them, represents the activated local features;

[0120] Step 4: The residual block is the core component of ResNet. It contains two convolutional layers. The output of the residual block is the sum of the output of the second convolutional layer, the output of the first convolutional layer, and the bias term within the residual block. The calculation formula is:

[0121] ;

[0122] Among them, represents the output of the residual block; represents the output of the first convolutional layer; represents the output of the second convolutional layer; represents the bias term of the residual block;

[0123] Step 5: Use a pooling layer (such as max pooling) to reduce the spatial dimension of the output of the residual block. The calculation formula is:

[0124] ;

[0125] Among them, represents the value at the coordinate (m, n) within the pooling window on the i-th channel of the output of the residual block before pooling, represents the value at the position j on the i-th channel of the output of the residual block after pooling;

[0126] Step 6: At the end of the network, use global average pooling to aggregate all features and extract global features. The calculation formula is:

[0127] ;

[0128] Among them, represents the global feature; D represents the depth of the feature map after multiple convolutional and pooling layers; represents the value at the position (i, j, k) of the output after the pooling layer;

[0129] Step 7: Pass the one-dimensional feature vector after global average pooling through a fully connected layer to obtain the third image global feature for classification or other tasks. The calculation formula is:

[0130] ;

[0131] Among them, represents the third image global feature; represents the one-dimensional feature vector after global average pooling; represents the weight of the fully connected layer in ResNet; represents the bias of the fully connected layer in ResNet.

[0132] An encoder is used to perform multi-scale feature extraction on the infrared image to obtain the third image multi-scale feature. The calculation formula is:

[0133] ;

[0134] Among them, represents the third image multi-scale feature of the first scale; represents the third image multi-scale feature of the second scale; represents the third image multi-scale feature of the third scale; represents the encoder for the infrared image; represents the infrared image.

[0135] S3: Pair and fuse the features of the three images in pairs respectively to obtain three shared semantic subspaces, and use the three shared semantic subspaces as different binary group features respectively; fuse all the features of the three images to obtain a shared semantic space, and use the shared semantic space as a triple node.

[0136] Specifically, the fusion process of the feature of the first image and the feature of the second image includes:

[0137] Use bilinear interpolation method to adjust the global feature of the first image and the global feature of the second image to the same dimension, expressed as:

[0138] ;

[0139] ;

[0140] Among them, represents the adjusted global feature of the first image; represents the adjusted global feature of the second image; represents the bilinear interpolation function;

[0141] The expansion of the bilinear interpolation function is:

[0142] Given a feature map F, represents the value of the feature map at position m, n, and it is expected to interpolate F from size M×N to size , where s is the scaling factor used to determine the target size;

[0143] For each position i, j in the feature map F, the calculation formula of its bilinear interpolation is:

[0144] ;

[0145] Among them, denotes the floor function; mod denotes the modulo operation; this formula estimates the value of the target position (i, j) by calculating the weighted sum of the four nearest neighbor points in the feature map F, and the weights are determined by the distances from each neighbor point to the target position, thus achieving a smooth interpolation effect.

[0146] The aligned first image global feature and the second image global feature are mapped to the first feature space through the first fully connected layer, which can be expressed as:

[0147] ;

[0148] where, denotes the first feature space, denotes the first fully connected layer;

[0149] In the first feature space, the aligned first image global feature and the second image global feature are weighted and summed to obtain the first fusion feature, and the calculation formula is:

[0150] ;

[0151] ;

[0152] where, denotes the first fusion feature; denotes the corresponding weight of the adjusted first image global feature; denotes the corresponding weight of the adjusted second image global feature;

[0153] The first fusion feature is subjected to a non-linear transformation (using the sigmoid activation function) to obtain the first activation feature ;

[0154] The feature pyramid network is used to perform multi-scale feature extraction and fusion on the first activation feature to obtain the first multi-scale fusion feature , which helps to detect objects of different sizes;

[0155] The first multi-scale fusion feature is mapped to the first shared semantic subspace through the second fully connected layer , which can be expressed as:

[0156] ;

[0157] where, denotes the second fully connected layer.

[0158] Meanwhile, a loss function is defined to train the network, which usually includes position loss, classification loss, and confidence loss. The calculation formula of the first loss function is:

[0159] ;

[0160] Among them, represents the first loss function; represents the first balance coefficient; represents the second balance coefficient; represents the third balance coefficient;

[0161] Use the first loss function to perform backpropagation on the feature pyramid network, and use an optimization algorithm (such as gradient-based optimization algorithms like SGD, Adam, etc.) to update the network weights;

[0162] During the training process, data augmentation techniques can also be used to improve the generalization ability of the model.

[0163] Through the above steps, the finally output first shared semantic subspace contains the fused features from the RGB image and the lidar point cloud image, which can provide richer information and enhance the performance of object detection and other related tasks.

[0164] The fusion process of the features of the first image and the features of the third image includes:

[0165] Perform convolution on the multi-scale features of the first image at each scale through a 3×3 convolution operation to obtain the first convolution feature. The calculation formula is:

[0166] ;

[0167] Among them, represents the i th scale of the first convolution feature; represents that the 3×3 convolution operation is used to extract deep features; represents the i th scale of the multi-scale features of the first image;

[0168] Perform convolution on the multi-scale features of the third image at each scale through a 3×3 convolution operation to obtain the second convolution feature. The calculation formula is:

[0169] ;

[0170] Among them, represents the i th scale of the second convolution feature; represents the i th scale of the multi-scale features of the third image;

[0171] Add the first convolution feature and the second convolution feature to obtain the common feature. The calculation formula is:

[0172] ;

[0173] Among them, represents the common feature of the i th scale;

[0174] Subtract the second convolutional feature from the first convolutional feature, and then divide by 2 to obtain the private feature of the RGB image. The calculation formula is:

[0175] ;

[0176] Among them, represents the private feature of the RGB image of the i th scale;

[0177] Subtract the first convolutional feature from the second convolutional feature, and then divide by 2 to obtain the private feature of the infrared image. The calculation formula is:

[0178] ;

[0179] Among them, represents the private feature of the infrared image of the i th scale;

[0180] Adopt the ECA attention mechanism to enhance the features of the common feature, the private feature of the RGB image, and the private feature of the infrared image respectively. The calculation formula is:

[0181] ;

[0182] ;

[0183] ;

[0184] Among them, represents the enhanced common feature of the i th scale; represents the enhanced private feature of the RGB image of the i th scale; represents the enhanced private feature of the infrared image of the i th scale; represents the ECA attention mechanism;

[0185] Fuse the enhanced common feature, the private feature of the RGB image, and the private feature of the infrared image at each scale respectively to obtain the second fusion feature of each scale. The calculation formula is:

[0186] ;

[0187] Among them, represents the second fusion feature of the i th scale;

[0188] Reconstruct the second fusion features at all scales through the decoder to obtain the second multi-scale fusion features , and the calculation formula is:

[0189] ;

[0190] wherein, represents the decoder; , , respectively represent the second fusion features at the 1st, 2nd, and 3rd scales; and use the second multi-scale fusion features as the second shared semantic subspace.

[0191] Meanwhile, construct an information perception loss function to guide the network training, and the information perception loss function is expressed as:

[0192] ;

[0193] wherein, represents the information perception loss function; represents the intensity loss; represents the gradient loss; represents the pixel loss based on information perception; represents the weight of; represents the weight of; represents the weight of;

[0194] Use an appropriate optimization algorithm (such as the Adam optimizer) to minimize the information perception loss function to train the network parameters.

[0195] The fusion process of the features of the second image and the features of the third image includes:

[0196] Perform the same spatial transformation on the global features of the second image and the global features of the third image, which can be expressed as:

[0197] ;

[0198] ;

[0199] wherein, represents the transformed global features of the second image; represents the transformed global features of the third image; represents the spatial transformation operation;

[0200] In this embodiment, the spatial transformation is an affine transformation, and the transformation process is as follows:

[0201] First, define the affine transformation matrix. The affine transformation matrix T is a 2×3 matrix used to represent transformation operations such as rotation, translation, and scaling. The affine transformation matrix is defined as:

[0202] ;

[0203] where a , b , c , d , e , f are all parameters to be determined;

[0204] Taking the global features of the second image as an example, each feature point (x, y) is transformed, and the transformed coordinates are (x', y'). The transformation formula is:

[0205] ;

[0206] where represents the linear transformation matrix used to implement transformations such as rotation, scaling, and shearing; represents the translation vector used to implement the translation of the image;

[0207] Each feature point in the global features of the third image is transformed through the above transformation formula;

[0208] Then, at least three sets of corresponding points are required to determine the six parameters to be determined. Specifically, matching feature point pairs (x1, y1)-(x2, y2) are found in the global features of the second image and the global features of the third image through the feature point matching algorithm SIFT;

[0209] Substitute these matching feature point pairs into the above transformation formula to obtain a system of equations, and solve this system of equations to determine the values of the six parameters;

[0210] After determining the affine transformation matrix T, each feature point in the global features of the second image and the global features of the third image is transformed through the transformation formula with the determined parameter values, and the global features of the second image and the global features of the third image are transformed into the same space.

[0211] Fuse the globally transformed second image features with the globally transformed third image features to obtain the third fused feature. The calculation formula is:

[0212] ;

[0213] where represents the third fused feature; represents the fusion operation, which can be simple data splicing, weighted average, or other complex fusion strategies;

[0214] Optimize the third fusion feature by using a data fitting term, a regularization term, and an argmin function to obtain an optimized feature. The calculation formula is:

[0215] ;

[0216] Where, represents the optimized feature; represents the data fitting term; represents the regularization term;

[0217] Extract and fuse the optimized features at different scales using multi-scale analysis to obtain a third multi-scale fusion feature. The calculation formula is:

[0218] ;

[0219] Where, represents the third multi-scale fusion feature; is the union symbol, used to perform a merging operation on the features after the feature extraction operation at different scales; represents the i th feature extraction operation at the scale;

[0220] In this embodiment, a difference of Gaussian pyramid is used for multi-scale analysis. The multi-scale analysis process is as follows:

[0221] Construct a Gaussian pyramid: Convolve the input optimized feature with Gaussian kernels of different standard deviations to obtain a series of images with different degrees of blurriness, forming a Gaussian pyramid. Among them, assuming the original optimized feature is I(X,Y), the calculation formula for the k-th layer Gaussian pyramid image G k (x,y) is:

[0222] ;

[0223] Where, represents the Gaussian kernel function; represents the standard deviation;

[0224] Construct a difference of Gaussian pyramid: Subtract two adjacent layers of Gaussian pyramid images to obtain a difference of Gaussian pyramid. The calculation formula is:

[0225] ;

[0226] Where, represents the k-th layer difference of Gaussian pyramid image; represents the (k + 1)-th Gaussian pyramid image;

[0227] Extract optimized features at different scales: For each layer of the Gaussian difference pyramid, extract corresponding features (such as edge features, texture features, etc.) according to specific requirements;

[0228] Fuse features at different scales: Fuse the extracted optimized features at different scales, and a weighted average method can be used to obtain the third multi-scale fusion feature.

[0229] Map the third multi-scale fusion feature to the third shared semantic subspace through an artificial neural network , which can be expressed as:

[0230] ;

[0231] where, represents the neural network.

[0232] Meanwhile, define a second loss function to train the neural network to ensure that the third shared semantic subspace can accurately reflect the information of the input features. The expression of the second loss function is:

[0233] ;

[0234] where, represents the second loss function; represents the semantic subspace representation of the ground truth, represents the regularization parameter; represents the complexity of the neural network;

[0235] Use an appropriate optimization algorithm (such as the gradient descent method) to minimize the second loss function and train the parameters of the neural network.

[0236] The all-fusion of the features of the three images to obtain the shared semantic space includes:

[0237] Adjust the global features of the second image and the global features of the third image to the same spatial reference system as the global features of the first image, which can be expressed as:

[0238] ;

[0239] ;

[0240] where, represents the adjusted global features of the second image; represents the adjusted global features of the third image; represents the transformation function for adjusting the global features of the second image to the same reference system as the global features of the first image; A transformation function for adjusting the third image global feature to the same reference system as the first image global feature; here, the two transformation functions , are both affine transformations, and their specific transformation processes are as shown in the above spatial transformation, with the only difference being the objects they are applied to.

[0241] Align the first image global feature with the adjusted second image global feature and the third image global feature, expressed as:

[0242] ;

[0243] Among them, represents the aligned first image global feature; represents the feature alignment function;

[0244] Perform weighted averaging on the feature-aligned first image global feature, second image global feature, and third image global feature to obtain a fourth fused feature, and the calculation formula is:

[0245] ;

[0246] ;

[0247] Among them, represents the fourth fused feature; represents the weight coefficient of; represents the weight coefficient of; represents the weight coefficient of;

[0248] Map the fourth fused feature to a second activation feature through a third fully connected layer ;

[0249] When the feature has multi-scale characteristics, use multi-scale analysis to extract and fuse the second activation features at different scales to obtain a fourth multi-scale fused feature, and the calculation formula is:

[0250] ;

[0251] Among them, represents the fourth multi-scale fused feature; represents the feature extraction operation at the i th scale; this multi-scale analysis also uses a Gaussian difference pyramid for multi-scale analysis, and the multi-scale analysis process is the same as the multi-scale analysis process of the above optimized feature, with the only difference being the objects they are applied to.

[0252] Optimize the fourth multi-scale fusion feature using a data fitting term, a regularization term, and an argmin function to obtain the shared semantic space. The calculation formula is as follows:

[0253] ;

[0254] Among them, represents the shared semantic space; represents the data fitting term; represents the regularization term.

[0255] Meanwhile, define a third loss function to train the third fully connected layer , to ensure that the shared semantic space can accurately reflect the information of the input features. The expression of the third loss function is:

[0256] ;

[0257] Among them, represents the third loss function; represents the feature representation of the ground truth, represents the regularization parameter; represents the complexity of the third fully connected layer;

[0258] Use an appropriate optimization algorithm (such as the Adam optimizer) to minimize the third loss function and train the parameters of the third fully connected layer.

[0259] S4: Connect each single-tuple node to each other and connect each single-tuple node to the triple-tuple node. Calculate the weights of each edge based on the feature space of the binary-tuple features and construct a graph structure.

[0260] Specifically, the single-tuple nodes of each image belong to the feature space of the corresponding image, which is expressed as: , , ; , , represent the feature spaces of the RGB image, the lidar point cloud image, and the infrared image respectively;

[0261] Fuse the feature spaces of each image to obtain the feature space to which the triple-tuple node belongs, which is expressed as: ;

[0262] Arrange the single-tuple nodes around the triple-tuple node and connect each single-tuple node to the triple-tuple node to obtain multiple first edges;

[0263] The feature space of the binary tuple features includes the feature space after the fusion of the feature spaces of any two images, and the feature space after the global feature comparison of the feature spaces of any two images; the fused feature space is used as the fusion weight;

[0264] Take the mean of the three obtained fusion weights to get the weight mean, and use the weight mean as the weight of the first edge connecting each single-tuple node and the triple-tuple node; the calculation formula is:

[0265] ;

[0266] ;

[0267] ;

[0268] ;

[0269] Among them, represents the weight mean; represents the fusion weight of the feature space of the RGB image and the feature space of the lidar point cloud image; represents the fusion weight of the feature space of the RGB image and the feature space of the infrared image; represents the fusion weight of the feature space of the lidar point cloud image and the feature space of the infrared image; represents the fusion function;

[0270] Use the weight adjustment function to adjust the weights of the feature spaces after the three global feature comparisons respectively, and obtain the weights of the second edges connecting the corresponding two single-tuple nodes respectively; expressed as:

[0271] ;

[0272] ;

[0273] ;

[0274] ;

[0275] ;

[0276] ;

[0277] Among them, represents the feature space after the first global feature comparison; represents the feature space after the second global feature comparison; represents the feature space after the third global feature comparison; represents the global feature comparison function; A tuple node representing an RGB image The weight of the second edge connected to the tuple node of the lidar point cloud image ; A tuple node representing an RGB image The weight of the second edge connected to the tuple node of the infrared image ; A tuple node representing the lidar point cloud image The weight of the second edge connected to the tuple node of the infrared image ; Represents a weight adjustment function;

[0278] Based on the triple node, each tuple node, multiple first edges and the weights of the first edges, multiple second edges and the weights of each second edge, construct the graph structure, and the graph structure is represented as:

[0279] G ={ V , E , W};

[0280] ;

[0281] ;

[0282] .

[0283] S5: Use a graph convolutional neural network to fuse the graph structure to obtain a fusion result.

[0284] Specifically, the use of a graph convolutional neural network to fuse the graph structure includes:

[0285] According to the degree matrix and adjacency matrix of the graph structure, construct the Laplacian matrix of the graph structure, which can be expressed as:

[0286] L = D - A;

[0287] Among them, L represents the Laplacian matrix of the graph structure; D represents the degree matrix of the graph structure; A represents the adjacency matrix of the graph structure;

[0288] Through each GCN layer in the graph convolutional neural network, update each node in the Laplacian matrix. The update formula for the l th GCN layer is:

[0289] ;

[0290] Among them, represents the l th node in the Laplacian matrix after the inodes; represents a GCN layer; represents the l weights of the represents the l weights from the j-th neuron to the i-th neuron in the represents the set of neurons in the previous layer l connected to the i-th neuron in the represents the l output of the j-th neuron in the

[0291] After each GCN layer, a non-linear activation is performed to obtain the representations of the nodes after activation of the corresponding GCN layer. The calculation formula is:

[0292] ;

[0293] where represents the l -th node in the Laplacian matrix updated after the i layer of the GCN layer; represents the ReLU activation function;

[0294] Aggregate the representations of all nodes after activation of the last GCN layer to obtain the global feature vector of the graph structure. The calculation formula is:

[0295] ;

[0296] where represents the global feature vector of the graph structure; V represents the number of nodes in the graph structure; represents the L -th node in the Laplacian matrix updated and activated after the i layer of the GCN layer;

[0297] Map the global feature vector of the graph structure to the fusion result through the mapping weights, which is expressed as:

[0298] ;

[0299] where represents the fusion result; represents the mapping weights.

[0300] At the same time, define the fourth loss function for training the GCN, which includes the loss term of the task objective (such as classification loss or regression loss) and the second regularization term. The expression of the fourth loss function is:

[0301] ;

[0302] Among them, represents the fourth loss function; represents the classification loss; represents the regularization parameter; represents the regression loss; represents the weight of the GCN;

[0303] Minimize the fourth loss function using an optimization algorithm (such as Adam) and update the network weights:

[0304] ;

[0305] Among them, represents the network weight of the updated GCN; represents the network weight of the GCN before update; represents the partial derivative; represents the learning rate.

[0306] S6: Update and optimize the graph structure based on the fusion result to obtain an optimized graph structure.

[0307] Specifically, the updating and optimizing the graph structure based on the fusion result includes:

[0308] Analyze the fusion result to obtain an update vector, and the update vector includes the nodes to be updated and their update directions (enhanced or weakened), and the calculation formula is:

[0309] ;

[0310] Among them, represents the analysis function; represents the update vector;

[0311] Adjust the node weight vectors of each node in the graph structure based on the update vector, and the calculation formula is:

[0312] ;

[0313] Among them, represents the node weight vectors of each node after adjustment; represents the original node weight vectors of each node; represents the initial weight matrix of each node; represents the Hadamard product;

[0314] According to the node weight vectors of each node after adjustment and the weights of the first edge / second edge, use the UpdateWeights function to update the weights of the corresponding edges, which is expressed as:

[0315] ;

[0316] Among them, represents the weights of each edge after update; represents the weights of each original edge;

[0317] According to the node weight vectors of each node after adjustment and the weights of each edge after update, use the AdjustTopology function to adjust the topological structure of the graph structure (such as adding or deleting nodes and weighted edges), and obtain the adjusted graph structure, which is expressed as:

[0318] ;

[0319] Among them, represents the adjusted graph structure, is the set of nodes in the adjusted graph structure, is the set of weighted edges in the adjusted graph structure;

[0320] Use the OptimizeGragh function to optimize the adjusted graph structure, and obtain the optimized graph structure, which is expressed as:

[0321] ;

[0322] Among them, represents the optimized graph structure, is the set of nodes in the optimized graph structure, is the set of weighted edges in the optimized graph structure.

[0323] The method for constructing a multi-tuple unmanned equipment multi-modal information system provided in this embodiment aims to construct a new multi-tuple information system in a more refined and systematic manner, which can better explore and utilize the potential connections between different modal data, improve the accuracy and efficiency of information processing, and thus enhance the perception ability and decision support ability of unmanned equipment in complex environments. This method is expected to promote the progress of unmanned equipment information processing technology and provide support for achieving a higher level of autonomy and intelligence.

[0324] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0325] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for constructing a multimodal information system of multi-group unmanned equipment, characterized in that: include: S1: receiving multimodal information collected by unmanned equipment, wherein the multimodal information includes RGB images, laser radar point cloud images, and infrared images; S2: extract features from the three images respectively, obtain features of the three images respectively, and use the features of the three images as tuple nodes of the corresponding images respectively; S3: Pair the features of the three images in pairs and fuse them to obtain three shared semantic subspaces, and use the three shared semantic subspaces as different bigram features; fuse all the features of the three images to obtain a shared semantic space, and use the shared semantic space as a triplet node; S4: Connect each tuple node to each other, and connect each tuple node to the triple node, calculate the weight of each edge based on the feature space of the two-tuple feature, and build a graph structure; S5: using a graph convolutional neural network to fuse the graph structure to obtain a fusion result; S6: Based on the fusion result, the graph structure is updated and optimized to obtain an optimized graph structure.

2. The method for constructing a multimodal information system of multi-group unmanned equipment according to claim 1, characterized in that: Using a convolutional neural network to extract global features from the RGB image to obtain a first image global feature, and using the first image global feature as a tuple node of the RGB image; Using an encoder to perform multi-scale feature extraction on the RGB image to obtain multi-scale features of a first image; Using PointNet to extract global features from the laser radar point cloud image to obtain a second image global feature, and using the second image global feature as a tuple node of the laser radar point cloud image; Using ResNet to extract global features of the infrared image to obtain a third image global feature, and using the third image global feature as a tuple node of the infrared image; An encoder is used to extract multi-scale features from the infrared image to obtain multi-scale features of a third image.

3. The method for constructing a multi-modal information system of multi-group unmanned equipment according to claim 2, characterized in that: The process of fusing the features of the first image with the features of the second image includes: Using a bilinear interpolation method to adjust the global features of the first image and the global features of the second image to the same dimension; Mapping the aligned first image global features and the second image global features to the first feature space through a first fully connected layer; Performing a weighted summation of the aligned first image global features and the second image global features in the first feature space to obtain a first fusion feature; Performing a nonlinear transformation on the first fused feature to obtain a first activated feature; Using a feature pyramid network to perform multi-scale feature extraction and fusion on the first activation feature to obtain a first multi-scale fusion feature; The first multi-scale fusion features are mapped to a first shared semantic subspace through a second fully connected layer.

4. The method for constructing a multimodal information system of multi-group unmanned equipment according to claim 2, characterized in that: The fusion process of the features of the first image and the features of the third image includes: Convolving the multi-scale features of the first image at each scale respectively through a 3×3 convolution operation to obtain a first convolution feature; Convolving the third image multi-scale features of each scale respectively through a 3×3 convolution operation to obtain a second convolution feature; Adding the first convolution feature and the second convolution feature to obtain a common feature; Subtract the second convolution feature from the first convolution feature, and divide the result by 2 to obtain a private feature of the RGB image; Subtract the first convolution feature from the second convolution feature, and divide the result by 2 to obtain a private feature of the infrared image; The ECA attention mechanism is used to enhance the common features, private features of RGB images, and private features of infrared images respectively; The enhanced common features, the private features of the RGB image, and the private features of the infrared image are fused at each scale to obtain the second fused features at each scale; The second fusion features of all scales are reconstructed through a decoder to obtain second multi-scale fusion features; and the second multi-scale fusion features are used as the second shared semantic subspace.

5. The method for constructing a multi-modal information system of multi-group unmanned equipment according to claim 2, characterized in that: The fusion process of the features of the second image and the features of the third image includes: Performing the same spatial transformation on the second image global feature and the third image global feature; Fusing the second image global feature after spatial transformation with the third image global feature to obtain a third fused feature; The third fusion feature is optimized by using a data fitting term, a regularization term and an argmin function to obtain an optimized feature; Use multi-scale analysis to extract optimized features of different scales and fuse them to obtain the third multi-scale fusion feature; The third multi-scale fusion feature is mapped to a third shared semantic subspace through an artificial neural network.

6. The method for constructing a multimodal information system of multi-group unmanned equipment according to claim 2, characterized in that: The method of fusing all the features of the three images to obtain a shared semantic space includes: Adjusting the second image global feature and the third image global feature to the same spatial reference system as the first image global feature; Perform feature alignment on the first image global feature, the adjusted second image global feature, and the third image global feature; Perform weighted averaging of the first image global features, the second image global features, and the third image global features that are feature aligned to obtain a fourth fusion feature; Mapping the fourth fusion feature into a second activation feature through a third fully connected layer; Use multi-scale analysis to extract second activation features of different scales and fuse them to obtain a fourth multi-scale fusion feature; The fourth multi-scale fusion feature is optimized by using a data fitting term, a regularization term, and an argmin function to obtain the shared semantic space.

7. The method for constructing a multimodal information system of multi-group unmanned equipment according to claim 2, characterized in that: The process of building a graph structure includes: The one-tuple node of each image belongs to the feature space of the corresponding image; the feature spaces of each image are fused to obtain the feature space to which the three-tuple node belongs; Arrange the unigram nodes around the triplet nodes, and connect each unigram node with the triplet node to obtain a plurality of first edges; The feature space of the binary feature includes the feature space after the feature spaces of any two images are fused, and the feature space after the feature spaces of any two images are compared with the global features; the fused feature space is used as the fusion weight; Taking the average of the three obtained fusion weights to obtain a weight average, and using the weight average as the weight of the first edge connecting each unigram node and the triplet node; Use the weight adjustment function to adjust the weights of the feature spaces after the comparison of the three global features, and obtain the weights of the second edges corresponding to the connection between the two unigram nodes; The graph structure is constructed based on the triplet nodes, each tuple node, a plurality of first edges and weights of the first edges, a plurality of second edges and weights of each second edge.

8. The method for constructing a multimodal information system of multi-group unmanned equipment according to claim 1, characterized in that: The adopting of a graph convolutional neural network to fuse the graph structure includes: Construct the Laplacian matrix of the graph structure based on the degree matrix and adjacency matrix of the graph structure; Through each GCN layer in the graph convolutional neural network, each node in the Laplacian matrix is ​​updated, and nonlinear activation is performed after each GCN layer to obtain the representation of each node after the corresponding GCN layer is activated; Aggregate the representations of all nodes after the activation of the last GCN layer to obtain the global feature vector of the graph structure; The global feature vector of the graph structure is mapped to the fusion result through mapping weights.

9. The method for constructing a multi-modal information system of multi-group unmanned equipment according to claim 1, characterized in that: The updating and optimizing of the graph structure based on the fusion result includes: Analyze the fusion result to obtain an update vector, where the update vector includes nodes that need to be updated and their update directions; Adjusting a node weight vector of each node in the graph structure based on the update vector; According to the adjusted node weight vector of each node and the weights of the first and second edges, the UpdateWeights function is used to update the weights of the corresponding edges; According to the adjusted node weight vectors of each node and the updated weights of each edge, the AdjustTopology function is used to adjust the topological structure of the graph structure to obtain the adjusted graph structure; The OptimizeGraph function is used to optimize the adjusted graph structure to obtain the optimized graph structure.

10. The method for constructing a multi-modal information system of multi-group unmanned equipment according to claim 1, characterized in that: The unmanned equipment includes drones and autonomous driving vehicles.

Citation Information

Patent Citations

  • Sensor data fusion method and device based on graph neural network, and storage medium

    CN118364432A

  • USE OF LANGUAGE MODELS IN AUTONOMOUS AND SEMI-AUTOMATIC SYSTEMS AND APPLICATIONS

    DE102024116258A1