Target matching method based on cross-domain multi-modal fusion coding
Through the target matching method of cross-domain multimodal fusion encoding, the problems of heterogeneity and large perspective differences in multimodal data are solved, and the deep feature fusion of multi-view images, lidar point clouds and text data is realized, which improves the accuracy and reliability of target matching and reduces the training data needs.
Patent Information
- Application Number
- CN202510767997.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
In cross-domain multi-unmanned equipment application scenarios, the heterogeneity of multimodal data and perspectives vary greatly. It is difficult for traditional methods to effectively integrate multi-view image data, lidar point cloud data and text data. In addition, existing model training requires a large amount of corresponding data, and data is scarce, resulting in difficulty in matching targets.
The target matching method based on cross-domain multimodal fusion encoding is adopted. By obtaining multi-view image data, lidar point cloud data and keyword text data, visual descriptors, viewpoint descriptors and keyword feature vectors are extracted respectively, and three-layer encoding is used to perform three-layer encoding, fuse multimodal features, and find the target matching result through cosine similarity matching.
It effectively realizes the deep feature fusion of three modalities: image, point cloud, and text, improves the accuracy and reliability of target matching, solves the problem of multimodal data heterogeneity, reduces the training data needs, and improves the generalization ability and application efficiency of the model.
Smart Images

Figure CN120277626A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of multi-unmanned equipment multi-modal information fusion, and particularly to a target matching method based on cross-domain multi-modal fusion coding. Background Art
[0002] In the application scenario of cross-domain multi-unmanned equipment, achieving accurate target matching is crucial. However, there are many challenges currently. On the one hand, multi-modal data (such as multi-view image data, lidar point cloud data, text data) has heterogeneity, and traditional methods are difficult to effectively fuse the deep features of these data; on the other hand, when cross-domain multi-unmanned equipment collects data, the perspective differences are large, and traditional modal alignment methods cannot handle it well. In addition, although the MLIP multi-modal fusion coding target matching model improved based on the GLIP model has high complexity capabilities, training requires a large amount of corresponding data of images, point clouds, and texts in three modalities, and currently, such publicly available data is extremely scarce, seriously restricting the training and application of the model. Summary of the Invention
[0003] Based on this, it is necessary to provide a target matching method based on cross-domain multi-modal fusion coding, and this method includes: S1: Obtain the keyword text data to be detected, as well as the multi-view image data and lidar point cloud data of multi-unmanned equipment; split the image feature matrix obtained by encoding the multi-view image data to obtain visual descriptors; perform visible point screening on the point-by-point feature vectors obtained by encoding the lidar point cloud data to obtain view point descriptors; convert the first keyword feature vector extracted from the keyword text data into a digital form to obtain the second keyword feature vector; S2: Fuse the visual descriptors and view point descriptors into a multi-view fusion 3D expression feature vector; perform a linear transformation on the second keyword feature vector to obtain the third keyword feature vector; and splice the multi-view fusion 3D expression feature vector and the third keyword feature vector to obtain the first joint expression vector; S3: Arrange three layers of transformer encoding blocks. The input of the first layer of transformer encoding block is the first joint expression vector, and the input of the remaining two layers of transformer encoding blocks is the splicing vector of the output of the previous layer and the third keyword feature vector. The third layer of transformer encoding block outputs a three-level joint expression vector; S4: Find the part with the highest cosine similarity between the three-level joint expression vector and the first keyword feature vector as the target matching result.
[0004] Preferably, encoding the multi-view image data includes: respectively passing the image data of each view through a convolutional neural network to obtain the image feature matrix corresponding to each view.
[0005] Preferably, splitting the image feature matrix includes: performing slicing or pooling operations on the image feature matrix corresponding to any one view in a specific dimension to obtain multiple visual descriptors corresponding to the view.
[0006] Preferably, encoding the lidar point cloud data includes: encoding the lidar point cloud data through PointNet to obtain a point-by-point feature vector for each point.
[0007] Preferably, performing visible point screening on the point-by-point feature vector includes: Calculating the Euclidean distance between a point in the lidar point cloud data i and any other point, and multiplying the Euclidean distance by the weight matrix to obtain a first product; Traversing the remaining points in the lidar point cloud data, calculating the first products corresponding to the point i and all other points, and summing all the calculated first products; Passing the summation result through an activation function to obtain the visibility score of the point i ; Comparing the visibility score of the point i with a set threshold. When the visibility score of the point i is greater than or equal to the set threshold, taking the point-by-point feature vector of the point i as a view point descriptor; otherwise, discarding it; Traversing all points in the lidar point cloud data, comparing the visibility scores of all points with the set threshold to obtain multiple view point descriptors.
[0008] Preferably, converting the first keyword feature vector extracted from the keyword text data into a digital form includes: Passing the keyword text data through a word vector model to obtain a first keyword feature vector; Splitting the first keyword feature vector into a sequence composed of individual numbers to obtain a second keyword feature vector.
[0009] Preferably, fusing the visual descriptor and the view point descriptor into a multi-view fusion 3D expression feature vector includes: Summing the products of all visual descriptors under any one view and a first fusion weight to obtain a second summation result; Summing the second summation results under all views to obtain the total visual descriptor; Summing the products of all view point descriptors and a second fusion weight to obtain the total view point descriptor; Adding the total visual descriptor and the total view point descriptor to obtain a multi-view fusion 3D expression feature vector.
[0010] Preferably, the splicing of the multi-view fusion 3D expression feature vector and the third keyword feature vector includes: Multiply the multi-view fusion 3D expression feature vector by the feature vector weight to obtain a second product; Add the second product and the third keyword feature vector to obtain a first joint expression vector.
[0011] Preferably, the splicing of the output of the first / second layer transformer encoding block and the third keyword feature vector includes: Multiply the output of the first / second layer transformer encoding block by the feature vector weight to obtain a third / fourth product; Add the third / fourth product and the third keyword feature vector to obtain the input of the second / third layer transformer encoding block.
[0012] Preferably, S4 includes: Split the three-level joint expression vector into multiple first feature blocks; Split the first keyword feature vector into multiple second feature blocks; Calculate the cosine similarity between any one first feature block and each second feature block to obtain a row / column of similarity values; Traverse all first feature blocks, and combine the row / column of similarity values calculated according to each first feature block into a similarity matrix; Use the first feature block with the highest similarity value to the second feature block in the similarity matrix as the target matching result.
[0013] Beneficial effects: Based on multi-view image data, lidar point cloud data, and keyword text data to be detected, the method respectively obtains a visual descriptor, a viewpoint descriptor, and a second keyword feature vector; fuses the visual descriptor and the viewpoint descriptor into a multi-view fusion 3D expression feature vector; linearly transforms the second keyword feature vector to obtain a third keyword feature vector; and splices the multi-view fusion 3D expression feature vector and the third keyword feature vector to obtain a first joint expression vector; passes the first joint expression vector through three layers of transformer encoding blocks to output a three-level joint expression vector; and finds the part with the highest cosine similarity between the three-level joint expression vector and the first keyword feature vector as the target matching result. The method adopts a multi-modal fusion coding method with keyword hints, effectively realizing the deep feature fusion of the three modalities of image, point cloud, and text, solving the problem of multi-modal data heterogeneity, and thus improving the accuracy and reliability of target matching. Description of the Drawings
[0014] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0015] Figure 1 It is a flowchart of the target matching method based on cross - domain multi - modal fusion coding in the embodiments of the present application. Specific embodiments
[0016] To make the above - mentioned objects, features, and advantages of the present application more clearly understandable, the following will give a detailed description of the specific embodiments of the present application in conjunction with the drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0017] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0018] As Figure 1 shown, this embodiment provides a target matching method based on cross - domain multi - modal fusion coding, and the method includes: S1: Obtain the keyword text data to be detected, the multi - view image data of multiple unmanned equipment, and the lidar point cloud data; split the image feature matrix obtained by encoding the multi - view image data to obtain visual descriptors; perform visible point screening on the point - by - point feature vectors obtained by encoding the lidar point cloud data to obtain viewpoint descriptors; convert the first keyword feature vector extracted from the keyword text data into a digital form to obtain the second keyword feature vector.
[0019] The multi - view image data contains image information obtained from different perspectives and can provide rich appearance features; the lidar point cloud data can accurately describe the spatial structure of the target; the keyword text data provides key clues for target matching from the semantic level.
[0020] In this embodiment, splitting the image feature matrix obtained by encoding the multi - view image data includes: The image data of each view is respectively passed through a convolutional neural network to obtain an image feature matrix corresponding to each view; Perform slicing or pooling operations on the image feature matrix corresponding to any one view in a specific dimension to obtain multiple visual descriptors corresponding to the view, which is expressed by the formula: ; where, represents the v th visual descriptor of the j th view; s represents the step size of slicing or pooling; represents the v th view corresponding image feature matrix; x represents the image feature matrix index in the height dimension; each visual descriptor corresponds to a local window, and the height range of the window ( x value range) is , that is, the j th window covers s consecutive rows in the height direction; y represents the image feature matrix index in the width dimension, y value range is , that is, traversing the entire width of the image feature matrix; represents a full slice operation on the channel dimension of the image feature matrix , that is, extracting all channel features at this spatial position ; n represents the width size of the image feature matrix , that is, the total number of elements in the width direction of the image feature matrix.
[0021] The meaning of this formula is: For the image feature matrix of the th view (dimension is , that is, height × width × channel), within the th local window (height range , width range ), perform pooling operations (average pooling) along the height and width dimensions, and retain all information for the channel dimension, finally generating the th visual descriptor of the th view .
[0022] The pooling step size determines the height range of the window, each window covers rows, and the final number of visual descriptors is .
[0023] Example: Assume the size of the image feature matrix is (i.e., height = 8, width = 10, channels = 256), and the pooling stride , then: The height range of each window is rows (such as corresponding to , corresponding to , and so on), and a total of visual descriptors ( ) are generated.
[0024] For each window, perform feature averaging in the channel dimension for the spatial region. Finally, the dimension of each visual descriptor is (consistent with the number of channels).
[0025] This operation realizes the extraction of feature expressions for local regions from multi-view images, providing fine-grained visual information for subsequent cross-modal fusion.
[0026] Furthermore, encoding the lidar point cloud data includes: encoding the lidar point cloud data through PointNet to obtain the per-point feature vector for each point.
[0027] Even further, screening the visible points from the per-point feature vectors includes: Calculating the Euclidean distance between a point i in the lidar point cloud data and any other point, and multiplying the Euclidean distance by the weight matrix to obtain a first product; Traverse the remaining points in the lidar point cloud data, calculate the first products corresponding to point i and all other points, and sum all the calculated first products; Pass the sum result through an activation function to obtain the visibility score of point i ; Compare the visibility score of point i with a set threshold. When the visibility score of point i is greater than or equal to the set threshold, use the per-point feature vector of point i as a viewpoint descriptor; otherwise, discard it; Traverse all the points in the lidar point cloud data, compare the visibility scores of all the points with the set threshold to obtain multiple viewpoint descriptors.
[0028] In this embodiment, the conversion of the first keyword feature vector extracted from the keyword text data into a digital form includes: The keyword text data is passed through a word vector model to obtain the first keyword feature vector; The first keyword feature vector is split into a sequence composed of individual numbers to obtain the second keyword feature vector.
[0029] This embodiment also provides a KGM keyword generation pre-training module, which needs to be pre-trained: 1. Data input: Input multi-view image data , lidar point cloud data , text data , and these data contain different modal information of the target.
[0030] 2. Feature extraction: Perform feature extraction on the multi-view image data using a convolutional neural network to obtain the image feature vector as described above .
[0031] Perform feature extraction on the lidar point cloud data using PointNet to obtain the point cloud feature vector as described in the PointNet encoding process above.
[0032] Perform feature extraction on the text data using a word vector model to obtain the text feature vector as described above.
[0033] 3. Acquisition of text information in the point cloud modality: Using the point cloud feature vector as the weight and the text feature vector as the value, input into the layer transformer encoding block for encoding. Assume that the input of the th layer transformer encoding block is and the output is , then: ; where is the multi-head attention mechanism and is the feed-forward neural network. After layers of encoding, the text information concerned by the point cloud modality is obtained.
[0034] 4. Acquisition of text information in the image modality: Using the image feature vector as the weight and the text feature vector as the value, input into the The layer transformer encoding block performs encoding. Similar to obtaining text information in the point cloud modality, the text information of interest in the image modality is obtained. .
[0035] 5. Keyword text data generation: The text information of interest in the point cloud modality and the text information of interest in the image modality are fused. For example . The fused information is then fed into layer transformer decoding block. Assume the input of the layer transformer decoding block is , and the output is , then: ; where is the decoding layer operation, is the output of the encoding stage. Finally, the generated keyword text data is obtained. These generated keyword text data serve as an input to the model, providing key semantic information for model training.
[0036] 6. Adaptive layer number adjustment: This module sets up an adaptive layer number adjustment module consisting of an input layer, an output layer, and a fully connected layer. Let the correlation metric between the image data and the text data be , and the correlation metric between the lidar point cloud data and the text data be . The adjusted layer number is calculated through the fully connected layer: ; where , are the first and second weight matrices of the fully connected layer respectively, is the first bias term of the fully connected layer, is the sigmoid activation function, is the maximum layer number limit. This module dynamically adjusts the layer numbers of the transformer encoding block and the decoding block by learning the correlations between the image data and the text data, and between the lidar point cloud data and the text data . When the correlation is high, the number of layers is appropriately reduced to improve efficiency; when the correlation is low, the number of layers is increased to better mine potential information, thereby optimizing the effect of keyword generation.
[0037] After training is completed, only the multi-view image data , lidar point cloud data are input. After dynamically adjusting the layer number The KGM keyword generation module can generate keywords that match the input data. These keywords are used for pre-training of the target matching model, which effectively solves the problem of insufficient training data.
[0038] Furthermore, in the KGM keyword generation pre-training module, an adaptive layer adjustment module can be set up. This module dynamically adjusts the number of layers of the transformer encoding and decoding blocks according to the strength of the correlation by learning the correlation between image data and text data, and between lidar point cloud data and text data. This optimization step aims to improve the effect of keyword generation while reducing unnecessary computational overhead. When the correlation is high, the number of layers is appropriately reduced to improve efficiency; when the correlation is low, the number of layers is increased to better mine potential information.
[0039] This module is integrated into the Transformer encoding block / decoding block stacking architecture and mainly acts in the cross-modal data interaction stage.
[0040] The specific implementation is as follows: 1. Input feature correlation calculation layer, Location: The cross-modal interaction layer before the multimodal features are fed into the Transformer encoding block. Function: By calculating the cosine similarity matrix of image-text and point cloud-text, a dynamic correlation coefficient (ranging from 0.6 to 1.2) is generated to quantify the semantic association strength of image-text and point cloud-text.
[0041] Calculate the cross-attention score for image and text features and normalize it to the correlation coefficient; The lidar point cloud and text features generate weighted similarity through a gating mechanism. 2. Dynamic layer controller, Position: Embedded between residual connections in a stack of Transformer encoding blocks. The workflow includes: Relevance threshold judgment: When the correlation coefficient is > 0.9, the layer compression mode is activated, only one layer of coding blocks (Layer 0) and one layer of residual layer (Layer 1) are retained, and the calculation of the subsequent 4 layers is skipped; When the correlation coefficient is <0.8, the full 6-layer encoding block is enabled and an additional feature decoupling layer is inserted after Layer 3; Otherwise, activate layer compression mode, retaining only three coding layers and one residual layer. The residual connection path is dynamically selected through a learnable gating network.
[0042] Residual gating formula: ; Wherein, represents the number of layers actually enabled in the encoding block (dynamically adjusted according to the correlation coefficient ); represents a binary gating function (taking values of 0 or 1), which controls whether the th layer participates in the calculation; : the feed-forward calculation module of the th layer Transformer encoding block; represents the hidden state of the th layer encoding block.
[0043] 3. Adaptive linkage of the decoding block, Location: Synchronously adjusted in the cross-attention module of the decoding block layer.
[0044] Implementation method: The correlation coefficient output by the encoding block is transmitted to the decoding block through a shared weight matrix.
[0045] The decoding block dynamically adjusts according to the received coefficient: When the correlation is high (correlation coefficient > 0.9), only 1 layer of the decoding block is retained for feature mapping; When the correlation is low (correlation coefficient < 0.8), 3 layers of the decoding block are enabled to strengthen semantic reconstruction; Otherwise, 2 layers of the decoding block are enabled.
[0046] Through the above adaptive layer adjustment process, the performance and efficiency of the model can be further improved.
[0047] S2: Fuse the visual descriptor and the viewpoint descriptor into a multi-view fusion 3D expression feature vector; linearly transform the second keyword feature vector to obtain a third keyword feature vector; and concatenate the multi-view fusion 3D expression feature vector and the third keyword feature vector to obtain a first joint expression vector.
[0048] Specifically, the fusing of the visual descriptor and the viewpoint descriptor into a multi-view fusion 3D expression feature vector includes: Sum the product of all visual descriptors and the first fusion weight under any one view to obtain a second summation result; Sum the second summation results under all views to obtain the total visual descriptor; Sum the product of all viewpoint descriptors and the second fusion weight to obtain the total viewpoint descriptor; Add the total visual descriptor and the total viewpoint descriptor to obtain a multi-view fusion 3D expression feature vector.
[0049] Further, the linear transformation of the second keyword feature vector includes: Multiplying the second key feature vector by a linear weight matrix to obtain a fifth product; Adding the fifth product to a bias term to obtain the third keyword feature vector.
[0050] Furthermore, the concatenation of the multi-view fusion 3D expression feature vector and the third keyword feature vector includes: Multiplying the multi-view fusion 3D expression feature vector by a feature vector weight to obtain a second product; Adding the second product to the third keyword feature vector to obtain a first joint expression vector.
[0051] In this embodiment, the value of the feature vector weight is 3.
[0052] This can ensure the dominant position of the multi-view fusion 3D expression feature vector in the fusion process and give full play to its main description role for the target features.
[0053] S3: Arrange three layers of transformer encoding blocks. The input of the first layer of transformer encoding blocks is the first joint expression vector, and the input of the remaining two layers of transformer encoding blocks is the concatenated vector of the output of the previous layer and the third keyword feature vector. The output of the third layer of transformer encoding blocks is the three-level joint expression vector.
[0054] Specifically, the multi-head attention mechanism in the transformer encoding block is used to further integrate and optimize the vector input into the encoding block. Through three layers of transformer encoding blocks, the information of features is further mined and fused, gradually strengthening the fusion effect of multi-modal features.
[0055] Further, the concatenation of the output of the first / second layer of transformer encoding blocks and the third keyword feature vector includes: Multiplying the output of the first / second layer of transformer encoding blocks by a feature vector weight to obtain a third / fourth product; Adding the third / fourth product to the third keyword feature vector to obtain the input of the second / third layer of transformer encoding blocks.
[0056] The advantages of performing three concatenations and transformer encoding are: Step - by - step enhanced multi - modal feature fusion: During the process of three - cycle splicing and encoding, starting from the second layer, the keyword feature vector is weighted - spliced with the output of the previous - layer Transformer encoding block, and the information is further integrated and optimized through the Transformer encoding block. This process gradually enhances the fusion effect of multi - modal features, enabling the finally generated joint expression vector to more comprehensively and accurately reflect the multi - modal features of the target.
[0057] Improve the accuracy and robustness of target matching: Through multiple splicings and encodings, the model can learn richer feature representations, so that in the target - matching stage, it can more accurately identify the part that is most similar to the keyword feature vector to be detected. This helps to improve the accuracy and robustness of target matching, especially in cross - domain multi - unmanned equipment application scenarios, where good matching effects can still be maintained in the face of challenges such as large perspective differences and high data heterogeneity.
[0058] Reduce training time: Although the three - cycle splicing increases the computational amount, thanks to the efficient processing ability of the Transformer encoding block and the weight setting (with relatively small weights) of the keyword feature vector during the splicing process, overall, this method reduces unnecessary computational overhead while ensuring the fusion effect, thus reducing the training time to a certain extent.
[0059] S4: Find the part with the highest cosine similarity between the three - level joint expression vector and the first keyword feature vector as the target - matching result.
[0060] Specifically, this step includes: Split the three - level joint expression vector into multiple first - level feature blocks; Split the first keyword feature vector into multiple second - level feature blocks; Calculate the cosine similarity between any one first - level feature block and each second - level feature block to obtain a row / column of similarity values; Traverse all first - level feature blocks, and combine the row / column of similarity values calculated according to each first - level feature block into a similarity matrix; Take the first - level feature block with the highest similarity value to the second - level feature block in the similarity matrix as the target - matching result.
[0061] After the model training is completed, the feature vector is processed by the trained three - layer Transformer encoding block and then performs target matching with the target to be recognized, and the final target - matching result can be obtained, which greatly improves the convenience and efficiency of the application.
[0062] The target - matching method based on cross - domain multi - modal fusion encoding provided in this embodiment has the following beneficial effects: 1. The proposed KGM keyword generation pre-training module can generate corresponding keyword text data only using image and point cloud data through training, which is used to train the target matching model, solves the problem of insufficient data volume required for training the MLIP multi-modal fusion coding target matching model, ensures the high generalization ability of the model, and enables the target to be recognized to be added at any time without retraining.
[0063] 2. The multi-modal fusion coding method with keyword prompting is adopted, which effectively realizes the deep feature fusion of three modalities: image, point cloud, and text, reduces the training time at the same time, and solves the problem of multi-modal data heterogeneity.
[0064] 3. By fusing the visual descriptor and the viewpoint descriptor to obtain the multi-view fusion 3D expression feature vector, it can handle the problems of huge perspective differences and cross-domain when collecting data from multiple unmanned equipment across domains, and improves the accuracy and reliability of target matching.
[0065] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0066] The above-described embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A target matching method based on cross-domain multi-modal fusion coding, characterized in that, Including: S1: Obtain keyword text data to be detected, multi-view image data of multiple unmanned equipment, and lidar point cloud data; Split the image feature matrix obtained by encoding the multi-view image data to obtain visual descriptors; perform visible point screening on the per-point feature vectors obtained by encoding the lidar point cloud data to obtain viewpoint descriptors; convert the first keyword feature vector extracted from the keyword text data into a digital form to obtain a second keyword feature vector; S2: Fuse the visual descriptors and the viewpoint descriptors into a multi-view fusion 3D expression feature vector; perform a linear transformation on the second keyword feature vector to obtain a third keyword feature vector; and concatenate the multi-view fusion 3D expression feature vector and the third keyword feature vector to obtain a first joint expression vector; S3: Arrange three layers of transformer encoding blocks. The input of the first layer of transformer encoding block is the first joint expression vector, and the input of the remaining two layers of transformer encoding blocks is the concatenation vector of the output of the previous layer and the third keyword feature vector. The third layer of transformer encoding block outputs a three-level joint expression vector; S4: Find the part with the highest cosine similarity between the three-level joint expression vector and the first keyword feature vector as the target matching result.
2. The target matching method based on cross-domain multimodal fusion coding according to claim 1, characterized in that, Encoding the multi-view image data includes: passing the image data of each view through a convolutional neural network respectively to obtain an image feature matrix corresponding to each view.
3. The target matching method based on cross-domain multimodal fusion coding according to claim 1, wherein Splitting the image feature matrix includes: performing slicing or pooling operations on the image feature matrix corresponding to any one view in a specific dimension to obtain multiple visual descriptors corresponding to the view.
4. The object matching method based on cross-domain multimodal fusion coding according to claim 1, characterized in that, Encoding the lidar point cloud data includes: passing the lidar point cloud data through PointNet for point cloud data encoding to obtain a per-point feature vector for each point.
5. The target matching method based on cross-domain multimodal fusion coding according to claim 1, wherein, Performing visible point screening on the per-point feature vectors includes: Calculate the Euclidean distance between a point in the lidar point cloud data and any other point, and multiply the Euclidean distance by the weight matrix to obtain a first product; i Traverse the remaining points in the lidar point cloud data and calculate the points i The first product corresponding to all the other points, and sum all the calculated first products; The sum result is passed through an activation function to obtain the visibility score of point i ; Compare the visibility score of the point i with a set threshold. When the visibility score of the point i is greater than or equal to the set threshold, use the point-by-point feature vector of the point i as a viewpoint descriptor; otherwise, discard it. Traverse all points in the lidar point cloud data, compare the visibility scores of all points with a set threshold to obtain multiple viewpoint descriptors.
6. The target matching method based on cross-domain multimodal fusion coding according to claim 1, characterized in that The conversion of the first keyword feature vector extracted from the keyword text data into a digital form includes: Pass the keyword text data through a word vector model to obtain the first keyword feature vector; Split the first keyword feature vector into a sequence composed of individual numbers to obtain a second keyword feature vector.
7. The object matching method based on cross-domain multimodal fusion coding according to claim 1, characterized in that, The fusion of the visual descriptors and the viewpoint descriptors into a multi-view fusion 3D expression feature vector includes: Sum the products of all visual descriptors under any one view and the first fusion weight to obtain a second summation result; Sum the second summation results under all views to obtain the total visual descriptor; Sum the products of all viewpoint descriptors and the second fusion weight to obtain the total viewpoint descriptor; Add the total visual descriptor and the total viewpoint descriptor to obtain a multi-view fusion 3D expression feature vector.
8. The target matching method based on cross-domain multimodal fusion coding according to claim 1, wherein The concatenation of the multi-view fusion 3D expression feature vector and the third keyword feature vector includes: Multiply the multi-view fusion 3D expression feature vector by the feature vector weight to obtain a second product; Add the second product to the third keyword feature vector to obtain the first joint expression vector.
9. The target matching method based on cross-domain multimodal fusion coding according to claim 1, characterized in that, The concatenation of the output of the first-layer / second-layer Transformer encoding block and the third keyword feature vector includes: Multiply the output of the first-layer / second-layer Transformer encoding block by the feature vector weight to obtain the third product / fourth product; Add the third product / fourth product to the third keyword feature vector to obtain the input of the second-layer / third-layer Transformer encoding block.
10. The target matching method based on cross-domain multi-modal fusion coding according to claim 1, characterized in that S4 Includes: Split the three-level joint expression vector into multiple first feature blocks; Split the first keyword feature vector into multiple second feature blocks; Calculate the cosine similarity between any one first feature block and each second feature block to obtain a row / column of similarity values; Traverse all the first feature blocks, and combine the row / column of similarity values calculated according to each first feature block into a similarity matrix; Take the first feature block with the highest similarity value to the second feature block in the similarity matrix as the target matching result.
Citation Information
Patent Citations
Data pair acquisition method, device and equipment, server, cluster of server and medium
CN116911353A
Digital twin modeling method and device based on multiple modes and feature fusion method
CN118552962A
Multi-modal data fusion sensing method and system based on space-time correlation
CN119418300A
Method and system for constructing place recognition model based on text-point cloud matching
CN119760647A
Position identification model construction method and system based on multi-view cross-modal matching
CN119887911A