Missing information supplementary modeling method based on multi-modal fusion
By feature encoding, fusion and reconstruction of unmanned equipment RGB images and point cloud images, the problem of information loss in traditional multimodal fusion technology is solved, more efficient information processing and decision-making support is achieved, and the perception and modeling capabilities of unmanned equipment are improved.
Patent Information
- Application Number
- CN202510511969.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Traditional multimodal information fusion technology fails to fully utilize and integrate the intrinsic connections and complementarity between different modal data, resulting in the loss of useful information and affecting the perception and decision-making support capabilities of unmanned equipment in complex environments.
By acquiring unmanned RGB images and point cloud images, features are extracted using color components and spatial position encoding, feature fusion and reconstruction are performed, frequency domain features are decomposed, and the source domain and target domain are aligned to perform point cloud supplementary feature extraction and three-dimensional surface reconstruction.
It has improved the perception and decision-making support capabilities of unmanned equipment in complex environments, enhanced the perception and modeling accuracy of the model, and promoted the advancement of unmanned equipment information processing technology.
Smart Images

Figure CN120279191A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of unmanned equipment information processing, and particularly to a method for supplementing missing information modeling based on multimodal fusion. Background Art
[0002] Traditional multimodal information fusion technologies have some limitations. These technologies often fail to fully utilize and integrate the internal connections and complementarities between different modal data when processing data of different modalities, resulting in the loss of some useful information. In addition, due to the relative independence of each modal information, the correlation between them is poor, which is particularly obvious when performing cross-modal tasks because of the lack of in-depth consideration of the inertia between different modal data; moreover, when applied to unmanned equipment, due to these limitations of traditional fusion methods, it is impossible to accurately and effectively perform subsequent tasks (such as modeling tasks, etc.), thus unable to improve the perception ability and decision-making support ability of unmanned equipment in complex environments. Summary of the Invention
[0003] Based on this, it is necessary to provide a method for supplementing missing information modeling based on multimodal fusion, and this method includes: S1: Obtain the RGB image and point cloud image of the unmanned equipment; S2: Encode the RGB image using color components to obtain the first feature; encode the point cloud image using spatial position encoding to obtain the second feature; S3: Align the feature spaces of the first feature and the second feature, and perform feature fusion to obtain a joint feature; perform image reconstruction based on the joint feature to obtain a reconstructed RGB image; S4: Perform frequency domain transformation on the point cloud image and the reconstructed RGB image respectively, and decompose them into color high-frequency features, color low-frequency features, point cloud high-frequency features, and point cloud low-frequency features; S5: Use the color high-frequency features and point cloud high-frequency features as the source domain, use the color low-frequency features and point cloud low-frequency features as the target domain, align and fuse the source domain and the target domain to obtain point cloud supplementary features; S6: Extract local geometric and color features and perform three-dimensional surface reconstruction on the point cloud supplementary features, combine the extracted features and the reconstructed three-dimensional surface to obtain a three-dimensional model.
[0004] Preferably, the encoding of the RGB image using color components includes: Decompose the RGB image into three two-dimensional matrices according to color channels; For each two-dimensional matrix, encode it according to the distribution characteristics of its corresponding pixel values to obtain the corresponding encoded matrix; Fuse the three encoded matrices to obtain a high-dimensional feature matrix; Perform global average pooling on the high-dimensional feature matrix to obtain the first feature.
[0005] Preferably, the encoding of the point cloud image using spatial position encoding includes: Perform spatial position encoding on the point cloud image to obtain the high-dimensional representation of each point; Input the high-dimensional representations of each point into a multi-layer perceptron to output the feature representation of each point; For any point, obtain the neighborhood set through the K-nearest neighbor algorithm or ball radius search, and aggregate the feature representations of the points in the neighborhood set to obtain the local feature representation of the corresponding point; Aggregate the local feature representations of all points in the point cloud image to obtain the second feature.
[0006] Preferably, the alignment of the feature spaces of the first feature and the second feature includes: Project the first feature into the shared space through a convolutional neural network; Project the second feature into the shared space through a multi-layer perceptron; Use the nearest neighbor interpolation method to upsample the projected first feature to the dimension of the second feature to align the dimensions of the first feature and the second feature.
[0007] Preferably, the process of obtaining the joint feature is: Perform linear fusion on the upsampled first feature and the projected second feature to obtain the joint feature.
[0008] Preferably, the image reconstruction based on the joint feature includes: Calculate the adjacency matrix between the RGB image and the point cloud image through spatial neighbors; Input the adjacency matrix and the joint feature into a multi-layer perceptron to obtain the preliminary reconstruction feature; Pass the preliminary reconstruction feature through the Linear function to obtain the reconstruction weight; Pass the reconstruction weight through the softmax activation function and multiply it element-wise with the preliminary reconstruction feature to obtain the reconstruction feature; Convert the reconstruction feature into the reconstructed RGB image through a decoder.
[0009] Preferably, in S3, it further includes: constructing a total loss function based on the weighted combination of the modality alignment loss, distribution alignment loss, and image reconstruction loss between the first feature and the second feature, and minimizing the total loss function to optimize the reconstruction process; The modality alignment loss is based on the first feature and the second feature and is calculated using the NT-Xent loss function; The distribution alignment loss is based on the first feature and the second feature and is calculated using the maximum mean discrepancy loss function; The image reconstruction loss is based on the first feature and the second feature, and is calculated using the mean square error.
[0010] Preferably, S4 includes: Perform wavelet transform on the reconstructed RGB image using the 2D-BTSA function, and decompose it into a first vector matrix of different frequency bands; Perform wavelet transform on the point cloud image using the 3D-TDMF-FE function, and decompose it into a second vector matrix of different frequency bands; Take the first vector matrix in the high-frequency band as the color high-frequency feature, and take the second vector matrix in the low-frequency band as the color low-frequency feature; Take the second vector matrix in the high-frequency band as the point cloud high-frequency feature, and take the second vector matrix in the low-frequency band as the point cloud low-frequency feature.
[0011] Preferably, the aligning and fusing the source domain and the target domain to obtain the point cloud supplementary feature includes: Take any one feature in the source domain as the anchor feature, take the feature in the target domain that comes from the same image as the anchor feature as the positive sample feature, and take the features in the target domain other than the positive sample feature as the negative sample features; calculate the triplet loss based on the anchor feature, the positive sample feature, and the negative sample features; Align the source domain and the target domain by minimizing the triplet loss to obtain the aligned features; Input the aligned features into a multi-layer perceptron to obtain the point cloud supplementary feature.
[0012] Preferably, S6 includes: The point cloud supplementary feature includes: the three-dimensional coordinates of the points in the point cloud and the color feature, and the color feature includes: the red channel component, the green channel component, and the blue channel component; For any point in the point cloud image, search for its neighborhood point set using the sphere radius; Based on the point and its neighborhood point set, calculate the local normal vector of the point through the principal component analysis method; calculate the curvature of the point based on the ratio of the eigenvalues of the covariance matrix of the neighborhood points; Calculate the color gradient based on the color feature of the point and the three-dimensional coordinates of the point; Integrate the three-dimensional coordinates of the point, the local normal vector, the curvature, the red channel component, the green channel component, the blue channel component, and the color gradient to obtain a high-dimensional feature vector; Input the high-dimensional feature vector into a multi-layer perceptron to generate the high-dimensional feature of the point; Based on each neighborhood point of the point and the distance from the neighborhood point to the neighborhood center point, calculate the second joint feature using point cloud neighborhood interpolation; Input the second joint feature and the local normal vector into the Poisson surface reconstruction function to generate the local three-dimensional surface corresponding to the point; Traverse all points in the point cloud image, obtain the high-dimensional features of all points, and generate the three-dimensional surface of the point cloud; Combine the high-dimensional features of all points and the three-dimensional surface of the point cloud to obtain a three-dimensional model.
[0013] Beneficial effects: By fusing multi-modal information, the potential connections between different modal data can be better explored and utilized, thereby enhancing the perception ability and decision support ability of unmanned equipment in complex environments. This new method is expected to promote the progress of unmanned equipment information processing technology and provide support for achieving a higher level of autonomy and intelligence. Description of the Drawings
[0014] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 It is a flowchart of the missing information supplementary modeling method based on multi-modal fusion in the embodiments of the present application. Detailed Embodiments
[0016] To make the above-mentioned objects, features, and advantages of the present application more obvious and understandable, the following will give a detailed description of the specific embodiments of the present application with reference to the drawings. Many specific details are set forth in the following description to fully understand the present application. However, the present application can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0017] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0018] As Figure 1 shown, this embodiment provides a missing information supplementary modeling method based on multi-modal fusion, and the method includes: S1: Obtain the RGB image and the point cloud image of the unmanned equipment.
[0019] S2: Encode the RGB image using color components to obtain the first feature; encode the point cloud image using spatial position encoding to obtain the second feature.
[0020] Further, the encoding of the RGB image using color components includes: Decompose the RGB image into three two-dimensional matrices according to color channels; For each two-dimensional matrix, perform encoding according to the distribution characteristics of its corresponding pixel values to obtain a corresponding encoded matrix; Linearly fuse the three encoded matrices to obtain a high-dimensional feature matrix; Perform global average pooling on the high-dimensional feature matrix to obtain the first feature.
[0021] The encoding method based on the pixel features of color components can effectively extract the significant information of the image and enhance the robustness and accuracy of task processing.
[0022] Further, the encoding of the point cloud image using spatial position encoding includes: Perform spatial position encoding on the point cloud image to obtain high-dimensional representations of each point; Input the high-dimensional representations of each point into a multi-layer perceptron to output the feature representation of each point; For any point, obtain a neighborhood set through the K-nearest neighbor algorithm or sphere radius search, and aggregate the feature representations of the points in the neighborhood set (aggregation can be performed by max pooling or weighted summation) to obtain the local feature representation of the corresponding point; Aggregate (aggregation can be performed by global pooling or attention mechanism) the local feature representations of all points in the point cloud image to obtain the second feature.
[0023] S3: Align the feature spaces of the first feature and the second feature, and perform feature fusion to obtain a joint feature; perform image reconstruction based on the joint feature to obtain a reconstructed RGB image.
[0024] Further, the alignment of the feature spaces of the first feature and the second feature includes: Project the first feature into the shared space through a convolutional neural network; Project the second feature into the shared space through a multi-layer perceptron; Use the nearest neighbor interpolation method to upsample the projected first feature to the dimension of the second feature to align the dimensions of the first feature and the second feature.
[0025] Further, the process of obtaining the joint feature is: Linearly fuse the upsampled first feature and the projected second feature to obtain the joint feature.
[0026] Even further, the image reconstruction based on the joint feature includes: Calculate the adjacency matrix between the RGB image and the point cloud image through spatial neighbors; Input the adjacency matrix and the combined features into a multi-layer perceptron to obtain preliminary reconstructed features; Pass the preliminary reconstructed features through a Linear function to obtain reconstruction weights; Pass the reconstruction weights through a softmax activation function and perform element-wise multiplication with the preliminary reconstructed features to obtain reconstructed features; Convert the reconstructed features into a reconstructed RGB image through a decoder.
[0027] Use a decoder to perform deconvolution operations, and fusing context features can enhance details.
[0028] In this embodiment, step S3 further includes: constructing a total loss function by weighted combination of the modality alignment loss, distribution alignment loss, and image reconstruction loss between the first feature and the second feature, and minimizing the total loss function to optimize the reconstruction process; The modality alignment loss is based on the first feature and the second feature and is calculated using the NT-Xent loss function; The distribution alignment loss is based on the first feature and the second feature and is calculated using the maximum mean discrepancy loss function; The image reconstruction loss is based on the first feature and the second feature and is calculated using the mean squared error.
[0029] S4: Perform frequency domain transformation on the point cloud image and the reconstructed RGB image respectively, and decompose them into color high-frequency features, color low-frequency features, point cloud high-frequency features, and point cloud low-frequency features.
[0030] Specifically, this step includes: Perform wavelet transform on the reconstructed RGB image using a 2D-BTSA function (two-dimensional texture spectrum analyzer, emphasizing the extraction of multi-dimensional texture information of two-dimensional images) and decompose it into a first vector matrix of different frequency bands; specifically extract features from the following three aspects: Horizontal detail function: Capture horizontal edges and texture information in the image; Vertical detail function: Capture vertical edges and texture information in the image; Diagonal detail function: Capture diagonal edges and texture information in the image.
[0031] Perform wavelet transform on the point cloud image using a 3D-TDMF-FE function (three-dimensional multi-frequency domain feature extractor, highlighting the ability to capture multi-frequency features of three-dimensional point clouds, including reflection intensity) and decompose it into a second vector matrix of different frequency bands; specifically extract features from the following three aspects: Horizontal detail function: Capture horizontal edges and texture information of the point cloud; Vertical detail function: Capture vertical edges and texture information of the point cloud; Reflection intensity function: Capturing the reflection intensity information of the point cloud.
[0032] The first vector matrix in the high-frequency band is used as the color high-frequency feature, representing local detail information (such as edges, textures, and details); the second vector matrix in the low-frequency band is used as the color low-frequency feature, representing global characteristics (such as shapes and illumination distributions); The second vector matrix in the high-frequency band is used as the point cloud high-frequency feature, representing local geometric information (such as surfaces and point cloud density changes); the second vector matrix in the low-frequency band is used as the point cloud low-frequency feature, representing the overall shape characteristics and illumination reflection intensity.
[0033] The vector matrix in the high-frequency band usually corresponds to the detail coefficients of the wavelet transform, and these coefficients reflect the rapid changes of the signal in the local area (such as edge, texture, and detail information); The vector matrix in the low-frequency band usually corresponds to the approximation coefficients, and these coefficients reflect the overall trend and global structure of the signal (such as overall shape, illumination distribution, and reflection intensity).
[0034] Through frequency-domain analysis, the source domain and target domain of the point cloud and RGB image are divided. Two-dimensional and three-dimensional wavelet transforms are used to extract local and global information from multiple angles. Through feature alignment and optimization, efficient modality fusion is achieved, providing high-quality feature inputs for subsequent tasks.
[0035] S5: Use the color high-frequency feature and the point cloud high-frequency feature as the source domain, use the color low-frequency feature and the point cloud low-frequency feature as the target domain, align and fuse the source domain and the target domain to obtain the point cloud supplementary feature.
[0036] Specifically, the aligning and fusing the source domain and the target domain to obtain the point cloud supplementary feature includes: Take any one feature in the source domain as the anchor feature, take the feature in the target domain that comes from the same image as the anchor feature as the positive sample feature, and take the features in the target domain other than the positive sample feature as the negative sample features; calculate the triplet loss based on the anchor feature, the positive sample feature, and the negative sample features; Align the source domain and the target domain by minimizing the triplet loss to obtain the aligned features; Input the aligned features into a multi-layer perceptron to obtain the point cloud supplementary feature.
[0037] Furthermore, the calculation formula of the triplet loss is: ; ; ; Among them, represents the triplet loss; represents the Euclidean distance from the anchor feature to the positive sample feature; Represents the Euclidean distance from the anchor feature to the negative sample feature; Represents the boundary, ensuring that the distance from the anchor feature to the negative sample feature is at least greater than the distance to the positive sample feature ; When holds, the loss is 0; when holds, the loss is positive, indicating that the alignment goal has not been achieved and further optimization is required.
[0038] Training process: In each batch, triplets are constructed in the following way: 1. Anchor sample: Randomly sample from the source domain features , 2. Positive sample selection: Find the one in the target domain that matches the anchor feature , 3. Negative sample selection: Randomly sample from the target domain, which does not match the anchor feature ; Calculate the loss for each triplet , and average it within the current batch: ; Among them, represents the triplet loss of the i-th triplet, is the number of triplets in the current batch; Use backpropagation to optimize the network weights and minimize the overall loss function.
[0039] After training is completed, evaluate the alignment effect of the source domain and target domain features 1. Visualize the alignment result: Use t-SNE to map the features to a two-dimensional space and observe whether the distributions of the source domain and target domain features tend to be consistent. 2. Feature similarity measurement: Calculate the average distance from the anchor feature to the positive sample feature and the average distance to the negative sample , and verify .
[0040] Through the triplet loss function, the features of the source domain and target domain can be efficiently aligned in the shared space. By dynamically adjusting the distance differences between the anchor, positive sample, and negative sample, the relevance and separability of high-frequency and low-frequency features are ensured, providing a higher-quality feature basis for subsequent tasks.
[0041] In the supplemented feature of the fused point cloud, RGB supplementation enriches the surface texture and color attributes of the point cloud, making the model more realistic in visualization and texture mapping; RGB information provides semantic context for each point in the point cloud, such as the distinction of object edges or materials; The point cloud provides a three-dimensional geometric structure, and RGB supplements texture details. The combination of the two can improve the integrity and expressiveness of modeling.
[0042] For the results of aligning RGB and point cloud in the source domain (local features) and the target domain (global features, low-frequency information), further utilize the fused features to complete 3D modeling. The source domain provides local texture and geometric details, and the target domain provides global structure and illumination information. The combination of the two lays the foundation for accurate 3D model construction.
[0043] S6: Extract local geometric and color features from the supplementary features of the point cloud and perform 3D surface reconstruction. Combine the extracted features and the reconstructed 3D surface to obtain a 3D model.
[0044] Specifically, this step includes: The supplementary features of the point cloud include: the three-dimensional coordinates of the points in the point cloud and the color features. The color features include: the red channel component, the green channel component, and the blue channel component; each channel component is further divided into a high-frequency part and a low-frequency part. The high-frequency part provides local texture details, and the low-frequency part provides global color and illumination information.
[0045] For any point in the point cloud image, use the spherical radius to search for its neighborhood point set.
[0046] Based on the point and its neighborhood point set, calculate the local normal vector of the point through the principal component analysis method. The calculation formula is: ; Among them, represents the local normal vector of the i-th point; represents the neighborhood point set of the i-th point; represents the i-th neighborhood point; n represents the number of neighborhood points; represents the square of the Euclidean norm.
[0047] Calculate the curvature of the point based on the ratio of the eigenvalues of the covariance matrix of the neighborhood points. The calculation formula is: ; Among them, represents the curvature of the i-th point; , , respectively represent the eigenvalues of the covariance matrix of different neighborhood points, satisfying .
[0048] Calculate the color gradient based on the color feature of the point and the three-dimensional coordinates of the point. The calculation formula is: ; Among them, represents the color gradient of the i-th point; Denotes partial derivative; Denotes the red channel component of the i-th point; Denotes the green channel component of the i-th point; Denotes the blue channel component of the i-th point; Denotes the spatial abscissa of the i-th point; Denotes the spatial ordinate of the i-th point; Denotes the spatial vertical coordinate of the i-th point.
[0049] Integrate the three-dimensional coordinates, local normal vector, curvature, red channel component, green channel component, blue channel component and color gradient of the point to obtain a high-dimensional feature vector; Input the high-dimensional feature vector into a multi-layer perceptron to generate the high-dimensional feature of the point.
[0050] Based on each neighborhood point of the point and the distance from the neighborhood point to the neighborhood center point, use point cloud neighborhood interpolation to calculate the second joint feature and generate a continuous surface. The calculation formula is: ; Where, Denotes the second joint feature of the i-th point; Denotes the interpolation weight, which is the distance from the j-th neighborhood point to the neighborhood center point.
[0051] In this embodiment, it further includes: the generated surface can be globally corrected by using the low-frequency part provided by the target domain to ensure the overall shape and lighting consistency of the model.
[0052] Specifically, the global correction process is a key step in 3D modeling, aiming to ensure that the shape, lighting and structure of the model as a whole match the expected global features. The correction process is as follows: Define the global correction target: The goal of global correction is to make the preliminarily reconstructed model match the low-frequency part extracted from the data in terms of shape, lighting and structure. This may involve defining one or more optimization goals, such as minimizing the difference between the model and the low-frequency part, or maximizing the global consistency of the model.
[0053] Apply the optimization algorithm: To achieve global correction, it is usually necessary to apply an optimization algorithm to adjust the surface of the model. The algorithms include: parameterize the model surface, parameterize the model surface into a set of adjustable parameters (such as the positions of control points); define an energy function, which measures the difference between the model surface and the low-frequency part. The energy function may include multiple terms (such as shape consistency term, lighting consistency term and smoothness term); optimize the energy function, and use an algorithm (such as gradient descent, simulated annealing or genetic algorithm) to minimize the energy function, thereby adjusting the parameters of the model surface.
[0054] Iterative adjustment: Global correction is usually an iterative process that requires adjusting the model surface multiple times until the preset global consistency is achieved; in each iteration, the optimization algorithm updates the parameters of the model surface based on the difference between the current model surface and the low-frequency part.
[0055] Evaluation and refinement: During the iterative adjustment process, it is necessary to regularly evaluate the global consistency of the model and refine it as needed, including: Visual evaluation, checking whether the shape, lighting, and structure of the model match the expected global features through visualization tools; Quantitative evaluation, using quantitative metrics (such as mean squared error or structural similarity index) to evaluate the difference between the model and the low-frequency part.
[0056] Input the second joint feature and the local normal vector into the Poisson surface reconstruction function to generate the local three-dimensional surface corresponding to the points.
[0057] In this embodiment, texture mapping and detail enhancement are performed on the model surface through the high-frequency part: 1. Texture mapping: Project the high-frequency part onto the reconstructed three-dimensional surface to generate a textured model.
[0058] 2. Detail enhancement: Introduce a high-frequency loss function Optimize the details: ; where, represents the high-frequency part of the i-th point; M represents the number of points; represents the predicted high-frequency part of the i-th point; represents the square of the Euclidean norm.
[0059] Traverse all points in the point cloud image to obtain the high-dimensional features of all points and generate the three-dimensional surface of the point cloud; Combine the high-dimensional features of all points and the three-dimensional surface of the point cloud to obtain a three-dimensional model.
[0060] The finally output three-dimensional model includes: 1. Precise geometric information: The spatial shape constructed from the point cloud geometric data; 2. High-fidelity texture information: The color and texture supplemented by RGB data.
[0061] This model can be used for tasks such as scene reconstruction, target recognition, and visualization display.
[0062] The method for supplementing missing information modeling based on multi-modal fusion provided in this embodiment has the following beneficial effects: This method can model using the point cloud (3D) information supplemented by RGB, enhance the expression ability of the data for the point cloud image information, and supplement the spatial depth and three-dimensional structure. By a more refined method, different modal data are integrated to improve the accuracy and efficiency of information processing. Through the fusion of multi-modal information, the potential connections between different modal data can be better explored and utilized, thereby enhancing the perception ability and decision support ability of unmanned equipment in complex environments. This new method is expected to promote the progress of unmanned equipment information processing technology and provide support for achieving a higher level of autonomy and intelligence.
[0063] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0064] The above-described embodiments only express several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for supplementing and modeling missing information based on multimodal fusion, characterized in that Including: S1: Obtain the RGB image and point cloud image of the unmanned equipment; S2: Encode the RGB image using color components to obtain the first feature; Encode the point cloud image using spatial position encoding to obtain the second feature; S3: Align the feature spaces of the first feature and the second feature, and perform feature fusion to obtain the joint feature; Based on the joint feature, perform image reconstruction to obtain the reconstructed RGB image; S4: Perform frequency domain transformation on the point cloud image and the reconstructed RGB image respectively, and decompose them into color high-frequency features, color low-frequency features, point cloud high-frequency features, and point cloud low-frequency features; S5: Use the color high-frequency features and point cloud high-frequency features as the source domain, use the color low-frequency features and point cloud low-frequency features as the target domain, align and fuse the source domain and the target domain to obtain the point cloud supplementary features; S6: Extract local geometric and color features and perform 3D surface reconstruction on the point cloud supplementary features, and combine the extracted features and the reconstructed 3D surface to obtain the 3D model.
2. The method for supplementing and modeling missing information based on multimodal fusion according to claim 1, wherein The encoding of the RGB image using color components includes: Decompose the RGB image into three two-dimensional matrices according to color channels; For each two-dimensional matrix, encode it according to the distribution characteristics of its corresponding pixel values to obtain the corresponding encoded matrix; Fuse the three encoded matrices to obtain a high-dimensional feature matrix; Perform global average pooling on the high-dimensional feature matrix to obtain the first feature.
3. The method for supplementing and modeling missing information based on multimodal fusion according to claim 1, wherein The encoding of the point cloud image using spatial position encoding includes: Perform spatial position encoding on the point cloud image to obtain the high-dimensional representation of each point; Input the high-dimensional representation of each point into a multi-layer perceptron to output the feature representation of each point; For any point, obtain the neighborhood set through the K-nearest neighbor algorithm or sphere radius search, and aggregate the feature representations of the points in the neighborhood set to obtain the local feature representation of the corresponding point; Aggregate the local feature representations of all points in the point cloud image to obtain the second feature.
4. The method for supplementing and modeling missing information based on multimodal fusion according to claim 1, wherein The alignment of the feature spaces of the first feature and the second feature includes: Project the first feature into the shared space through a convolutional neural network; Project the second feature into the shared space through a multi-layer perceptron; Use the nearest neighbor interpolation method to upsample the projected first feature to the dimension of the second feature to align the dimensions of the first feature and the second feature.
5. The method for supplementing and modeling missing information based on multimodal fusion according to claim 4, wherein The process of obtaining the joint feature is: Linearly fuse the upsampled first feature and the projected second feature to obtain the joint feature.
6. The method for supplementing and modeling missing information based on multimodal fusion according to claim 1, wherein The image reconstruction based on the joint feature includes: Calculate the adjacency matrix between the RGB image and the point cloud image through spatial neighborhood; Input the adjacency matrix and the joint feature into a multi-layer perceptron to obtain the preliminary reconstruction feature; Pass the preliminary reconstruction feature through the Linear function to obtain the reconstruction weight; Pass the reconstruction weight through the softmax activation function, and multiply it element-wise with the preliminary reconstruction feature to obtain the reconstruction feature; Convert the reconstruction feature into the reconstructed RGB image through a decoder.
7. The method for supplementing and modeling missing information based on multimodal fusion according to claim 1, wherein In S3, it also includes: Based on the modal alignment loss, distribution alignment loss, and image reconstruction loss between the first feature and the second feature, perform weighted combination to construct the total loss function, and minimize the total loss function to optimize the reconstruction process; The modality alignment loss is based on the first feature and the second feature, and is calculated using the NT-Xent loss function; The distribution alignment loss is based on the first feature and the second feature, and is calculated using the maximum mean discrepancy loss function; The image reconstruction loss is based on the first feature and the second feature, and is calculated using the mean squared error.
8. The method for supplementing and modeling missing information based on multimodal fusion according to claim 1, wherein S4 It includes: Performing wavelet transform on the reconstructed RGB image using the 2D-BTSA function to decompose it into a first vector matrix of different frequency bands; Performing wavelet transform on the point cloud image using the 3D-TDMF-FE function to decompose it into a second vector matrix of different frequency bands; Taking the first vector matrix of the high-frequency band as the color high-frequency feature, and taking the second vector matrix of the low-frequency band as the color low-frequency feature; Taking the second vector matrix of the high-frequency band as the point cloud high-frequency feature, and taking the second vector matrix of the low-frequency band as the point cloud low-frequency feature.
9. The method for supplementing and modeling missing information based on multimodal fusion according to claim 1, wherein The aligning and fusing the source domain and the target domain to obtain the point cloud supplementary feature includes: Taking any feature in the source domain as the anchor feature, taking the feature in the target domain that comes from the same image as the anchor feature as the positive sample feature, and taking the features in the target domain other than the positive sample feature as the negative sample features; calculating the triplet loss based on the anchor feature, the positive sample feature, and the negative sample features; Aligning the source domain and the target domain by minimizing the triplet loss to obtain the aligned features; Inputting the aligned features into a multi-layer perceptron to obtain the point cloud supplementary feature.
10. The method for supplementing and modeling missing information based on multimodal fusion according to claim 1, wherein S6 includes: The point cloud supplementary feature includes: the three-dimensional coordinates of the points in the point cloud and the color feature, and the color feature includes: the red channel component, the green channel component, and the blue channel component; For any point in the point cloud image, using the spherical radius to search for its neighborhood point set; Based on the point and its neighborhood point set, calculating the local normal vector of the point through the principal component analysis method; calculating the curvature of the point based on the ratio of the eigenvalues of the covariance matrix of the neighborhood points; Calculating the color gradient based on the color feature of the point and the three-dimensional coordinates of the point; Integrating the three-dimensional coordinates of the point, the local normal vector, the curvature, the red channel component, the green channel component, the blue channel component, and the color gradient to obtain a high-dimensional feature vector; Inputting the high-dimensional feature vector into a multi-layer perceptron to generate the high-dimensional feature of the point; Based on each neighborhood point of the point and the distance from the neighborhood point to the neighborhood center point, calculating the second joint feature using point cloud neighborhood interpolation; Inputting the second joint feature and the local normal vector into the Poisson surface reconstruction function to generate the local three-dimensional surface corresponding to the point; Traversing all the points in the point cloud image to obtain the high-dimensional features of all the points and generate the three-dimensional surface of the point cloud; Combining the high-dimensional features of all the points and the three-dimensional surface of the point cloud to obtain a three-dimensional model.
Citation Information
Patent Citations
Adaptive method for point cloud classification field based on comparative learning
CN117454223A
Construction method and device of 3D target detection model based on LiDAR point cloud and RGB image
CN118212405A
Frequency and space mixed domain multi-modal fusion three-dimensional target detection framework building method
CN118506352A
Three-dimensional face defect complementing method based on multi-modal deep learning
CN118657910A
Deep multimodal cross-layer intersecting fusion method, terminal device, and storage medium
US11120276B1
Cited By
Business process missing activity repairing method and system based on missing perception and medium
CN121543036A
Method, system, and media for missing activity repair of business process based on missing awareness
CN121543036B