Multimodal Autoencoder Parts Intelligent Detection Method and Device for Fusing 3D Point Clouds
Through the multimodal autoencoder method that integrates 3D point clouds, the volume, weight and 3D point cloud data of spare parts are used for feature extraction and fusion, which solves the problems of large computing resources and poor real-time performance of traditional detection methods, and realizes efficient and accurate detection of spare parts models.
Patent Information
- Application Number
- CN202510315537.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The existing technology consumes a lot of computing resources and has poor real-time performance in factory spare parts detection, and traditional methods can only accept single-modal data, which limits the accuracy of spare parts model detection.
The multimodal autoencoder method that fuses 3D point clouds is adopted. By obtaining the volume, weight and 3D point cloud data of the spare parts, multimodal feature extraction and feature fusion are performed, and the detection is performed using the autoencoder model to simplify the network structure to reduce the computational complexity.
It improves the accuracy and real-time performance of spare parts detection, is suitable for application scenarios with limited computing resources, reduces overfitting, and improves the generalization ability of the model.
Smart Images

Figure CN119848569B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and particularly to a multi-modal autoencoder intelligent detection method and device for spare parts that fuse 3D point clouds. Background Art
[0002] With the rapid development of computer vision technology and the increasing demand for factory intelligence, how to achieve efficient and intelligent detection and classification of spare parts in factory workshops has become an urgent problem to be solved in factories. Due to the refinement and high standardization of the manufacturing industry, there are extremely high requirements for the accuracy of spare part model detection. Traditional detection methods often rely on complex deep learning models. Although these models can provide high-precision detection results, they often face problems such as large consumption of computing resources and poor real-time performance in practical applications. At the same time, traditional image detection methods can often only accept single-modal data, which greatly limits the accuracy of spare part model detection.
[0003] Therefore, there is a need for a multi-modal autoencoder intelligent detection method for spare part models in factories that fuses 3D point clouds. This method can not only collect the weight and volume data of spare parts, but also integrate the 3D point cloud depth information of spare parts, providing more useful information for judging the spare part model, thereby improving the accuracy of spare part model detection. At the same time, the lightweight method has lower computational complexity and faster response speed, and is easy to be deployed on edge devices, which makes this method have broad application prospects in industrial production lines, intelligent manufacturing, automated detection and other scenarios. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a multi-modal autoencoder intelligent detection method and device for spare parts that fuse 3D point clouds. By obtaining the volume, weight information and multi-modal data of 3D point clouds of spare parts for spare part detection, it can effectively extract the multi-modal data features of spare parts, thereby helping to improve the detection accuracy of spare parts. The network structure of the autoencoder model used is simple, with a small number of parameters. The trained model has low computational overhead and memory occupancy during the inference process, and is more efficient in real-time detection and deployment, especially suitable for application scenarios with limited computing resources.
[0005] On the one hand, the multi-modal autoencoder intelligent detection method for spare parts that fuse 3D point clouds is as follows:
[0006] S1, obtain the volume data, weight data and 3D point cloud data of spare parts of each model; perform normalization processing on the weight data and volume data of spare parts, and convert the 3D point cloud data into a 2D depth image;
[0007] S2. Crop the 2D depth image and then perform feature extraction to obtain a feature map. Fuse the feature map with the normalized weight data and volume data to obtain the multi-modal features of the spare parts.
[0008] S3. Grayscale the multi-modal features of the spare parts to obtain a grayscale multi-modal feature map.
[0009] S4. Divide the grayscale multi-modal feature map into a training set and a test set. Use the training set to train the constructed autoencoder model to obtain a trained autoencoder model.
[0010] S5. Input the test set into the trained autoencoder model to obtain the feature codes of spare parts of each model.
[0011] S6. Process the spare parts to be detected through S1 - S3 to obtain the grayscale multi-modal feature map of the spare parts to be detected, and input it into the trained autoencoder model to obtain the feature code of the spare parts to be detected. Compare the feature code of the spare parts to be detected with the feature codes of spare parts of each model to detect the model information of the spare parts to be detected.
[0012] Preferably, the conversion of 3D point cloud data into 2D depth image is as follows:
[0013] Represent the point (x, y, z) of the 3D point cloud data in the lidar coordinate system as a set of points in the polar coordinate system, i.e., (r, θ, φ). The conversion formula is expressed as:
[0014]
[0015] Among them, based on the lidar coordinate system, r represents the distance from the point to the origin, θ represents the angle of the point in the horizontal direction, and φ represents the angle of the point in the vertical direction; (x, y, z) respectively represent the coordinates of the point on each coordinate axis in the spatial Cartesian coordinate system.
[0016] Map the polar coordinate values to the grayscale image. The pixel grayscale value of the grayscale image represents the depth information of the point cloud. The depth information is the distance from each point in the 3D point cloud data in the three-dimensional space to the origin of the coordinate system.
[0017] Preferably, the cropping and then feature extraction of the 2D depth image is as follows:
[0018] Crop the 2D depth image to a unified size.
[0019] Input the cropped image into the convolution kernel for feature extraction. The size of the convolution kernel is 3×3.
[0020] Input the output of the convolution kernel into the Relu activation function.
[0021] Preferably, the feature map is subjected to feature fusion with the normalized weight data and volume data, specifically as follows:
[0022] Subtract the normalized spare part weight value from the normalized spare part volume value, and add the obtained difference to the extracted feature map pixel by pixel to obtain a multi-modal feature containing the volume information, weight information, and 3D point cloud information of the spare part.
[0023] Preferably, the multi-modal feature of the spare part is gray-scaled to obtain a multi-modal feature gray-scale map, specifically as follows:
[0024] The multi-modal feature of the spare part is input into a 1×1 convolutional module for dimensionality reduction processing to obtain a single-channel gray-scale map, that is, the multi-modal feature gray-scale map.
[0025] Preferably, the loss function of the autoencoder model is expressed as:
[0026]
[0027] where Loss represents the loss function of the autoencoder model; || || represents the Euclidean norm; f() represents a mapping function that maps the multi-modal feature of the spare part to the encoded feature; max(,) represents taking the larger value; α represents a preset margin used to limit the minimum distance between the encoded features obtained from spare parts of different models; f(a) and f(b) respectively represent the encoded features obtained by mapping the multi-modal features of spare parts of different models; f(c k ) and f(c j ) respectively represent the encoded features obtained by mapping the multi-modal features of each component of the same model.
[0028] Preferably, the feature similarity is used to compare the feature encoding of the spare part to be detected with the feature encodings of the spare parts of each model; the feature similarity is expressed as:
[0029] g(x,y) = cos(x,y) + 1;
[0030]
[0031] where g(x,y) represents the feature similarity; cos(x,y) represents the cosine value; x represents the encoded feature of the spare part to be detected, y represents the encoded features of the spare parts of each model; || || represents the Euclidean norm; n represents the total number of features in the feature vector; x i represents the i-th encoded feature in x; y i represents the i-th encoded feature in y.
[0032] Preferably, the feature codes of the spare parts to be detected are compared with the feature codes of the spare parts of each model, and the model information of the spare parts to be detected is obtained through detection, which is specifically as follows:
[0033] The feature similarity between the feature code of the spare part to be detected and the feature codes of the spare parts of each model is calculated respectively. If the feature similarity corresponding to a certain model of spare part is greater than the preset similarity threshold, the model detection of the spare part to be detected is successful, and the model of the spare part to be detected is the spare part of that model; otherwise, the model detection fails.
[0034] Preferably, if the model detection fails, the feature code of the spare part to be detected is added to the feature codes of the spare parts of each model for the next model detection.
[0035] On the other hand, the multi-modal autoencoder spare part intelligent detection device integrating 3D point cloud includes the following:
[0036] The data acquisition and processing module is used to acquire the volume data, weight data, and 3D point cloud data of the spare parts of each model; perform normalization processing on the weight data and volume data of the spare parts, and convert the 3D point cloud data into a 2D depth image;
[0037] The multi-modal feature acquisition module is used to crop the 2D depth image and then extract features to obtain a feature map, and perform feature fusion on the feature map with the normalized weight data and volume data to obtain the multi-modal features of the spare parts;
[0038] The multi-modal feature grayscale image acquisition module is used to perform grayscale processing on the multi-modal features of the spare parts to obtain a multi-modal feature grayscale image;
[0039] The model training module is used to divide the multi-modal feature grayscale image into a training set and a test set; use the training set to train the constructed autoencoder model to obtain a trained autoencoder model;
[0040] The feature code acquisition module for spare parts of each model is used to input the test set into the trained autoencoder model to obtain the feature codes of the spare parts of each model;
[0041] The spare part to be detected detection module is used to process the spare part to be detected through the data acquisition and processing module, the multi-modal feature acquisition module, and the multi-modal feature grayscale image acquisition module to obtain the multi-modal feature grayscale image of the spare part to be detected, and input it into the trained autoencoder model to obtain the feature code of the spare part to be detected, and compare the feature code of the spare part to be detected with the feature codes of the spare parts of each model to detect the model information of the spare part to be detected.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] (1) The present invention performs spare part detection by acquiring the volume, weight information, and multi-modal data of 3D point clouds of spare parts, and can effectively extract the multi-modal data features of spare parts, thereby helping to improve the detection accuracy of spare parts;
[0044] (3) The network structure of the autoencoder model of the present invention is simple, with a small number of parameters. The trained model has low computational overhead and memory occupancy during the inference process, is more efficient in real-time detection and deployment, and is particularly suitable for application scenarios with limited computing resources; it avoids the drawbacks of traditional image detection methods that obtain detection accuracy by deepening the network depth and increasing the model complexity;
[0045] (3) The autoencoder model of the present invention can effectively reduce the overfitting phenomenon and improve the generalization ability of the model through optimized feature learning and data compression capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The present invention will be further described in detail below with reference to the accompanying drawings;
[0047] Figure 1 is a flowchart of the intelligent detection method for spare parts of the multi-modal autoencoder integrating 3D point clouds according to an embodiment of the present invention;
[0048] Figure 2 is an architecture diagram of the detection model of the intelligent detection method for spare parts of the multi-modal autoencoder integrating 3D point clouds according to an embodiment of the present invention;
[0049] Figure 3 is a code diagram of the autoencoder of the intelligent detection method for spare parts of the multi-modal autoencoder integrating 3D point clouds according to an embodiment of the present invention;
[0050] Figure 4 is a structural block diagram of the intelligent detection device for spare parts of the multi-modal autoencoder integrating 3D point clouds according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The present invention will be further described below through specific embodiments.
[0052] As Figure 1 and Figure 2 shown, the intelligent detection method for spare parts of the multi-modal autoencoder integrating 3D point clouds is as follows:
[0053] S1, obtain the volume data, weight data, and 3D point cloud data of each model of spare parts; perform normalization processing on the weight data and volume data of the spare parts, and convert the 3D point cloud data into a 2D depth image.
[0054] Specifically, the weight of the spare parts is measured by an electronic balance in milligrams; the volume of the spare parts is calculated by scanning with a 3D scanner in cubic centimeters; the 3D point cloud of the spare parts is obtained by scanning with a laser 3D scanner. Then, normalization operations are performed on the weight and volume data of the spare parts, and the 3D point cloud of the spare parts is converted into a 2D depth image. The normalization operation for the weight of the spare parts is defined by the following formula:
[0055]
[0056] where max represents the maximum mass among all current spare parts, min represents the minimum mass among all current spare parts, represents the average mass among all current spare parts, m represents the mass of the spare part to be detected currently, and m′ represents the normalized mass of the spare part.
[0057] The normalization operation for the volume of the spare parts is defined by the following formula:
[0058]
[0059] where max represents the maximum volume among all current spare parts, min represents the minimum volume among all current spare parts, represents the average volume among all current spare parts, v represents the volume of the spare part to be detected currently, and v′ represents the normalized volume of the spare part.
[0060] The main function of data normalization is to help improve the training effect of the model and accelerate the convergence process in machine learning and deep learning model training.
[0061] The conversion operation method for converting the 3D point cloud of the spare parts into a 2D depth image is as follows:
[0062] This method first represents the point (x, y, z) in the lidar coordinate system as a set of points in the polar coordinate system, that is, (r, θ, φ), and the conversion formula is specifically:
[0063]
[0064] where, based on the radar coordinate system, r represents the distance from the point to the origin, θ represents the angle of the point in the horizontal direction, and φ represents the angle of the point in the vertical direction. (x, y, z) represents the coordinates of the point in the spatial Cartesian coordinate system.
[0065] In practical applications, since the projection process is continuous, the pixel points of the 2D image may not exactly correspond to the projection positions of the 3D points. Therefore, interpolation processing is usually required to map the discrete points in the 3D point cloud to the 2D image grid. Common interpolation methods include nearest neighbor interpolation, bilinear interpolation, etc.
[0066] S2. Crop the 2D depth image and then perform feature extraction to obtain a feature map. Fuse the feature map with the normalized weight data and volume data to obtain the multi-modal features of the spare part.
[0067] Specifically, uniformly crop the 2D depth image to a size of S×S pixels, where S is defined as:
[0068] S = 1.2×max(height max , width max );
[0069] where width max , height max represent the maximum width and height values of the spare part size respectively, and max(,) means taking the larger value. The cropped spare part image retains the original information of the spare part and, to a great extent, removes the interference of background information.
[0070] After that, input the cropped image into a 3×3 convolution kernel for feature extraction, and then input the output feature into the Relu activation function to introduce a non-linear relationship, enabling the network to represent more complex features, thereby improving the expressiveness and generalization ability of the model. The convolution kernel has a stride of 1 and a padding of 1 to ensure that the output feature map has the same size as the input image.
[0071] The specific method of feature fusion between the feature map and the normalized weight data and normalized volume data is as follows: Subtract the normalized spare part weight value from the normalized spare part volume value, and then add the obtained difference to the feature map pixel by pixel, thereby obtaining a multi-modal feature map containing the volume information, weight information, and 3D point cloud information of the spare part. Compared with traditional deep learning methods, through this multi-modal feature fusion module, the spare part detection model can utilize more comprehensive spare part information, thereby improving the accuracy of spare part detection. This innovative design significantly enhances the model's understanding of spare part features and helps improve the detection performance of the model.
[0072] S3. Grayscale the multi-modal features of the spare part to obtain a multi-modal feature grayscale map.
[0073] Use a 1×1 convolution kernel to grayscale the multi-modal features of the spare part and perform channel dimensionality reduction to obtain a single-channel grayscale map, which can reduce the number of model parameters, prevent overfitting, and at the same time help reduce training parameters and contribute to the lightweight of the model.
[0074] S4. Divide the multi-modal feature grayscale images into a training set and a test set; use the training set to train the constructed auto-encoder model to obtain a trained auto-encoder model.
[0075] The construction code of the auto-encoder model in this embodiment can be seen in Figure 3 As shown, design the following loss function based on the construction code of the auto-encoder model to train and build the auto-encoder model. The formula of the loss function is expressed as:
[0076]
[0077] where, || || represents the Euclidean distance, f() represents a mapping function that maps the multi-modal features of spare parts into encoded features, max represents taking the larger value, α is a preset margin, and α is used to limit the minimum distance between the encoded features obtained from different model spare parts. f(a) and f(b) respectively represent the encoded features obtained by mapping the multi-modal features of different model spare parts, and f(c k ) and f(c j ) respectively represent the encoded features obtained by mapping the multi-modal features of each component of the same model.
[0078] The advantage of using the designed loss function to train the built Auto-encoder network in this embodiment is that it can make the mapping distances between the encoded features obtained from the multi-modal features of the same model components as close as possible, while ensuring that the mapping distances between the encoded features obtained from the multi-modal features of different model components are as far as possible within a suitable range. The benefit of such a design is that it can make the feature encodings of the same model spare parts as similar as possible and make the feature encodings of different model spare parts as different as possible.
[0079] S5. Input the test set into the trained auto-encoder model to obtain the feature encodings of each model of spare parts.
[0080] Specifically, input the grayscale images of the multi-modal features of each model of spare parts into the trained auto-encoder model to obtain the feature encodings of each model of spare parts, and use the obtained feature encodings of each model to construct a feature vector database for each model of spare parts. The feature vector database established in this embodiment is a database of the feature encodings of each model of spare parts in the training samples, and each model of spare parts stores one feature encoding in the database for the subsequent detection of the spare parts to be detected.
[0081] S6. Process the spare parts to be detected through S1 - S3 to obtain the multi - modal feature grayscale image of the spare parts to be detected, and input it into the trained auto - encoder model to obtain the feature encoding of the spare parts to be detected. Compare the feature encoding of the spare parts to be detected with the feature encodings of spare parts of each model to detect the model information of the spare parts to be detected.
[0082] After the database is established, obtain the weight, volume data and 3D point cloud of the spare parts to be detected, and input the data of the spare parts to be detected into the network. The data processing process is the same as that of the spare parts used for training. According to these data, obtain the grayscale image of the multi - modal features of the spare parts to be detected. Then input the grayscale image of the multi - modal features of the spare parts to be detected into the trained Auto - encoder model to obtain the encoded features of the spare parts to be detected. Finally, compare the encoded features of the spare parts to be detected with the encoded features of spare parts of each model stored in the database. If the feature similarity between the encoded features of the spare parts to be detected and the encoded features of a certain model of spare parts in the database is not lower than the set threshold, output the model of the spare parts, and the detection is successful. Otherwise, if the encoded features of this spare part to be detected have not been stored in the database, assign a model name to this spare part to be detected and store the encoded features of this spare part to be detected in the database.
[0083] Specifically, the calculation method of the feature similarity between the encoded features of different spare parts in this embodiment refers to the cosine similarity calculation formula, and the cosine similarity calculation formula is as follows:
[0084] g(x,y) = cos(x,y)+1;
[0085] In the above formula, cos(x,y) is expressed as:
[0086]
[0087] Among them, x represents the encoded feature obtained by the Auto - encoder model according to the multi - modal features of the spare parts to be detected, and y represents the encoded features of spare parts of each model stored in the database. i represents the i - th feature in the feature vector. i is an integer from 1 to n, indicating a specific dimension of the feature vector, and n represents the total number of features in the feature vector, that is, how many features or dimensions there are in the vector. x i is the i - th feature of the feature vector x, and y iIt is the i-th feature of the eigenvector y. || || is the Euclidean norm. Since the value range of cos(x, y) is [-1, 1], the value range of g(x, y) is [0, 2]. The closer the calculation result of g(x, y) is to 2, the higher the feature similarity between x and y. If the set feature similarity threshold is θ (θ is greater than 1), when g(x, y) is greater than or equal to the set similarity threshold, there exists the coding feature of the spare parts of the same model as the spare parts to be detected in the database, the detection is successful, and the model name of this spare part is output. Otherwise, the detection fails. The coding feature of this spare part to be detected has not been stored in the database, so a model name is given to this spare part to be detected, and the coding feature of this spare part to be detected is stored in the database.
[0088] Such as Figure 4 As shown, the present invention also discloses a multi-modal autoencoder spare parts intelligent detection device integrating 3D point cloud, including:
[0089] A data acquisition and processing module 401, configured to acquire the volume data, weight data, and 3D point cloud data of spare parts of each model; perform normalization processing on the weight data and volume data of the spare parts, and convert the 3D point cloud data into a 2D depth image;
[0090] A multi-modal feature acquisition module 402, configured to crop the 2D depth image and then perform feature extraction to obtain a feature map, and perform feature fusion on the feature map with the normalized weight data and volume data to obtain the multi-modal features of the spare parts;
[0091] A multi-modal feature grayscale image acquisition module 403, configured to perform grayscale processing on the multi-modal features of the spare parts to obtain a multi-modal feature grayscale image;
[0092] A model training module 404, configured to divide the multi-modal feature grayscale image into a training set and a test set; use the training set to train the constructed autoencoder model to obtain a trained autoencoder model;
[0093] A feature coding acquisition module 405 for each model of spare parts, configured to input the test set into the trained autoencoder model to obtain the feature codes of spare parts of each model;
[0094] A spare part to be detected detection module 406, configured to process the spare part to be detected through the data acquisition and processing module, the multi-modal feature acquisition module, and the multi-modal feature grayscale image acquisition module to obtain the multi-modal feature grayscale image of the spare part to be detected, and input it into the trained autoencoder model to obtain the feature code of the spare part to be detected, and compare the feature code of the spare part to be detected with the feature codes of spare parts of each model to detect the model information of the spare part to be detected.
[0095] The specific implementation of the intelligent detection device for spare parts of the multi-modal autoencoder integrating 3D point cloud is the same as the intelligent detection method for spare parts of the multi-modal autoencoder integrating 3D point cloud, which will not be repeated in this embodiment.
[0096] The above are only specific embodiments of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification of the present invention using this concept shall fall within the scope of infringement of the protection scope of the present invention.
Claims
1. A multi-modal autoencoder intelligent detection method for spare parts integrating 3D point clouds, characterized in that, It includes the following steps: S1. Obtain the volume data, weight data, and 3D point cloud data of spare parts of each model; perform normalization processing on the weight data and volume data of the spare parts, and convert the 3D point cloud data into a 2D depth image; S2. Crop the 2D depth image and then perform feature extraction to obtain a feature map, and perform feature fusion on the feature map with the normalized weight data and volume data to obtain the multi-modal features of the spare parts; S3. Perform grayscale processing on the multi-modal features of the spare parts to obtain a multi-modal feature grayscale image; S4. Divide the multi-modal feature grayscale image into a training set and a test set; use the training set to train the constructed autoencoder model to obtain a trained autoencoder model; S5. Input the test set into the trained autoencoder model to obtain the feature codes of spare parts of each model; S6. Process the spare parts to be detected through S1 - S3 to obtain the multi-modal feature grayscale image of the spare parts to be detected, and input it into the trained autoencoder model to obtain the feature code of the spare parts to be detected. Compare the feature code of the spare parts to be detected with the feature codes of spare parts of each model to detect the model information of the spare parts to be detected; The loss function of the autoencoder model is expressed as: Among them, Loss represents the loss function of the autoencoder model; || || represents the Euclidean norm; f() represents a mapping function that maps the multi-modal features of spare parts into encoded features; max(,) represents taking the larger value; α represents a preset margin used to limit the minimum distance between the encoded features obtained from spare parts of different models; f(a) and f(b) respectively represent the encoded features obtained by mapping the multi-modal features of spare parts of different models; f(c k ) and f(c j ) respectively represent the encoded features obtained by mapping the multi-modal features of each component of the same model; Use feature similarity to perform feature comparison between the feature code of the spare parts to be detected and the feature codes of spare parts of each model; the feature similarity is expressed as: g(x,y) = cos(x,y) + 1; Among them, g(x, y) represents the feature similarity; cos(x, y) represents the cosine value; x represents the coding feature of the spare parts to be detected, and y represents the coding features of spare parts of each model; || || represents the Euclidean norm; n represents the total number of features in the feature vector; x i represents the i-th coding feature in x; y i represents the i-th coding feature in y; The process of comparing the feature code of the spare parts to be detected with the feature codes of spare parts of each model to detect the model information of the spare parts to be detected is as follows: Calculate the feature similarity between the feature code of the spare parts to be detected and the feature codes of spare parts of each model respectively. If the feature similarity corresponding to a certain model of spare parts is greater than the preset similarity threshold, the model detection of the spare parts to be detected is successful, and the model of the spare parts to be detected is the spare parts of that model; otherwise, the model detection fails.
2. The multi-modal autoencoder spare part intelligent detection method for fusing 3D point clouds according to claim 1, wherein, The process of converting the 3D point cloud data into a 2D depth image is as follows: Express the point (x, y, z) of the 3D point cloud data in the lidar coordinate system as a set of points in the polar coordinate system, that is, (r, θ, φ), and the conversion formula is expressed as: Where, based on the lidar coordinate system, r represents the distance from the point to the origin, θ represents the angle of the point in the horizontal direction, and φ represents the angle of the point in the vertical direction; (x, y, z) respectively represent the coordinates of the point on each coordinate axis in the spatial Cartesian coordinate system; Map the polar coordinate values to a grayscale image; the pixel grayscale value of the grayscale image represents the depth information of the point cloud; the depth information is the distance from each point in the 3D point cloud data in three-dimensional space to the origin of the coordinate system.
3. The intelligent detection method for multi-modal autoencoder spare parts of fused 3D point clouds according to claim 1, wherein The process of cropping the 2D depth image and then performing feature extraction is as follows: Crop the 2D depth image to a unified size; Input the cropped image into a convolution kernel for feature extraction; the size of the convolution kernel is 3×3; Input the output of the convolution kernel into the Relu activation function.
4. The intelligent detection method for multi-modal autoencoder spare parts of fused 3D point cloud according to claim 1, wherein The process of performing feature fusion on the feature map with the normalized weight data and volume data is as follows: Subtract the normalized spare part weight value from the normalized spare part volume value, and add the obtained difference to the extracted feature map pixel by pixel to obtain a multi-modal feature containing the volume information, weight information, and 3D point cloud information of the spare part.
5. The intelligent detection method for multi-modal autoencoder spare parts of fused 3D point cloud according to claim 1, characterized in that The multi-modal feature of the spare part is gray-scaled to obtain a gray-scale multi-modal feature map, which is specifically as follows: Input the multi-modal feature of the spare part into a 1×1 convolutional module for dimensionality reduction to obtain a single-channel gray-scale map, that is, the gray-scale multi-modal feature map.
6. The intelligent detection method for multi-modal autoencoder spare parts of fused 3D point cloud according to claim 1, wherein If the model detection fails, add the feature encoding of the spare part to be detected to the feature encodings of the spare parts of each model for the next model detection.
7. A multi-modal autoencoder spare part intelligent detection device integrating 3D point cloud, including the following: A data acquisition and processing module, which is used to acquire the volume data, weight data, and 3D point cloud data of the spare parts of each model; perform normalization processing on the weight data and volume data of the spare parts, and convert the 3D point cloud data into a 2D depth image; A multi-modal feature acquisition module, which is used to crop the 2D depth image and then extract features to obtain a feature map, and fuse the feature map with the normalized weight data and volume data to obtain the multi-modal feature of the spare part; A gray-scale multi-modal feature map acquisition module, which is used to gray-scale the multi-modal feature of the spare part to obtain a gray-scale multi-modal feature map; A model training module, which is used to divide the gray-scale multi-modal feature map into a training set and a test set; use the training set to train the constructed autoencoder model to obtain a trained autoencoder model; A feature encoding acquisition module for each model of spare parts, which is used to input the test set into the trained autoencoder model to obtain the feature encodings of the spare parts of each model; A spare part to be detected module, which is used to process the spare part to be detected through the data acquisition and processing module, the multi-modal feature acquisition module, and the gray-scale multi-modal feature map acquisition module to obtain the gray-scale multi-modal feature map of the spare part to be detected, and input it into the trained autoencoder model to obtain the feature encoding of the spare part to be detected. Compare the feature encoding of the spare part to be detected with the feature encodings of the spare parts of each model to detect the model information of the spare part to be detected; The loss function of the autoencoder model is expressed as: Among them, Loss represents the loss function of the autoencoder model; || || represents the Euclidean norm; f() represents a mapping function that maps the multi-modal features of spare parts into encoded features; max(,) represents taking the larger value; α represents a preset margin used to limit the minimum distance between the encoded features obtained from spare parts of different models; f(a) and f(b) respectively represent the encoded features obtained by mapping the multi-modal features of spare parts of different models; f(c k ) and f(c j ) respectively represent the encoded features obtained by mapping the multi-modal features of each component of the same model; Use feature similarity to compare the feature encoding of the spare part to be detected with the feature encodings of the spare parts of each model; the feature similarity is expressed as: g(x,y) = cos(x,y) + 1; Among them, g(x, y) represents the feature similarity; cos(x, y) represents the cosine value; x represents the coding feature of the spare part to be detected, and y represents the coding features of spare parts of each model; || || represents the Euclidean norm; n represents the total number of features in the feature vector; x i represents the i-th coding feature in x; y i represents the i-th coding feature in y; The feature encoding of the spare part to be detected is compared with the feature encodings of the spare parts of each model to detect the model information of the spare part to be detected, which is specifically as follows: Calculate the feature similarity between the feature encoding of the spare part to be detected and the feature encodings of the spare parts of each model respectively. If the feature similarity corresponding to a certain model of spare part is greater than the preset similarity threshold, the model detection of the spare part to be detected is successful, and the model of the spare part to be detected is the spare part of that model; otherwise, the model detection fails.
Citation Information
Patent Citations
Rainfall runoff prediction method based on multi-modal fusion
CN117035148A
Part defect image detection method based on self-supervised learning
CN118351107A
Cited By
Precise aeronautical part assembly method based on multi-modal fusion perception and action generation
CN122263307A