Image and point cloud feature fusion method and system based on conditional generative adversarial network

Through the method of generating adversarial networks based on conditions, the problem of feature fusion between image and point cloud data in unmanned mining areas is solved, the performance and robustness of the intelligent driving system are improved, and the feature extraction and fusion process is simplified.

CN120356048AActive Publication Date: 2025-07-22CHONGQING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510427500.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In the prior art, the feature fusion method of image and point cloud data is difficult to be effectively carried out in complex environments, resulting in insufficient accuracy and robustness of object detection and large computing resources.

Method used

The method of generating adversarial networks based on conditions is adopted, and point cloud and RGB image data are acquired, preprocessing and feature extraction are performed, and target synthetic point cloud data is generated using the conditional generation adversarial network model, and feature fusion is performed through dual Gaussian distribution.

Benefits of technology

It realizes effective feature fusion between images and point cloud data, improves the performance of intelligent driving systems in unmanned mining areas, and simplifies the feature extraction and fusion process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356048A_ABST
    Figure CN120356048A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image data fusion, and provides an image and point cloud feature fusion method and system based on a conditional generative adversarial network, and the method comprises the steps: obtaining first point cloud data and RGB image data, and carrying out the preprocessing of the first point cloud data, and obtaining second point cloud data; calling a target conditional generative adversarial network model, and generating target synthetic point cloud data based on the second point cloud data and the RGB image data; according to the target synthetic point cloud data and the second point cloud data, generating a double-Gaussian distribution synthetic point cloud feature, and according to the target synthetic point cloud data, the second point cloud data and the double-Gaussian distribution synthetic point cloud feature, processing to obtain a fused point cloud feature, thereby solving a problem of how to perform effective feature fusion on the image data and the point cloud data. And the performance of the intelligent driving system in the unmanned mining area is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data fusion, and particularly to an image and point cloud feature fusion method and system based on a conditional generative adversarial network. Background Art

[0002] The introduction of unmanned mining trucks and unmanned bulldozers in open-pit coal mines has brought great convenience to unmanned operations. An intelligent driving system usually needs to detect obstacles in the mining area, such as operating personnel, stockpiles, retaining walls, etc., to ensure safe operation. Therefore, in order to reduce labor costs while improving operation efficiency, it is necessary to promote the intelligent development of target detection in the mining area. Currently, target detection methods can be divided into three types. The first is the target detection method based on images, the second is the target detection method based on radar point clouds, and the third is the target detection method based on the fusion of images and point clouds. However, the mining area usually has harsh natural conditions, which will seriously affect the detection effect based on images. And the target detection based only on radar point clouds does not utilize appearance information such as the texture, color, and shape of the target. Through the target detection method based on the fusion of images and point clouds, the advantages of both modalities can be utilized simultaneously to make up for the deficiencies of a single modality, thereby improving the accuracy and robustness of target detection.

[0003] Currently, the main target detection methods based on the fusion of images and point clouds are as follows: One is to map a two-dimensional image into three-dimensional space through depth estimation, and then encode it into a Bird's Eye View (BEV) feature, which is fused with the BEV feature of the point cloud data; Another is to project each point in the lidar point cloud into the output of an image semantic segmentation network, attach the segmentation result to the corresponding lidar point, so as to add the semantic information of the image to the point cloud, and finally input the point cloud fused with the image features into a three-dimensional (3D) target detection network based on lidar for detection; There is also a method that uses a BEV map, a front view of the point cloud, and three primary color (Red Green Blue, RGB) Figure 3 feature maps for fusion, and fully utilizes the position information of the point cloud data and the semantic information of the image data for feature fusion. Although the above several methods have achieved good recognition effects after experimental verification, due to the sparsity and noise problems of the point cloud data, complex preprocessing is required in the early stage of model training, and it is difficult to align the features of different modalities, which requires a large amount of computing resources, and important feature information is inevitably lost in complex processing and calculations.

[0004] Therefore, how to effectively fuse image data and point cloud data is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide an image and point cloud feature fusion method and system based on a conditional generative adversarial network to solve the problem of how to effectively fuse image data and point cloud data in the prior art; that is, embodiments of the present invention can improve the performance of intelligent driving systems in unmanned mining areas.

[0006] According to one aspect of the present invention, there is provided an image and point cloud feature fusion method based on a conditional generative adversarial network. The image and point cloud feature fusion method based on a conditional generative adversarial network includes: obtaining first point cloud data and RGB image data, and preprocessing the first point cloud data to obtain second point cloud data; calling a target conditional generative adversarial network model, and generating target synthetic point cloud data based on the second point cloud data and the RGB image data; generating a double Gaussian distribution synthetic point cloud feature according to the target synthetic point cloud data and the second point cloud data, and processing the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud feature to obtain a fused point cloud feature.

[0007] In one embodiment, the calling a target conditional generative adversarial network model and generating target synthetic point cloud data based on the second point cloud data and the RGB image data includes: extracting point cloud features based on the second point cloud data, and extracting RGB image features based on the RGB image data; training a conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain a target conditional generative adversarial network model; and processing the second point cloud data and the RGB image data based on the target conditional generative adversarial network model to generate target synthetic point cloud data.

[0008] In one embodiment, the training a conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain a target conditional generative adversarial network model includes: obtaining a random noise vector, generating synthetic point cloud features based on the random noise vector, the point cloud features, and the RGB image features, and obtaining synthetic point cloud data according to the synthetic point cloud features; performing a discrimination process based on the synthetic point cloud features and the point cloud features to obtain a discrimination result; obtaining random point cloud features, generating a first loss value based on the synthetic point cloud features, generating a third loss value based on the synthetic point cloud data and the second point cloud data, generating a fourth loss value based on the synthetic point cloud features and the random point cloud features, and performing a weighted sum of the first loss value, the third loss value, and the fourth loss value to obtain a target loss value; and optimizing the model parameters in the conditional generative adversarial network model in a direction of reducing the target loss value to obtain the target conditional generative adversarial network model.

[0009] In one embodiment, generating a double Gaussian distribution synthetic point cloud feature based on the target synthetic point cloud data and the second point cloud data, and processing to obtain a fused point cloud feature based on the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud feature includes: respectively performing Poisson disk sampling on the target synthetic point cloud data and the second point cloud data to obtain sampled synthetic point cloud data and third point cloud data; respectively performing feature extraction on the sampled synthetic point cloud data and the third point cloud data to obtain sampled synthetic point cloud features and third point cloud features; respectively performing BEV projection on the sampled synthetic point cloud features and the third point cloud features to obtain BEV synthetic point cloud features and BEV point cloud features; performing double Gaussian distribution sampling on the BEV synthetic point cloud features and the BEV point cloud features to obtain the double Gaussian distribution synthetic point cloud feature; and performing stacking processing on the BEV synthetic point cloud features, the BEV point cloud features, and the double Gaussian distribution synthetic point cloud feature to obtain the fused point cloud feature.

[0010] In one embodiment, performing weighted summation on the first loss value, the third loss value, and the fourth loss value to obtain a target loss value, where the target loss value is:

[0011]

[0012] where α is the first weight, β is the second weight, γ is the third weight, is the first loss value, is the third loss value, P G is the synthetic point cloud data, is the second point cloud data, L contrast is the fourth loss value.

[0013] In one embodiment, performing double Gaussian distribution sampling on the BEV synthetic point cloud features and the BEV point cloud features to obtain the double Gaussian distribution synthetic point cloud feature, where the double Gaussian distribution synthetic point cloud feature is:

[0014] F sample ~p(F)

[0015]

[0016] where π1 is the first mixing coefficient, π2 is the second mixing coefficient, π1 + π2 = 1, (μ1, Σ1) is the mean and covariance of the BEV point cloud feature, (μ2, Σ2) is the mean and covariance of the BEV synthetic point cloud feature, p(F) is the double Gaussian distribution synthetic point cloud data, is the Gaussian distribution.

[0017] According to another aspect of the present invention, there is provided an image and point cloud feature fusion system based on a conditional generative adversarial network. The image and point cloud feature fusion system based on the conditional generative adversarial network includes: a preprocessing module, an image data generation module, and an image feature fusion module. Among them, the preprocessing module is used to obtain first point cloud data and RGB image data, and preprocess the first point cloud data to obtain second point cloud data; the image data generation module is used to call a target conditional generative adversarial network model to generate target synthetic point cloud data based on the second point cloud data and the RGB image data; the image feature fusion module is used to generate a double Gaussian distribution synthetic point cloud feature according to the target synthetic point cloud data and the second point cloud data, and process the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud feature to obtain a fused point cloud feature.

[0018] In one embodiment, the image data generation module includes: a feature extraction module, a model training module, and a generation module. Among them, the feature extraction module is used to extract point cloud features based on the second point cloud data and extract RGB image features based on the RGB image data; the model training module is used to train a conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain a target conditional generative adversarial network model; the generation module is used to process the second point cloud data and the RGB image data based on the target conditional generative adversarial network model to generate target synthetic point cloud data.

[0019] In one embodiment, the model training module includes: a first generation unit, a discriminant unit, a calculation unit, and a parameter optimization unit. Among them, the first generation unit is used to obtain a random noise vector, generate a synthetic point cloud feature based on the random noise vector, the point cloud features, and the RGB image features, and obtain synthetic point cloud data according to the synthetic point cloud feature; the discriminant unit is used to perform discriminant processing based on the synthetic point cloud feature and the point cloud feature to obtain a discriminant result; the calculation unit is used to obtain random point cloud features, generate a first loss value based on the synthetic point cloud feature, generate a third loss value based on the synthetic point cloud data and the second point cloud data, generate a fourth loss value based on the synthetic point cloud feature and the random point cloud features, and perform weighted summation on the first loss value, the third loss value, and the fourth loss value to obtain a target loss value; the parameter optimization unit is used to optimize the model parameters in the conditional generative adversarial network model in the direction of reducing the target loss value to obtain the target conditional generative adversarial network model.

[0020] In one embodiment, the image feature fusion module includes a first sampling module, a feature extraction module, a projection module, a second sampling module, and a fusion module. Among them, the first sampling module is configured to perform Poisson disk sampling on the target synthetic point cloud data and the second point cloud data respectively to obtain sampled synthetic point cloud data and a third point cloud data; the feature extraction module is configured to perform feature extraction on the sampled synthetic point cloud data and the third point cloud data respectively to obtain sampled synthetic point cloud features and third point cloud features; the projection module is configured to perform BEV projection on the sampled synthetic point cloud features and the third point cloud features respectively to obtain BEV synthetic point cloud features and BEV point cloud features; the second sampling module is configured to perform double Gaussian distribution sampling on the BEV synthetic point cloud features and the BEV point cloud features to obtain the double Gaussian distribution synthetic point cloud features; the fusion module is configured to perform stacking processing on the BEV synthetic point cloud features, the BEV point cloud features, and the double Gaussian distribution synthetic point cloud features to obtain the fused point cloud features.

[0021] In summary, in the embodiments of the present invention, by acquiring the first point cloud data and the RGB image data, preprocessing the first point cloud data to obtain the second point cloud data, calling the target conditional generative adversarial network model, generating the target synthetic point cloud data based on the second point cloud data and the RGB image data, generating the double Gaussian distribution synthetic point cloud features according to the target synthetic point cloud data and the second point cloud data, and processing to obtain the fused point cloud features according to the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud features, the problem of how to effectively fuse the features of the image data and the point cloud data is solved, and the performance of the intelligent driving system in the unmanned mining area is improved. At the same time, by training the conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain the target conditional generative adversarial network model, and processing the second point cloud data and the RGB image data based on the target conditional generative adversarial network model to generate the target synthetic point cloud data, the data modality is unified, thereby simplifying the subsequent feature extraction and fusion process. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In the following description of the exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present invention are disclosed. In the drawings:

[0023] Figure 1 A flowchart showing the method for fusing image and point cloud features based on a conditional generative adversarial network disclosed in the embodiments of the present application is shown;

[0024] Figure 2 Shows Figure 1 A flowchart of the steps of step S120 shown;

[0025] Figure 3 shows Figure 2 a schematic flowchart of step S122 shown;

[0026] Figure 4 shows Figure 2 a schematic structural diagram of the conditional generative adversarial network model in the step S122 shown;

[0027] Figure 5 shows Figure 1 a schematic flowchart of step S130 shown;

[0028] Figure 6 shows a schematic structural diagram of an image and point cloud feature fusion system based on a conditional generative adversarial network disclosed in an embodiment of the present application. Detailed implementation manners

[0029] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0030] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0031] As used herein, the term "comprising" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0032] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".

[0033] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0034] It should be noted that the execution subject of the method for fusing image and point cloud features based on a conditional generative adversarial network provided in the embodiments of the present invention can be one or more electronic devices, and the present invention does not make any limitations in this regard; among them, the electronic device can be a terminal (i.e., a client) or a server. Then, when the execution subject includes multiple electronic devices, and at least one terminal and at least one server are included in the multiple electronic devices, the method for fusing image and point cloud features based on a conditional generative adversarial network provided in the embodiments of the present invention can be jointly executed by the terminal and the server. Correspondingly, the terminals mentioned herein may include, but are not limited to: smart phones, tablet computers, laptop computers, desktop computers, smart watches, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, and so on. The server mentioned herein can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, and so on.

[0035] Based on the above description, the embodiments of the present invention propose a method for fusing image and point cloud features based on a conditional generative adversarial network. The method for fusing image and point cloud features based on a conditional generative adversarial network can be executed by the above-mentioned electronic devices (terminals or servers); or, the method for fusing image and point cloud features based on a conditional generative adversarial network can be jointly executed by the terminal and the server. For the convenience of description, hereinafter, the method for fusing image and point cloud features based on a conditional generative adversarial network executed by an electronic device will be used as an example for illustration.

[0036] Please refer to Figure 1 , which is a schematic flowchart of the method for fusing image and point cloud features based on a conditional generative adversarial network disclosed in the embodiments of the present application. The method for fusing image and point cloud features based on a conditional generative adversarial network solves the problem of how to effectively fuse image data and point cloud data, thereby improving the performance of the intelligent driving system in an unmanned mining area. It should be noted that the method for fusing image and point cloud features based on a conditional generative adversarial network in the embodiments of the present application is not limited to Figure 1 the steps and sequences in the flowchart shown. According to different requirements, the steps in the shown flowchart can be added, removed, or the order can be changed. In the embodiments of the present application, as Figure 1As shown, the process of the image and point cloud feature fusion method based on conditional generative adversarial network at least includes the following steps.

[0037] S110. Obtain the first point cloud data and RGB image data, and preprocess the first point cloud data to obtain the second point cloud data.

[0038] In an embodiment of the present invention, a lidar (Light Laser Detection and Ranging, LiDAR) with high-precision and long-distance detection capabilities is installed on a vehicle in an unmanned mining area. The lidar is used to collect the first point cloud data P of the entire scene centered on the vehicle itself. Where represents the i1-th first point cloud data, and the first point cloud data may include clear surface point cloud data of target recognition obstacles. At the same time, a camera is installed on the vehicle in the unmanned mining area, and the camera is used to collect RGB image data X. Where represents the i2-th RGB image data.

[0039] Since the first point cloud data covers the entire unmanned mining area at 360°, while the RGB image data only includes the target recognition obstacles within the viewing angle and the surrounding part of the environment, it is necessary to preprocess the first point cloud data to the second point cloud data. Where represents the i3-th second point cloud data, so that the viewing angle areas of the second point cloud data and the RGB image data are basically coincident. Among them, the preprocessing may include noise reduction, cropping, and adjusting the viewing angle.

[0040] S120. Invoke the target conditional generative adversarial network model, and generate the target synthetic point cloud data based on the second point cloud data and the RGB image data.

[0041] As Figure 2 shown, in an embodiment of the present invention, Figure 2 The step S120 at least includes the following steps:

[0042] S121. Extract point cloud features based on the second point cloud data, and extract RGB image features based on the RGB image data.

[0043] S122. Train the conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain the target conditional generative adversarial network model.

[0044] As Figure 3 shown, in an embodiment of the present invention, Figure 3 The step S122 at least includes the following steps:

[0045] S1221. Obtain a random noise vector, generate synthetic point cloud features based on the random noise vector, the point cloud features, and the RGB image features, and obtain synthetic point cloud data according to the synthetic point cloud features.

[0046] In the embodiments of the present invention, as Figure 4 shown, the conditional generative adversarial network model 200 includes a generator G and a discriminator D. The generator can be a network structure including an encoder-decoder. The encoder can include convolutional layers and pooling layers. The convolutional layers are used to capture local features, and the pooling layers are used to reduce the spatial dimension of the features. At the same time, the pooling layers are also used to retain important semantic information, such as color, texture, and shape features in the RGB image features. The decoder performs deconvolution on the features extracted by the encoder and finally outputs a tensor with a dimension of [(point_size * 3), 1] through a fully connected layer, where point_size is the number of generated point clouds, and 3 refers to the three-dimensional coordinates (x, y, z) of each point cloud.

[0047] The generator receives the random noise vector and the RGB image feature F X as inputs, generates initial synthetic point cloud features, and the initial synthetic point cloud features are processed through a fully connected layer to obtain initial synthetic point cloud data The initial synthetic point cloud data is as follows:

[0048]

[0049] where F X is the RGB image feature, is the random noise vector, is the initial synthetic point cloud data.

[0050] At the same time, the point cloud features are used as a supervision signal. A cross-attention mechanism is introduced into the generator. By applying the supervision signal to the initial synthetic point cloud features generated by the generator, the generator can dynamically adjust the geometric structure of the point cloud to make it closer to the point cloud features. When the conditional generative adversarial network model is trained, it is no longer necessary to use the point cloud features as a supervision signal. Based on the random noise vector and the RGB image features, synthetic point cloud features can be directly generated, and synthetic point cloud data can be obtained according to the synthetic point cloud features. The cross-attention mechanism is as follows:

[0051]

[0052] where Q is the query vector, and the query vector Q is from the point cloud features extracted from it, K is the key vector, and the key vector K is extracted from the RGB image feature F X extracted from it, V is the value vector, and the value vector V is extracted from the RGB image feature F X extracted from it, QK T is the similarity between Q and K, d k is the dimension of the key, d k is used as a scaling factor to avoid excessive numerical values.

[0053] After introducing the cross-attention mechanism, the initial synthetic point cloud feature is weighted and corrected to obtain the corrected synthetic point cloud feature, and the corrected synthetic point cloud feature is processed through a fully connected layer to obtain the corrected synthetic point cloud data to guide the generator G to generate a more realistic point cloud. Among them, the synthetic point cloud feature can include the initial synthetic point cloud feature and the corrected synthetic point cloud feature The synthetic point cloud data P G can include the initial synthetic point cloud data and the corrected synthetic point cloud data The corrected synthetic point cloud data is as follows:

[0054]

[0055] Among them, is a random noise vector, and CrossAttention(Q, K, V) is the cross-attention mechanism.

[0056] S1222. Perform discrimination processing based on the synthetic point cloud feature and the point cloud feature to obtain a discrimination result.

[0057] In the embodiment of the present invention, the initial synthetic point cloud feature is processed through a fully connected layer to obtain the initial synthetic point cloud data The discriminator D receives the second point cloud data and the corrected synthetic point cloud data, and performs discrimination processing on the authenticity of the corrected synthetic point cloud data to obtain a discrimination result The discrimination result of the discriminator D is as follows:

[0058]

[0059] Taking the discrimination result as 0 or 1 as an example, when is 1, it indicates that the generation effect of the generator G is completely real; when is 0, it indicates that the generation effect of the generator G is not good.

[0060] S1223. Obtain the random point cloud features, generate a first loss value based on the synthetic point cloud features, generate a third loss value based on the synthetic point cloud data and the second point cloud data, generate a fourth loss value based on the synthetic point cloud features and the random point cloud features, and perform weighted summation on the first loss value, the third loss value, and the fourth loss value to obtain the target loss value.

[0061] In the embodiment of the present invention, based on the synthetic point cloud features generate a first loss value The first loss value is as follows:

[0062]

[0063] where G is the generator, is the synthetic point cloud feature.

[0064] Based on the point cloud feature and the synthetic point cloud feature generate a second loss value The second loss value is as follows:

[0065]

[0066] where D is the discriminator, is the point cloud feature, is the synthetic point cloud feature.

[0067] In order to measure the point cloud synthesis effect, the Chamfer Distance (CD) index is introduced. The performance of the generator is optimized by measuring the bidirectional nearest neighbor distance between the synthetic point cloud data and the second point cloud data through CD. Based on the synthetic point cloud data P G and the second point cloud data generate a third loss value The third loss value is as follows:

[0068]

[0069] where P G is the synthetic point cloud data, is the second point cloud data, x is the value of the abscissa in P G and y is the value of the ordinate in .

[0070] The CD evaluates similarity by calculating the shortest distance between points, and the optimization direction is to make the synthetic point cloud as geometrically close to the real point cloud as possible. However, less consideration is given to the semantic information of the point cloud and the relative relationship between features. To further narrow the semantic consistency between the synthetic point cloud features and the point cloud features in the feature space, a contrastive loss is introduced, random point cloud features are obtained, and positive sample pairs (synthetic point cloud features and point cloud features) and negative sample pairs (synthetic point cloud features and random point cloud features) are constructed. Based on the synthetic point cloud features and random point cloud features generate a fourth loss value L contrast , and the fourth loss value L contrast is as follows:

[0071]

[0072] wherein, is the point cloud feature, is the synthetic point cloud feature, is the random point cloud feature (which is also the point cloud feature of other randomly selected scenes), τ is the temperature coefficient, τ is used to control the distribution smoothness, sim() is the cosine similarity, and both k and K are positive integers.

[0073] For the first loss value the third loss value and the fourth loss value L contrast perform weighted summation to obtain the target loss value L, and the target loss value L is as follows:

[0074]

[0075] wherein, α is the first weight, β is the second weight, γ is the third weight, and α, β, and γ can all be set by developers or users or automatically generated adaptively, and the present invention does not make any limitations in this regard. The loss value can include the first loss value the second loss value the third loss value the fourth loss value L contrast and the target loss value L.

[0076] S1224. Optimize the model parameters in the conditional generative adversarial network model in the direction of reducing the target loss value to obtain the target conditional generative adversarial network model.

[0077] In the embodiment of the present invention, during the backpropagation process, the chain rule is used to calculate the first gradient of the target loss value L with respect to the generator parameter θ G , and the first gradient is as follows:

[0078]

[0079] Among them, is the first loss value, is the third loss value, L contrast is the fourth loss value, θ G is the generator parameter, L is the target loss value, α is the first weight, β is the second weight, and γ is the third weight.

[0080] Calculate the second loss value For the discriminator parameter θ D of the second gradient, the second gradient is as follows:

[0081]

[0082] Among them, is the second loss value, is the synthetic point cloud feature, is the point cloud feature, θ D is the discriminator parameter.

[0083] In the direction of reducing the target loss value, use the gradient descent method to optimize the model parameters in the conditional generative adversarial network model. The gradient descent method is as follows:

[0084]

[0085]

[0086] Among them, θ G is the generator parameter, θ D is the discriminator parameter, L is the target loss value, is the second loss value, η1 is the first learning rate, and η2 is the second learning rate.

[0087] S123. Based on the target conditional generative adversarial network model, process the second point cloud data and the RGB image data to generate target synthetic point cloud data.

[0088] S130. Generate a double Gaussian distribution synthetic point cloud feature according to the target synthetic point cloud data and the second point cloud data, and process the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud feature to obtain a fused point cloud feature.

[0089] As Figure 5 shown, in the embodiment of the present invention, Figure 5 the step S130 at least includes the following steps:

[0090] S131. Perform Poisson disk sampling on the target synthetic point cloud data and the second point cloud data respectively to obtain sampled synthetic point cloud data and third point cloud data.

[0091] In the embodiment of the present invention, the texture information on the surface of the mine area obstacles is an important feature for distinguishing the surrounding mine soil. In order to generate higher-quality point cloud data that retains the texture information in the image data, the Poisson disk sampling method is adopted. Poisson disk sampling is respectively performed on the target synthetic point cloud data and the second point cloud data, and the Poisson disk sampling process is as follows: Select an initial point p0 as the sampling point, add the initial point p0 to the sampled point set S, and add the initial point p0 to the active point set A. In each iteration, randomly select a point p from the active point set A, randomly generate k candidate points within a spherical region centered at the point p with a radius of 2r. For each candidate point q, check the distance between the candidate point q and all points in the sampled point set S. When the following conditions are met:

[0092]

[0093] where S is the sampled point set, p is a randomly selected point from the active point set A, q is the candidate point, and r is the radius.

[0094] When the above conditions are met, the candidate point q is added to the active point set A and the sampled point set S. When none of the k candidate points meet the above conditions, the point p is removed from the active point set A, and the above iteration is repeated until the active point set A is empty. The sampled synthetic point cloud data and the third point cloud data

[0095] S132. Feature extraction is respectively performed on the sampled synthetic point cloud data and the third point cloud data to obtain the sampled synthetic point cloud feature and the third point cloud feature.

[0096] In the embodiment of the present invention, feature extraction is respectively performed on the sampled synthetic point cloud data and the third point cloud data through a 3D backbone network, and the high-dimensional point cloud data is mapped to a low-dimensional feature space to obtain the sampled synthetic point cloud feature and the third point cloud feature so as to capture the local and global information of the point cloud for subsequent tasks. The multilayer perceptron (MLP) enhances the expression ability of the features through nonlinear transformation and improves the robustness of the model to noise and changes. Among them, the backbone can be PointNet, PointNet++, VoxelNet, PointPillars, etc.

[0097] S133. BEV projection is respectively performed on the sampled synthetic point cloud feature and the third point cloud feature to obtain the BEV synthetic point cloud feature and the BEV point cloud feature.

[0098] In the embodiments of the present invention, the sampled synthetic point cloud features are compressed and projected along the Z-axis onto the BEV plane, and the point features within each BEV grid cell are aggregated to obtain the BEV synthetic point cloud features. The third point cloud features are compressed and projected along the Z-axis onto the BEV plane, and the point features within each BEV grid cell are aggregated to obtain the BEV point cloud features. For each point (x, y, z) in the point cloud, its corresponding grid cell is determined according to its x and y coordinates, and the x and y coordinates of the point are normalized to the grid index. The normalization formula is as follows:

[0099]

[0100] where x min is the minimum value of the x coordinate of the point cloud to be projected, ymin is the minimum value of the y coordinate of the point cloud to be projected, and resolution is the resolution of the grid.

[0101] S134. Perform double Gaussian distribution sampling on the BEV synthetic point cloud features and the BEV point cloud features to obtain the double Gaussian distribution synthetic point cloud features.

[0102] In the embodiments of the present invention, the BEV synthetic point cloud features and the BEV point cloud features are sampled through double Gaussian distribution sampling to obtain the double Gaussian distribution synthetic point cloud features. The double Gaussian distribution (Mixture of Gaussians, MoG) is a probability model and is often used to model complex data distributions. In point cloud processing tasks, the double Gaussian distribution can be used to model the spatial distribution of point cloud features. The mean and covariance of the BEV point cloud features are (μ1, Σ1), and the mean and covariance of the BEV synthetic point cloud features are (μ2, Σ2). The double Gaussian distribution probability model is constructed as the following formula:

[0103]

[0104] where π1 is the first mixing coefficient, π2 is the second mixing coefficient, satisfying π1 + π2 = 1, (μ1, Σ1) are the mean and covariance of the BEV point cloud features, (μ2, Σ2) are the mean and covariance of the BEV synthetic point cloud features, and p(F) is the double Gaussian distribution synthetic point cloud data, and is the Gaussian distribution.

[0105] The double Gaussian distribution synthetic point cloud features sampled from the double Gaussian distribution are as the following formula:

[0106] F sample ~p(F) Formula (17)

[0107] The double Gaussian distribution synthetic point cloud features F sampleIt represents the multimodal distribution characteristics that can capture the BEV point cloud features and the BEV synthetic point cloud features, so as to better fuse the information of the two, be used for subsequent generation of diverse fusion features, and enhance the generalization ability of the model.

[0108] S135. Stack the BEV synthetic point cloud features, the BEV point cloud features, and the double Gaussian distribution synthetic point cloud features to obtain the fused point cloud features.

[0109] In the embodiment of the present invention, the double Gaussian distribution synthetic point cloud feature F sample , F sample has the shape of (B, C, H, W). Stack the BEV synthetic point cloud features, the BEV point cloud features, and the double Gaussian distribution synthetic point cloud features in the channel dimension. The shape of the stacked features is (B, 3*C, H, W). After passing through a 3*3 two-dimensional convolutional layer for the stacked features, the shape of the output fused point cloud feature is (B, C out , H, W). Among them, C out can be determined according to the requirements for the fused point cloud features. The BEV synthetic point cloud features, the BEV point cloud features, and the double Gaussian distribution synthetic point cloud features are spatially aligned. After stacking, a convolutional operation is directly performed in the spatial dimension, combining the RGB image data and the first point cloud data, and enhancing the feature expression ability.

[0110] In summary, in the method for fusing image and point cloud features based on a conditional generative adversarial network in this application, the first point cloud data and the RGB image data are obtained, the first point cloud data is preprocessed to obtain the second point cloud data, the target conditional generative adversarial network model is called, the target synthetic point cloud data is generated based on the second point cloud data and the RGB image data, the double Gaussian distribution synthetic point cloud feature is generated according to the target synthetic point cloud data and the second point cloud data, and the fused point cloud feature is obtained according to the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud feature, thus solving the problem of how to effectively fuse the image data and the point cloud data, and improving the performance of the intelligent driving system in the unmanned mining area. At the same time, by training the conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain the target conditional generative adversarial network model, and processing the second point cloud data and the RGB image data based on the target conditional generative adversarial network model to generate the target synthetic point cloud data, the unified data modality is realized, thus simplifying the subsequent feature extraction and fusion process.

[0111] Please refer to Figure 6 , which is a schematic structural diagram of the image and point cloud feature fusion system based on a conditional generative adversarial network disclosed in the embodiment of the present application. In one embodiment, as Figure 6As shown in the figure, the present application provides an image and point cloud feature fusion system 100 based on a conditional generative adversarial network. The image and point cloud feature fusion system 100 based on a conditional generative adversarial network can at least include: a preprocessing module 110, an image data generation module 130, and an image feature fusion module 150. Among them, there is information interaction between the preprocessing module 110 and the image data generation module 130, and there is information interaction between the image data generation module 130 and the image feature fusion module 150.

[0112] The preprocessing module 110 is used to obtain the first point cloud data and RGB image data, and preprocess the first point cloud data to obtain the second point cloud data. Specifically, a lidar (Light Laser Detection and Ranging, LiDAR) with high-precision and long-distance detection capabilities is installed on a vehicle in an unmanned mining area. The lidar is used to collect the first point cloud data P of the entire scene centered on the vehicle itself. where represents the i1-th first point cloud data, and the first point cloud data can include clear surface point cloud data of target recognition obstacles. At the same time, a camera is installed on the vehicle in the unmanned mining area, and the camera is used to collect RGB image data X. where represents the i2-th RGB image data. Since the first point cloud data covers the entire unmanned mining area at 360°, while the RGB image data only includes the target recognition obstacles within the viewing angle and the surrounding part of the environment, it is necessary to preprocess the first point cloud data to the second point cloud data where represents the i3-th second point cloud data, so that the viewing angle areas of the second point cloud data and the RGB image data are basically coincident. Among them, the preprocessing can include noise reduction, cropping, and adjusting the viewing angle.

[0113] The image data generation module 130 is used to call a target conditional generative adversarial network model to generate target synthetic point cloud data based on the second point cloud data and the RGB image data. Among them, the image data generation module 130 can at least include: a feature extraction module 131, a model training module 133, and a generation module 135.

[0114] The feature extraction module 131 is used to extract point cloud features based on the second point cloud data and extract RGB image features based on the RGB image data.

[0115] The model training module 133 is used to train a conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain a target conditional generative adversarial network model. Among them, the model training module 133 can at least include: a first generation unit 1331, a discriminant unit 1333, a calculation unit 1335, and a parameter optimization unit 1336.

[0116] The first generation unit 1331 is used to obtain a random noise vector, generate synthetic point cloud features based on the random noise vector, the point cloud features, and the RGB image features, and obtain synthetic point cloud data according to the synthetic point cloud features. Specifically, a conditional generative adversarial network model includes a generator G and a discriminator D. The generator can be a network structure including an encoder-decoder. The encoder can include convolutional layers and pooling layers. The convolutional layers are used to capture local features, while the pooling layers are used to reduce the spatial dimension of the features. At the same time, the pooling layers are also used to retain important semantic information, such as color, texture, and shape features in the RGB image features. The decoder performs deconvolution on the features extracted by the encoder, and finally outputs a tensor with a dimension of [(point_size * 3), 1] through a fully connected layer, where point_size is the number of generated point clouds, and 3 refers to the three-dimensional coordinates (x, y, z) of each point cloud.

[0117] The generator receives the random noise vector and the RGB image feature F X as inputs, generates initial synthetic point cloud features, and the initial synthetic point cloud features are processed through a fully connected layer to obtain initial synthetic point cloud data Initial synthetic point cloud data as the following formula:

[0118]

[0119] where F X is the RGB image feature, is the random noise vector, is the initial synthetic point cloud data.

[0120] At the same time, the point cloud feature is used as a supervision signal. A cross-attention mechanism is introduced into the generator. By applying the supervision signal to the initial synthetic point cloud features generated by the generator, the generator can dynamically adjust the geometric structure of the point cloud to make it closer to the point cloud feature. When the conditional generative adversarial network model is completed training, the point cloud feature does not need to be used as a supervision signal anymore. Based on the random noise vector and the RGB image feature, synthetic point cloud features can be directly generated, and synthetic point cloud data can be obtained according to the synthetic point cloud features. The cross-attention mechanism is as the following formula:

[0121]

[0122] Among them, Q is the query vector, and the query vector Q is extracted from the point cloud features K is the key vector, and the key vector K is extracted from the RGB image feature F X V is the value vector, and the value vector V is extracted from the RGB image feature F X QK T is the similarity between Q and K, and d k is the dimension of the key, and d k is used as a scaling factor to avoid excessive numerical values.

[0123] After introducing the cross-attention mechanism, the initial synthetic point cloud features are weighted and corrected to obtain the corrected synthetic point cloud features, and the corrected synthetic point cloud features are processed through a fully connected layer to obtain the corrected synthetic point cloud data to guide the generator G to generate a point cloud that is more in line with the real situation. Among them, the synthetic point cloud features can include the initial synthetic point cloud features and the corrected synthetic point cloud features The synthetic point cloud data P G can include the initial synthetic point cloud data and the corrected synthetic point cloud data The corrected synthetic point cloud data is as follows:

[0124]

[0125] Among them, is a random noise vector, and CrossAttention(Q, K, V) is the cross-attention mechanism.

[0126] The discriminant unit 1333 is used to perform discriminant processing based on the synthetic point cloud features and the point cloud features to obtain a discriminant result. Specifically, the initial synthetic point cloud features are processed through a fully connected layer to obtain the initial synthetic point cloud data The discriminator D receives the second point cloud data and the corrected synthetic point cloud data, and performs discriminant processing on the authenticity of the corrected synthetic point cloud data to obtain a discriminant result The discriminant result of the discriminator D is as follows:

[0127]

[0128] Taking the discriminant result as 0 or 1 as an example, when is 1, it means that the generation effect of the generator G is completely real; when is 0, it means that the generation effect of the generator G is not good.

[0129] The calculation unit 1335 is used to obtain the random point cloud features, generate the first loss value based on the synthesized point cloud features, generate the third loss value based on the synthesized point cloud data and the second point cloud data, generate the fourth loss value based on the synthesized point cloud features and the random point cloud features, and perform weighted summation on the first loss value, the third loss value and the fourth loss value to obtain the target loss value. Specifically, based on the synthesized point cloud features Generate the first loss value The first loss value As shown in the following formula:

[0130]

[0131] Where G is the generator, is the synthesized point cloud feature.

[0132] Based on the point cloud feature and the synthesized point cloud feature Generate the second loss value The second loss value As shown in the following formula:

[0133]

[0134] Where D is the discriminator, is the point cloud feature, is the synthesized point cloud feature.

[0135] In order to measure the point cloud synthesis effect, the Chamfer Distance (CD) index is introduced. The performance of the generator is optimized by measuring the bidirectional nearest neighbor distance between the synthesized point cloud data and the second point cloud data through CD. Based on the synthesized point cloud data P G and the second point cloud data Generate the third loss value The third loss value As shown in the following formula:

[0136]

[0137] Where P G is the synthesized point cloud data, is the second point cloud data, x is the value of the abscissa in P G and y is the value of the ordinate in

[0138] ​The CD evaluates similarity by calculating the shortest distance between points. The optimization direction is to make the synthetic point cloud as geometrically close to the real point cloud as possible, but less consideration is given to the semantic information of the point cloud and the relative relationship between features. To further narrow the semantic consistency between the synthetic point cloud features and the point cloud features in the feature space, a contrastive loss is introduced, and positive sample pairs (synthetic point cloud features and point cloud features) and negative sample pairs (synthetic point cloud features and random point cloud features) are constructed. Based on the synthetic point cloud features and random point cloud features generate a fourth loss value L contrast The fourth loss value L contrast is as follows:

[0139]

[0140] where is the point cloud feature, is the synthetic point cloud feature, is the random point cloud feature (i.e., the point cloud feature of other randomly selected scenes), τ is the temperature coefficient, τ is used to control the distribution smoothness, sim() is the cosine similarity, and both k and K are positive integers.

[0141] For the first loss value the third loss value and the fourth loss value L contrast perform a weighted sum to obtain the target loss value L. The target loss value L is as follows:

[0142]

[0143] where α is the first weight, β is the second weight, γ is the third weight, and α, β, and γ can all be set by developers or users or automatically generated adaptively. The present invention does not make any limitations in this regard.

[0144] The parameter optimization unit 1336 is used to optimize the model parameters in the conditional generative adversarial network model in the direction of reducing the target loss value to obtain the target conditional generative adversarial network model. Specifically, in the backpropagation process, the chain rule is used to calculate the first gradient of the target loss value L with respect to the generator parameter θ G The first gradient is as follows:

[0145]

[0146] where is the first loss value, is the third loss value, L contrast is the fourth loss value, θ G is the generator parameter, L is the target loss value, α is the first weight, β is the second weight, and γ is the third weight.

[0147] Calculate the second loss value For the discriminator parameter θ D The second gradient, and the second gradient is as follows:

[0148]

[0149] Wherein, is the second loss value, is the synthetic point cloud feature, is the point cloud feature, and θ D is the discriminator parameter.

[0150] In the direction of reducing the target loss value, use the gradient descent method to optimize the model parameters in the conditional generative adversarial network model. The gradient descent method is as follows:

[0151]

[0152]

[0153] Wherein, θ G is the generator parameter, θ D is the discriminator parameter, L is the target loss value, is the second loss value, η1 is the first learning rate, and η2 is the second learning rate.

[0154] The generation module 135 is configured to process the second point cloud data and the RGB image data based on the target conditional generative adversarial network model to generate target synthetic point cloud data.

[0155] The image feature fusion module 150 is configured to generate a double Gaussian distribution synthetic point cloud feature according to the target synthetic point cloud data and the second point cloud data, and process the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud feature to obtain a fused point cloud feature. Among them, the image feature fusion module 150 may at least include: a first sampling module 151, a feature extraction module 153, a projection module 155, a second sampling module 157, and a fusion module 159.

[0156] The first sampling module 151 is used to perform Poisson disk sampling on the target synthetic point cloud data and the second point cloud data respectively, to obtain sampled synthetic point cloud data and third point cloud data. Specifically, the texture information on the surface of the mining area obstacles is an important feature for distinguishing it from the surrounding ore soil. In order to generate higher-quality point cloud data that retains the texture information in the image data, the Poisson disk sampling method is adopted. Poisson disk sampling is performed on the target synthetic point cloud data and the second point cloud data respectively, and the Poisson disk sampling process is as follows: Select an initial point p0 as the sampling point, add the initial point p0 to the sampled point set S, and add the initial point p0 to the active point set A. In each iteration, randomly select a point p from the active point set A, randomly generate k candidate points within a spherical region centered at point p with a radius of 2r. For each candidate point q, check the distance between the candidate point q and all points in the sampled point set S. When the following conditions are met:

[0157]

[0158] where S is the sampled point set, p is a randomly selected point from the active point set A, q is the candidate point, and r is the radius.

[0159] When the above conditions are met, add the candidate point q to the active point set A and the sampled point set S. When none of the k candidate points meet the above conditions, remove the point p from the active point set A, and repeat the above iteration until the active point set A is empty. After the first sampling process, the sampled synthetic point cloud data and the third point cloud data

[0160] The feature extraction module 153 is used to perform feature extraction on the sampled synthetic point cloud data and the third point cloud data respectively, to obtain sampled synthetic point cloud features and third point cloud features. Specifically, the 3D backbone network is used to perform feature extraction on the sampled synthetic point cloud data and the third point cloud data respectively, map the high-dimensional point cloud data to a low-dimensional feature space, and obtain the sampled synthetic point cloud features and the third point cloud features so as to capture the local and global information of the point cloud for subsequent tasks. The multilayer perceptron (MLP) enhances the expression ability of the features through non-linear transformation, and improves the robustness of the model to noise and changes. Among them, the Backbone can be PointNet, PointNet++, VoxelNet, PointPillars, etc.

[0161] The projection module 155 is used to perform BEV projection on the sampled synthetic point cloud features and the third point cloud features respectively, to obtain BEV synthetic point cloud features and BEV point cloud features. Specifically, the sampled synthetic point cloud features Compress and project it onto the BEV plane along the Z-axis, and aggregate the point features within each BEV grid cell to obtain the BEV synthetic point cloud feature. Compress and project the third point cloud feature onto the BEV plane along the Z-axis, and aggregate the point features within each BEV grid cell to obtain the BEV point cloud feature. For each point (x, y, z) in the point cloud, determine the grid cell it belongs to according to its x and y coordinates, and normalize the x and y coordinates of the point to the grid index. The normalization formula is as follows:

[0162]

[0163] where x min is the minimum value of the x coordinate of the point cloud to be projected, y min is the minimum value of the y coordinate of the point cloud to be projected, and resolution is the resolution of the grid.

[0164] The second sampling module 157 is used to perform double Gaussian distribution sampling on the BEV synthetic point cloud feature and the BEV point cloud feature to obtain the double Gaussian distribution synthetic point cloud feature. Specifically, sample the BEV synthetic point cloud feature and the BEV point cloud feature through double Gaussian distribution sampling to obtain the double Gaussian distribution synthetic point cloud feature. The mixture of Gaussians (MoG) is a probability model commonly used to model complex data distributions. In point cloud processing tasks, the mixture of Gaussians can be used to model the spatial distribution of point cloud features. The mean and covariance of the BEV point cloud feature are (μ1, Σ1), and the mean and covariance of the BEV synthetic point cloud feature are (μ2, Σ2). Construct the double Gaussian distribution probability model as the following formula:

[0165]

[0166] where π1 is the first mixing coefficient, π2 is the second mixing coefficient, satisfying π1 + π2 = 1, (μ1, Σ1) are the mean and covariance of the BEV point cloud feature, and (μ2, Σ2) are the mean and covariance of the BEV synthetic point cloud feature.

[0167] The double Gaussian distribution synthetic point cloud feature sampled from the double Gaussian distribution is as the following formula:

[0168] F sample ~p(F) Formula (17)

[0169] The double Gaussian distribution synthetic point cloud feature F sample represents the multimodal distribution characteristics that can capture the BEV point cloud feature and the BEV synthetic point cloud feature, so as to better fuse the information of both, be used to generate diverse fusion features subsequently, and enhance the generalization ability of the model.

[0170] The fusion module 159 is used to stack the BEV synthetic point cloud features, the BEV point cloud features, and the double Gaussian distribution synthetic point cloud features to obtain the fused point cloud features. Specifically, the double Gaussian distribution synthetic point cloud feature F sample , F sample has a shape of (B, C, H, W). The BEV synthetic point cloud features, the BEV point cloud features, and the double Gaussian distribution synthetic point cloud features are stacked in the channel dimension, and the shape of the stacked features is (B, 3*C, H, W). After passing through a 3*3 two-dimensional convolutional layer for the stacked features, the shape of the output fused point cloud features is (B, C out , H, W). Among them, C out can be determined according to the requirements for the fused point cloud features. The BEV synthetic point cloud features, the BEV point cloud features, and the double Gaussian distribution synthetic point cloud features are spatially aligned, and a convolutional operation is directly performed in the spatial dimension after stacking, combining the RGB image data and the first point cloud data to enhance the feature expression ability.

[0171] In summary, in the image and point cloud feature fusion system based on the conditional generative adversarial network of the present application, the first point cloud data and the RGB image data are obtained through the preprocessing module 110, the first point cloud data is preprocessed to obtain the second point cloud data, the image data generation module 130 calls the target conditional generative adversarial network model, and the target synthetic point cloud data is generated based on the second point cloud data and the RGB image data. The image feature fusion module 150 generates the double Gaussian distribution synthetic point cloud features according to the target synthetic point cloud data and the second point cloud data, and processes to obtain the fused point cloud features according to the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud features, thereby solving the problem of how to effectively fuse the image data and the point cloud data, and improving the performance of the intelligent driving system in the unmanned mining area.

[0172] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one implementation manner", "one preferred implementation manner" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0173] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. An image and point cloud feature fusion method based on a conditional generative adversarial network, characterized in that The method for fusing image and point cloud features based on a conditional generative adversarial network includes: Obtain first point cloud data and RGB image data, and preprocess the first point cloud data to obtain second point cloud data; Call a target conditional generative adversarial network model, and generate target synthetic point cloud data based on the second point cloud data and the RGB image data; Generate a double Gaussian distribution synthetic point cloud feature according to the target synthetic point cloud data and the second point cloud data, and process the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud feature to obtain a fused point cloud feature.

2. The method for fusing image and point cloud features based on conditional generative adversarial network according to claim 1, wherein The step of calling a target conditional generative adversarial network model and generating target synthetic point cloud data based on the second point cloud data and the RGB image data includes: Extract point cloud features based on the second point cloud data, and extract RGB image features based on the RGB image data; Train a conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain a target conditional generative adversarial network model; Based on the target conditional generative adversarial network model, process the second point cloud data and the RGB image data to generate target synthetic point cloud data.

3. The method for fusing image and point cloud features based on a conditional generative adversarial network according to claim 2, characterized in that The step of training a conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain a target conditional generative adversarial network model includes: Obtain a random noise vector, generate a synthetic point cloud feature based on the random noise vector, the point cloud features, and the RGB image features, and obtain synthetic point cloud data according to the synthetic point cloud feature; Perform a discrimination process based on the synthetic point cloud feature and the point cloud feature to obtain a discrimination result; Obtain random point cloud features, generate a first loss value based on the synthetic point cloud feature, generate a third loss value based on the synthetic point cloud data and the second point cloud data, generate a fourth loss value based on the synthetic point cloud feature and the random point cloud features, and perform a weighted sum of the first loss value, the third loss value, and the fourth loss value to obtain a target loss value; Optimize the model parameters in the conditional generative adversarial network model in the direction of reducing the target loss value to obtain the target conditional generative adversarial network model.

4. The method for fusing image and point cloud features based on conditional generative adversarial network according to claim 3, wherein The step of generating a double Gaussian distribution synthetic point cloud feature according to the target synthetic point cloud data and the second point cloud data, and processing the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud feature to obtain a fused point cloud feature includes: Perform Poisson disk sampling on the target synthetic point cloud data and the second point cloud data respectively to obtain sampled synthetic point cloud data and third point cloud data; Extract features from the sampled synthetic point cloud data and the third point cloud data respectively to obtain sampled synthetic point cloud features and third point cloud features; Perform BEV projection on the sampled synthetic point cloud features and the third point cloud features respectively to obtain BEV synthetic point cloud features and BEV point cloud features; Performing double Gaussian distribution sampling on the BEV synthetic point cloud feature and the BEV point cloud feature to obtain the double Gaussian distribution synthetic point cloud feature; Stacking the BEV synthetic point cloud feature, the BEV point cloud feature, and the double Gaussian distribution synthetic point cloud feature to obtain the fused point cloud feature.

5. The method for fusing image and point cloud features based on conditional generative adversarial network according to claim 3, wherein Performing weighted summation on the first loss value, the third loss value, and the fourth loss value to obtain the target loss value, where the target loss value is: where α is the first weight, β is the second weight, and γ is the third weight, is the first loss value, is the third loss value, P G is the synthesized point cloud data, is the second point cloud data, L contrast is the fourth loss value.

6. The method for fusing image and point cloud features based on conditional generative adversarial network according to claim 5, characterized in that Performing double Gaussian distribution sampling on the BEV synthetic point cloud feature and the BEV point cloud feature to obtain the double Gaussian distribution synthetic point cloud feature, where the double Gaussian distribution synthetic point cloud feature is: F sample ~p(F) Where, π1 is the first mixing coefficient, π2 is the second mixing coefficient, π1 + π2 = 1, (μ1, Σ1) are the mean and covariance of the BEV point cloud features, (μ2, Σ2) are the mean and covariance of the BEV synthetic point cloud features, and p(F) is the double Gaussian distribution synthetic point cloud data. is a Gaussian distribution.

7. An image and point cloud feature fusion system based on a conditional generative adversarial network, characterized in that The image and point cloud feature fusion system based on the conditional generative adversarial network includes: a preprocessing module, an image data generation module, and an image feature fusion module, where The preprocessing module is used to obtain the first point cloud data and the RGB image data, and preprocess the first point cloud data to obtain the second point cloud data; The image data generation module is used to call the target conditional generative adversarial network model to generate the target synthetic point cloud data based on the second point cloud data and the RGB image data; The image feature fusion module is used to generate the double Gaussian distribution synthetic point cloud feature according to the target synthetic point cloud data and the second point cloud data, and process the target synthetic point cloud data, the second point cloud data, and the double Gaussian distribution synthetic point cloud feature to obtain the fused point cloud feature.

8. The image and point cloud feature fusion system based on conditional generative adversarial network according to claim 7, wherein The image data generation module includes: a feature extraction module, a model training module, and a generation module, where The feature extraction module is used to extract features from the second point cloud data to obtain point cloud features, and extract features from the RGB image data to obtain RGB image features; The model training module is used to train the conditional generative adversarial network model based on the point cloud features and the RGB image features to obtain the target conditional generative adversarial network model; The generation module is used to process the second point cloud data and the RGB image data based on the target conditional generative adversarial network model to generate the target synthetic point cloud data.

9. The image and point cloud feature fusion system based on conditional generative adversarial network according to claim 8, wherein The model training module includes: a first generation unit, a discriminant unit, a calculation unit, and a parameter optimization unit, where The first generation unit is used to obtain a random noise vector, generate synthetic point cloud features based on the random noise vector, the point cloud features, and the RGB image features, and obtain synthetic point cloud data according to the synthetic point cloud features; The discriminant unit is used to perform discriminant processing based on the synthetic point cloud features and the point cloud features to obtain a discriminant result; The calculation unit is used to obtain random point cloud features, generate a first loss value based on the synthetic point cloud features, generate a third loss value based on the synthetic point cloud data and the second point cloud data, generate a fourth loss value based on the synthetic point cloud features and the random point cloud features, and perform weighted summation on the first loss value, the third loss value, and the fourth loss value to obtain the target loss value; The parameter optimization unit is configured to optimize the model parameters in the conditional generative adversarial network model in the direction of reducing the target loss value to obtain the target conditional generative adversarial network model.

10. The image and point cloud feature fusion system based on conditional generative adversarial network according to claim 9, wherein The image feature fusion module includes a first sampling module, a feature extraction module, a projection module, a second sampling module, and a fusion module, where The first sampling module is configured to perform Poisson disk sampling on the target synthetic point cloud data and the second point cloud data respectively to obtain sampled synthetic point cloud data and third point cloud data; The feature extraction module is configured to perform feature extraction on the sampled synthetic point cloud data and the third point cloud data respectively to obtain sampled synthetic point cloud features and third point cloud features; The projection module is configured to perform BEV projection on the sampled synthetic point cloud features and the third point cloud features respectively to obtain BEV synthetic point cloud features and BEV point cloud features; The second sampling module is configured to perform double Gaussian distribution sampling on the BEV synthetic point cloud features and the BEV point cloud features to obtain the double Gaussian distribution synthetic point cloud features; The fusion module is configured to perform stacking processing on the BEV synthetic point cloud features, the BEV point cloud features, and the double Gaussian distribution synthetic point cloud features to obtain the fused point cloud features.

Citation Information

Patent Citations

  • Three-dimensional target detection method based on multi-modal fusion and deformable attention

    CN117975436A

  • Large-scale point cloud completion method and device

    CN119559093A

  • Method and apparatus for fusing image with point cloud

    EP4481592A1