Point cloud network training method, point cloud data processing method and device

Through the multimodal joint training method, cross-modal alignment is achieved using regional-level text features, and the representation ability of point cloud network is enhanced, which solves the problem of low accuracy of three-dimensional feature extraction in the existing technology, and realizes efficient three-dimensional point cloud understanding.

CN120299054AActive Publication Date: 2025-07-11PENG CHENG LAB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510781004.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The self-supervised learning method based on image-point cloud in the prior art is susceptible to limited receptive fields, limited context information, and potential noise points or wrong pixel semantics, resulting in a low accuracy of three-dimensional feature extraction and affecting the accuracy of three-dimensional scene understanding.

Method used

A multimodal joint training method is introduced, and cross-modal alignment is achieved using regional-level text features. Through the self-supervised comparison learning framework of image-text-point cloud, the representation ability of point cloud network is enhanced, the cross-modal feature distance of the same semantic region is narrowed and the features of different semantic regions are pushed apart, and the global feature blurring is avoided.

Benefits of technology

It significantly improves the feature extraction accuracy of 3D point clouds and improves the accuracy of downstream 3D point cloud understanding, without relying on manual annotation point cloud data to achieve low-cost and high-efficiency three-dimensional spatial understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299054A_ABST
    Figure CN120299054A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a point cloud network training method and device and a point cloud data processing method and device, and relates to the technical field of unmanned driving and intelligent transportation. According to the method, three-dimensional point cloud features of laser radar point cloud data are obtained by using a point cloud network, two-dimensional image features of a multi-view image are obtained by using an image encoder, mask division is performed on the multi-view image by using a visual basic model to obtain a plurality of object areas, text features corresponding to each object area are generated by using a text encoder, and the text features of the multi-view image are obtained. And performing image-text-point cloud contrast learning according to the text features, the two-dimensional region pixel features corresponding to each object region and the three-dimensional region point cloud features corresponding to each object region, generating a semantic region contrast loss value, and training the point cloud network. Multi-modal data is used for joint training, region-level text features are used for realizing fine-grained cross-modal feature alignment, a self-supervised contrast learning framework based on image-text-point cloud enhances the characterization capability of a point cloud network, and the understanding precision of downstream three-dimensional point cloud is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of driverless and intelligent transportation, and particularly to a point cloud network training method, a point cloud data processing method, and an apparatus. Background Art

[0002] The three-dimensional scene understanding based on lidar point cloud is widely used in scenarios such as autonomous driving and robot navigation. It uses algorithms or models to extract features from the input three-dimensional point cloud, and performs semantic understanding of the three-dimensional scene content according to the extraction results. The models related to the three-dimensional scene understanding based on lidar point cloud usually undergo a supervised learning training process based on manually annotated point clouds. However, the annotation cost of annotating point cloud data is relatively high, which limits the applicable scope of the three-dimensional scene understanding.

[0003] In the related art, the self-supervised learning method based on image-point cloud trains a three-dimensional point cloud network through a point-by-point comparison learning method between pixel points and point cloud points. However, this point-level learning method is easily affected by the limited receptive field, limited context information, and potential noise points or incorrect pixel semantics, resulting in a low accuracy of three-dimensional feature extraction and affecting the accuracy of three-dimensional scene understanding. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a point cloud network training method, a point cloud data processing method, and an apparatus, so as to improve the accuracy of the point cloud network in extracting features from lidar point cloud data.

[0005] To achieve the above purpose, a first aspect of the embodiments of this application proposes a point cloud network training method, which is executed by a self-supervised training system. The self-supervised training system at least includes an image encoder, a text encoder, a visual foundation model, and the point cloud network. Except for the point cloud network, other network parameters are frozen. The method includes: Obtain lidar point cloud data and corresponding multi-view images, input the lidar point cloud data into the point cloud network for feature extraction to obtain three-dimensional point cloud features, and input the multi-view images into the image encoder for feature extraction to obtain two-dimensional image features; Input the multi-view images into the visual foundation model for mask division to obtain at least one object region, and use the text encoder to generate text features corresponding to each object region; Associate the two-dimensional image features and the three-dimensional point cloud features according to a preset spatial projection relationship to obtain a point-pixel association result, and divide the point-pixel association result according to the object region to obtain two-dimensional region pixel features and three-dimensional region point cloud features corresponding to each object region; Perform image-text-point cloud contrastive learning based on the text features, the two-dimensional region pixel features, and the three-dimensional region point cloud features to generate a semantic region contrast loss value, and adjust the parameters of the point cloud network based on the semantic region contrast loss value until the trained point cloud network is obtained.

[0006] In some embodiments, the multi-view image includes scene images corresponding to at least one view, and the step of inputting the multi-view image into the vision base model for mask division to obtain at least one object region includes: Input the scene images into the vision base model for mask division respectively to obtain mask images corresponding to each scene image, where each mask image includes at least one mask region; Perform region division in each scene image based on the mask region to obtain at least one object region corresponding to each scene image.

[0007] In some embodiments, the self-supervised training system further includes a text generation from image base model, and the step of using the text encoder to generate text features corresponding to each object region includes: For each scene image, input the corresponding object region into the text generation from image base model for text generation to obtain a guiding language description corresponding to each object region; Input the guiding language description into the text encoder for feature extraction to obtain the corresponding text features.

[0008] In some embodiments, the step of performing image-text-point cloud contrastive learning based on the text features, the two-dimensional region pixel features, and the three-dimensional region point cloud features to generate a semantic region contrast loss value includes: Perform average pooling on the two-dimensional region pixel features to obtain region point features, and perform average pooling on the three-dimensional region point cloud features to obtain region pixel features; Calculate the semantic region contrast loss value according to the region point features, the region pixel features, and the text features.

[0009] In some embodiments, the step of performing contrastive learning according to the region point features, the region pixel features, and the text features to calculate the semantic region contrast loss value includes: For each object region, calculate an indication parameter between the region pixel features and the text features and other object regions, and determine a plurality of negative samples from other object regions based on the indication parameter; Calculate a region loss value based on the negative samples, the region point features corresponding to the object region, and the region pixel features, and obtain the semantic region contrast loss value according to the mean value of the region loss value.

[0010] In some embodiments, calculating an indication parameter between the object region and other object regions according to the region pixel features and the text features includes: Select a comparison region one by one from other object regions, and obtain the region pixel features of the comparison region as comparison pixel features; Calculate a first text similarity between the text features and the region pixel features, and calculate a second text similarity between the text features and the comparison pixel features; Obtain the difference between the first text similarity and the second text similarity. If the difference is greater than or equal to a preset reference value, set the indication parameter to one, which is used to indicate that the comparison region is a negative sample of the object region.

[0011] In some embodiments, calculating the region loss value based on the negative samples, the region point features corresponding to the object region, and the region pixel features includes: For each object region, calculate a first semantic similarity between the region point features and the region pixel features; Obtain the region pixel features of each negative sample as sample pixel features, calculate a sample semantic similarity between the region point features and each sample pixel feature, and calculate the sum of the sample semantic similarities as a second semantic similarity; Obtain the sum of the first semantic similarity and the second semantic similarity as an intermediate reference value, and obtain the region loss value according to the quotient of the first semantic similarity and the intermediate reference value.

[0012] In some embodiments, the self-supervised training system further includes a two-dimensional mapping layer and a three-dimensional mapping layer to be trained. Before dividing the two-dimensional image features according to the object region to obtain two-dimensional region pixel features corresponding to each object region, the method further includes: Input the two-dimensional image features into the two-dimensional mapping layer for channel dimension mapping to update the two-dimensional image features; Input the three-dimensional point cloud features into the three-dimensional mapping layer for channel dimension mapping to update the three-dimensional point cloud features. The updated three-dimensional point cloud features and the two-dimensional image features have the same channel dimension.

[0013] To achieve the above object, a second aspect of the embodiments of the present application proposes a three-dimensional point cloud data processing method, including: Obtain initial point cloud data; Input the initial point cloud number into the point cloud network for feature extraction to obtain point cloud feature data, where the point cloud network is trained by the point cloud network training method described in any item of the first aspect.

[0014] To achieve the above object, a third aspect of the embodiments of the present application proposes a point cloud network training device, which is executed by a self-supervised training system. The self-supervised training system at least includes an image encoder, a text encoder, a vision foundation model, and the point cloud network. Except for the point cloud network, other network parameters are frozen. The device includes: An initial feature extraction module: used to obtain lidar point cloud data and corresponding multi-view images, input the lidar point cloud data into the point cloud network for feature extraction to obtain three-dimensional point cloud features, and input the multi-view images into the image encoder for feature extraction to obtain two-dimensional image features; A text description generation module: used to input the multi-view images into the vision foundation model for mask division to obtain at least one object region, and use the text encoder to generate text features corresponding to each object region; A mapping association module: used to associate the two-dimensional image features and the three-dimensional point cloud features according to a preset spatial projection relationship to obtain a point-pixel association result, and divide the point-pixel association result according to the object region to obtain two-dimensional region pixel features and three-dimensional region point cloud features corresponding to each object region; A contrast learning module: used to perform image-text-point cloud contrast learning according to the text features, the two-dimensional region pixel features, and the three-dimensional region point cloud features to generate a semantic region contrast loss value, and adjust the parameters of the point cloud network based on the semantic region contrast loss value until the trained point cloud network is obtained.

[0015] To achieve the above object, a fourth aspect of the embodiments of the present application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect or the second aspect is implemented.

[0016] To achieve the above object, a fifth aspect of the embodiments of the present application proposes a storage medium, which is a storage medium that stores a computer program. When the computer program is executed by a processor, the method described in the first aspect or the second aspect is implemented.

[0017] The point cloud network training method, point cloud data processing method and device proposed in the embodiment of the present application obtain laser radar point cloud data and corresponding multi-view images, input the laser radar point cloud data into the point cloud network for feature extraction, obtain three-dimensional point cloud features, and input the multi-view images into the image encoder for feature extraction to obtain two-dimensional image features, input the multi-view images into the visual basic model for mask division to obtain at least one object area, and use the text encoder to generate text features corresponding to each object area, associate the two-dimensional image features with the three-dimensional point cloud features according to the preset spatial projection relationship to obtain point-pixel association results, divide the point-pixel association results according to the object area, obtain the two-dimensional region pixel features and three-dimensional region point cloud features corresponding to each object area, perform image-text-point cloud comparison learning according to the text features, two-dimensional region pixel features and three-dimensional region point cloud features, generate semantic region comparison loss values, and adjust the parameters of the point cloud network based on the semantic region comparison loss values ​​until the trained point cloud network is obtained. The embodiment of the present application introduces multimodal joint training and uses region-level text features to achieve cross-modal alignment. Based on the self-supervised contrastive learning framework of image-text-point cloud, the representation ability of the point cloud network is enhanced, and the local semantic consistency is enhanced during the training process: by shortening the cross-modal feature distance of the same semantic area and pushing the features of different semantic areas apart, fine-grained cross-modal feature alignment is achieved, avoiding global feature blurring, and significantly improving the feature extraction accuracy of 3D point clouds, thereby improving the accuracy of downstream 3D point cloud understanding. There is no need to rely on manually annotated point cloud data, and low-cost and high-efficiency 3D space understanding is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram of the training process of the self-supervised training system provided in an embodiment of the present application.

[0019] Figure 2 It is a flow chart of the point cloud network training method provided in an embodiment of the present application.

[0020] Figure 3 This is a flowchart provided by an embodiment of the present application for inputting a multi-view image into a visual basic model for mask division to obtain at least one object area.

[0021] Figure 4 It is a schematic diagram of the object area provided in the embodiment of the present application.

[0022] Figure 5 This is a flowchart of using a text encoder to generate text features corresponding to each object area provided by an embodiment of the present application.

[0023] Figure 6 It is a flowchart of channel dimension mapping of two-dimensional image features and three-dimensional point cloud features provided in an embodiment of the present application.

[0024] Figure 7 It is a flowchart for generating a semantic region contrast loss value by performing image-text-point cloud contrast learning based on text features, two-dimensional region pixel features, and three-dimensional region point cloud features provided by an embodiment of the present application.

[0025] Figure 8 It is a flowchart for calculating a semantic region contrast loss value by performing contrast learning based on region point features, region pixel features, and text features provided by an embodiment of the present application.

[0026] Figure 9 It is a flowchart for calculating an indication parameter between other object regions based on region pixel features and text features provided by an embodiment of the present application.

[0027] Figure 10 It is a flowchart for calculating a region loss value based on negative samples, region point features corresponding to an object region, and region pixel features provided by an embodiment of the present application.

[0028] Figure 11 It is a structural block diagram of a point cloud network training device provided by another embodiment of the present application.

[0029] Figure 12 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0030] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0031] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the flowchart.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0033] First, several nouns involved in the present application are analyzed: Artificial Intelligence (AI): is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. AI is a branch of computer science. AI attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Research in this field includes robots, language recognition, image recognition, natural language processing and expert systems. AI can simulate the information process of human consciousness and thinking. AI is also a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0034] 3D scene understanding based on LiDAR point clouds is widely used in scenarios such as autonomous driving and robot navigation. It uses algorithms or models to extract features from the input 3D point clouds and semantically understand the 3D scene content based on the extraction results. 3D scene understanding related models based on LiDAR point clouds are usually trained in a supervised learning process based on manually annotated point clouds. However, the annotation cost of annotated point cloud data is high, which limits the scope of application of 3D scene understanding.

[0035] In the related art, the image-point cloud-based self-supervised learning method trains the 3D point cloud network through pixel-point-cloud point-by-point comparison learning. However, this point-level learning method is easily affected by limited receptive fields, limited context information, and potential noise points or erroneous pixel semantics, resulting in low accuracy of 3D feature extraction and affecting the accuracy of 3D scene understanding.

[0036] Based on this, the embodiments of the present application provide a point cloud network training method, a point cloud data processing method and a device, introduce multimodal joint training, and use regional text features to achieve cross-modal alignment. Based on the image-text-point cloud self-supervised contrast learning framework, the representation ability of the point cloud network is enhanced, and local semantic consistency is enhanced during the training process: by shortening the cross-modal feature distance of the same semantic area and pushing away the features of different semantic areas, fine-grained cross-modal feature alignment is achieved, avoiding global feature blurring, and significantly improving the feature extraction accuracy of the three-dimensional point cloud, thereby improving the accuracy of downstream three-dimensional point cloud understanding, without relying on manually annotated point cloud data, to achieve low-cost and high-efficiency three-dimensional space understanding.

[0037] The embodiments of the present application provide a point cloud network training method, a point cloud data processing method and a device, which are specifically illustrated by the following embodiments. First, the point cloud network training method in the embodiments of the present application is described.

[0038] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0039] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0040] The point cloud network training method provided by the embodiments of the present application relates to the field of unmanned driving and intelligent transportation technologies. The point cloud network training method provided by the embodiments of the present application can be applied to terminals, can also be applied to server sides, or can be a computer program running on terminals or server sides. For example, the computer program can be a native program or software module in an operating system; it can be a local (Native) application program (Application, APP), that is, a program that needs to be installed in the operating system to run, such as a client supporting point cloud network training, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module, or plug-in. Among them, the terminal communicates with the server through the network. The point cloud network training method can be executed by the terminal or the server, or jointly executed by the terminal and the server.

[0041] In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, etc. In addition, the terminal may also be an intelligent vehicle-mounted device. The intelligent vehicle-mounted device applies the point cloud network training method of this embodiment to provide relevant services and improve the driving experience. The server may be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; it may also be a service node in a blockchain system, and the service nodes in the blockchain system form a Peer To Peer (P2P) network. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). The terminal and the server can be connected through communication connection methods such as Bluetooth, Universal Serial Bus (USB), or network, and this embodiment does not limit this.

[0042] This application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0043] First, the self-supervised training system provided by the embodiments of this application is described. Refer to Figure 1 , Figure 1 which is a schematic diagram of the training process of the self-supervised training system provided by the embodiments of this application.

[0044] Among them, the self-supervised training system includes an image encoder, a text encoder, a vision foundation model, an image-to-text foundation model, a two-dimensional mapping layer, a three-dimensional mapping layer, and a point cloud network to be trained. Except for the point cloud network and the three-dimensional mapping layer, the model parameters corresponding to other networks are frozen (the frozen parameters are indicated by thick solid frames in the figure), that is to say, only the point cloud network and the three-dimensional mapping layer are fine-tuned during training (the parameters to be trained are indicated by thick dashed frames in the figure), and the parameters of other networks remain unchanged. Doing so can reduce the computational overhead during training and improve the training efficiency of the point cloud network on the one hand, and avoid the phenomenon of non-convergence of the loss function that may occur in joint training on the other hand.

[0045] In one embodiment, the dual encoders in the multi-modal model (Contrastive Language-Image Pre-training, CLIP) are used as the image encoder and the text encoder. Among them, the image encoder is the CLIP image encoder, which adopts the Vision Transformer (ViT) architecture or the ResNet architecture to map the input image data into a feature vector. The text encoder is the CLIP text encoder, which is based on the Transformer architecture and maps the input text data into a feature vector.

[0046] In one embodiment, the vision foundation model is the SAM (Segment Anything Model) model, which can perform mask segmentation on the input image data and output mask information unrelated to the category corresponding to the image data. In addition, the image-to-text foundation model is the BLIP-2 model, which can generate corresponding text descriptions according to the input image data.

[0047] In one embodiment, both the two-dimensional mapping layer and the three-dimensional mapping layer can be MLP network layers, which are used to align the input data in the channel dimension. The two-dimensional mapping layer is a network layer with frozen parameters, and the three-dimensional mapping layer is a network layer to be trained.

[0048] The following describes the point cloud network training method in the embodiments of the present application in combination with the above self-supervised training system.

[0049] Figure 2 is an optional flowchart of the point cloud network training method provided by the embodiments of the present application. Figure 2 The method in may include but is not limited to steps 110 to 140. At the same time, it can be understood that the order of steps 110 to 140 in this embodiment Figure 2 is not specifically limited, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.

[0050] Step 110: Obtain lidar point cloud data and corresponding multi-view images. Input the lidar point cloud data into a point cloud network for feature extraction to obtain 3D point cloud features, and input the multi-view images into an image encoder for feature extraction to obtain 2D image features.

[0051] In one embodiment, referring to Figure 1 , the lidar point cloud data is 3D spatial data obtained by scanning a target scene using a lidar. Specifically, the lidar emits laser pulses into the target scene and receives the reflected signals generated by the target scene. Based on the reflected signals, the distances and azimuths of the objects in the target scene are measured, and lidar point cloud data composed of a set of discrete points is generated. Each point in the lidar point cloud data usually contains information such as 3D coordinates and reflection intensity. The multi-view images are a set of 2D images taken of the target scene from different perspectives. That is to say, the multi-view images include scene images corresponding to multiple perspectives respectively. These scene images can be collected by multiple fixed cameras (such as vehicle surround view systems) or mobile cameras (such as drone aerial photography), which can provide rich visual information, including color, texture, lighting changes, etc.

[0052] Next, referring to Figure 1 , input the lidar point cloud data into the point cloud network for feature extraction to obtain 3D point cloud features , where N and D represent the number and feature dimension of the point cloud features respectively. In addition, input the multi-view images into the image encoder for feature extraction to obtain 2D image features corresponding to each scene image , where h, w, and C represent the size and feature dimension of the image features respectively.

[0053] Step 120: Input the multi-view images into a vision foundation model for mask division to obtain at least one object region, and use a text encoder to generate text features corresponding to each object region.

[0054] In one embodiment, considering that directly performing modal alignment of 2D image features and 3D point cloud features may lack sufficient receptive fields and be unable to perceive the context environment, which limits the accuracy of representation learning. Therefore, in the embodiments of the present application, region-level semantic feature extraction is performed on the multi-view images to guide the alignment process of multi-modal data, enhance the ability of point cloud representation learning, and thus improve the accuracy of downstream 3D point cloud understanding. Referring to Figure 3 , Figure 3 is a flowchart showing the input of multi-view images into a vision foundation model for mask division to obtain at least one object region provided by the embodiments of the present application, which specifically includes the following steps: Step 310: Input the scene images into the visual base model for mask division to obtain the mask images corresponding to each scene image.

[0055] In one embodiment, for each scene image, it is input into the visual base model for mask division to obtain the mask image corresponding to each scene image. Among them, the possible objects in the scene image are located using their corresponding masks. Therefore, the mask image includes mask regions corresponding to each possible object, and the mask region is used to identify the range occupied by the corresponding object in the scene image.

[0056] Step 320: Based on the mask regions, perform region division in each scene image to obtain at least one object region corresponding to each scene image.

[0057] In one embodiment, since the objects may be irregular, the mask regions may be irregular. To better understand the context information, in the embodiments of the present application, each mask region is expanded into a rectangular box. The rectangular box can be the smallest rectangle containing the corresponding object, or a slightly expanded rectangle of the smallest rectangle. This embodiment does not limit this.

[0058] After having the rectangular region corresponding to each mask region, use this rectangular region to intercept in the corresponding scene image, and take the local region of the scene image corresponding to the rectangular region as the object region. It can be understood that in this interception method, each object region contains the corresponding object and a part of the background information of the object.

[0059] In one embodiment, to avoid data redundancy, a minimum mask area can also be set. Only the mask regions larger than this mask area need to generate object regions, so as to prevent the generation of overly localized or meaningless data. In addition, descriptions that are irrelevant or meaningless to the pre-trained dataset category can also be filtered out, such as "cloudy sky", to generate meaningful and descriptive data, thereby enhancing the regional information of each category.

[0060] In one embodiment, refer to Figure 4 , Figure 4 is a schematic diagram of the object region provided by the embodiments of the present application. Figure 4 shows a mask image obtained by a scene image being subjected to mask division by a visual base model. It can be seen from this mask image that the mask image includes multiple mask regions such as the mask region of a car, the mask region of a traffic cone, the mask region of a pedestrian, and the mask region of a fire hydrant. Next, each mask region is expanded to obtain the corresponding rectangular box, and then positioned in the scene image according to this rectangular box to obtain the object region corresponding to each mask region. The object region contains the object and the corresponding background. For example, the object region corresponding to the traffic cone contains the traffic cone and its ground background.

[0061] Multiple object regions corresponding to each scene image are obtained through the above process. It can be understood that due to the difference in perspectives, the shapes and sizes of the same object in different scene images may be different, and the corresponding object regions are also related to the corresponding perspectives.

[0062] Next, for each object region, its corresponding semantic information is determined. Refer to Figure 5 , Figure 5 FIG. is a flowchart of using a text encoder to generate text features corresponding to each object region provided by an embodiment of the present application, which specifically includes the following steps: Step 510: For each scene image, input the corresponding object region into the image-to-text basic model for text generation to obtain a guiding language description corresponding to each object region.

[0063] In one embodiment, for each scene image, all its object regions are analyzed to obtain the semantic information it contains. Refer to Figure 1 , input the object region into the image-to-text basic model for content analysis, and use a text prompt to guide the image-to-text basic model to generate a guiding language description in combination with the object in the object region and its corresponding background. For example, the text prompt can be: "Question: What is the content of the image? Answer: ". Refer to Figure 4 , the guiding language descriptions generated for different object regions can be: "The rear of a yellow car on the road", "A close-up of a truck with large tires", "A large excavator parked in front of a person", "A traffic cone placed on the ground", etc.

[0064] It can be seen that the embodiment of the present application excavates the potential of the text modality through the guiding language description to capture meaningful semantic information, and can generate real and specific region-level text descriptions for each scene image, such as positional relationships, color attributes, etc., providing more fine-grained and comprehensive reference information for the subsequent contrast learning process, thereby enhancing the self-supervised representation learning ability.

[0065] Step 520: Input the guiding language description into the text encoder for feature extraction to obtain the corresponding text features.

[0066] In one embodiment, there is a corresponding guiding language description for each object region in the above process. Refer to Figure 1 , input the guiding language descriptions into the text encoder for feature extraction respectively to obtain the corresponding text features , where B and L respectively represent the number of categories in the pre-training dataset and the feature dimension of the text.

[0067] Next, a multi-modal data alignment process is performed based on text features. Before this, channel dimension mapping also needs to be performed on two-dimensional image features and three-dimensional point cloud features. In one embodiment, referring to Figure 6 , Figure 6 is a flowchart for performing channel dimension mapping on two-dimensional image features and three-dimensional point cloud features provided by an embodiment of the present application, specifically including the following steps: Step 610: Input the two-dimensional image features into a two-dimensional mapping layer for channel dimension mapping to update the two-dimensional image features.

[0068] Step 620: Input the three-dimensional point cloud features into a three-dimensional mapping layer for channel dimension mapping to update the three-dimensional point cloud features.

[0069] In one embodiment, referring to Figure 1 , the two-dimensional image features are updated through the two-dimensional mapping layer to obtain the updated two-dimensional image features , where L represents the updated channel dimension, which is consistent with the feature dimension of the text, and the three-dimensional image features are updated through the three-dimensional mapping layer to obtain the updated three-dimensional point cloud features . It can be seen that , and all have the same channel dimension, enabling the subsequent contrast learning process.

[0070] Step 130: Correlate the two-dimensional image features and the three-dimensional point cloud features according to a preset spatial projection relationship to obtain a point-pixel correlation result, and divide the point-pixel correlation result according to the object region to obtain the two-dimensional region pixel features and three-dimensional region point cloud features corresponding to each object region.

[0071] In one embodiment, since the lidar point cloud data and the multi-view images are both obtained for the target scene, the preset spatial projection relationship is obtained according to the extrinsic data (rotation matrix and translation vector) and intrinsic data (camera focal length, distortion, etc.) of the camera-lidar, and for each scene image, the points in the lidar point cloud data are projected onto the scene image to obtain the corresponding pixels.

[0072] In one embodiment, the points in the lidar point cloud data have corresponding point features in the corresponding three-dimensional point cloud features, and the pixels have pixel features in the corresponding two-dimensional image features. Therefore, according to the above projection relationship, the two-dimensional image features and the three-dimensional point cloud features can be correlated to obtain a point-pixel correlation result, expressed as: , where represents the point feature corresponding to the i-th point, represents the pixel feature corresponding to the i-th point in one of the scene images, and K represents the number of successfully paired points.

[0073] In one embodiment, since the object regions corresponding to each scene image have been divided in the above process, for a scene image, the point-pixel association result can be divided according to its object region, and the pixel features corresponding to the pixels within the object region in the point-pixel association result are used to form two-dimensional region pixel features, and then the corresponding three-dimensional region point cloud features are determined according to the two-dimensional region pixel features.

[0074] According to the above process, multiple sets of two-dimensional region pixel features and corresponding three-dimensional region point cloud features corresponding to each perspective are obtained.

[0075] Step 140: Perform image-text-point cloud contrast learning based on the text features, two-dimensional region pixel features, and three-dimensional region point cloud features to generate a semantic region contrast loss value, and adjust the parameters of the point cloud network based on the semantic region contrast loss value until the trained point cloud network is obtained.

[0076] In one embodiment, referring to Figure 7 , Figure 7 is a flowchart of performing image-text-point cloud contrast learning based on the text features, two-dimensional region pixel features, and three-dimensional region point cloud features provided by an embodiment of the present application to generate a semantic region contrast loss value, which specifically includes the following steps: Step 710: Perform average pooling on the two-dimensional region pixel features to obtain region point features, and perform average pooling on the three-dimensional region point cloud features to obtain region pixel features.

[0077] In one embodiment, different from the point-level contrast learning in the related art, the embodiment of the present application performs a region-level solution to promote information-rich feature learning. Specifically, for each object region, average pooling is performed on the two-dimensional region pixel features to obtain region point features , and average pooling is performed on the three-dimensional region point cloud features to obtain region pixel features . Among them, average pooling can be understood as an operation of calculating the average value. In addition, in the embodiment of the present application, all region point features can be represented as , and all region pixel features can be represented as , where M represents the total number of object regions in all perspectives.

[0078] Step 720: Calculate the semantic region contrast loss value according to the region point features, region pixel features, and text features.

[0079] In one embodiment, during the process of contrastive learning, since different parts of the same object in the target scenario may be marked as different mask regions, it is possible that these different parts with the same semantics of the same object are regarded as negative samples, thus pushing away these false negative samples and making it difficult for them to learn. Therefore, in the embodiment of the present application, when calculating the loss function, it is necessary to use the regional semantic information to filter these false negative samples, effectively enhance the perception of context information, alleviate the self-semantic conflict problem at the regional level, and enforce the consistency of regional representations to improve the accuracy of contrastive learning. Refer to Figure 8 , Figure 8 FIG. Figure 8 is a flowchart for calculating the semantic region contrast loss value through contrastive learning based on regional point features, regional pixel features, and text features provided by an embodiment of the present application, and specifically includes the following steps: Step 810: For each object region, calculate the indication parameter between it and other object regions according to the regional pixel feature and the text feature, and determine multiple negative samples from other object regions based on the indication parameter.

[0080] In one embodiment, it is necessary to determine the corresponding negative sample for each object region. During the process of determining the negative sample, it is necessary to use the indication parameter to judge whether other object regions outside the object region can be used as its negative sample. Refer to Figure 9 , Figure 9 FIG. Figure 9 is a flowchart for calculating the indication parameter between an object region and other object regions according to the regional pixel feature and the text feature provided by an embodiment of the present application, and specifically includes the following steps: Step 910: Select a comparison region one by one from other object regions, and obtain the regional pixel feature of the comparison region as the comparison pixel feature.

[0081] In one embodiment, assume that it is currently necessary to determine the negative sample of object region i. Select a comparison region one by one from other object regions, assume it is object region j. At this time, the text feature of object region i is , the regional pixel feature is , and the regional pixel feature of object region j is .

[0082] Step 920: Calculate the first text similarity between the text feature and the regional pixel feature, and calculate the second text similarity between the text feature and the comparison pixel feature.

[0083] In one embodiment, the first text similarity is the cosine similarity between the text feature and the regional pixel feature, denoted as , and the second text similarity is the cosine similarity between the text feature and the comparison pixel feature, denoted as , Step 930: Obtain the difference between the first text similarity and the second text similarity. If the difference is greater than or equal to a preset reference value, set the indication parameter to one.

[0084] In one embodiment, the difference is expressed as: , and the preset reference value is . If , then the indication parameter is 0, indicating that the object region j is not a negative sample of the object region i. Otherwise, the indication parameter is 1, indicating that the object region j is a negative sample of the object region i.

[0085] In this way, one or more negative samples corresponding to each object region are obtained.

[0086] Step 820: Calculate the region loss value based on the negative samples, the region point features and the region pixel features corresponding to the object region, and obtain the semantic region contrast loss value according to the mean value of the region loss value.

[0087] In one embodiment, referring to Figure 10 , Figure 10 is a flowchart for calculating the region loss value based on the negative samples, the region point features and the region pixel features corresponding to the object region provided by the embodiment of the present application, which specifically includes the following steps: Step 1010: For each object region, calculate the first semantic similarity between the region point feature and the region pixel feature.

[0088] In one embodiment, taking the object region i as an example, its region point feature is , and the region pixel feature is . Therefore, the first semantic similarity is expressed as: , where is a preset temperature parameter used to control the sharpness of the distribution of the first semantic similarity, which can be set according to the actual situation.

[0089] Step 1020: Obtain the region pixel feature of each negative sample as the sample pixel feature, calculate the sample semantic similarity between the region point feature and each sample pixel feature, and calculate the sum of the sample semantic similarities as the second semantic similarity.

[0090] In one embodiment, for the negative sample, its region pixel feature is called the sample pixel feature. Therefore, for the object region i, when the negative sample is the object region j, the sample semantic similarity is expressed as: . Here, an indication parameter is introduced. If it is 0, the calculated value of the sample semantic similarity is 0. Next, sum all the sample semantic similarities of the object region i, and the obtained second semantic similarity is expressed as: .

[0091] Step 1030: Obtain the sum of the first semantic similarity and the second semantic similarity as an intermediate reference value, and obtain the region loss value according to the quotient of the first semantic similarity and the intermediate reference value.

[0092] In one embodiment, the area loss value is expressed as:

[0093] Next, the semantic region contrast loss value is obtained based on the average of the region loss values , expressed as:

[0094] After obtaining the semantic region contrast loss value, the semantic region contrast loss value is minimized, and the parameters of the point cloud network and the three-dimensional mapping layer are adjusted until the trained point cloud network and the three-dimensional mapping layer are obtained.

[0095] The technical solution provided by the embodiment of the present application obtains laser radar point cloud data and corresponding multi-view images, inputs the laser radar point cloud data into the point cloud network for feature extraction, obtains three-dimensional point cloud features, and inputs the multi-view images into the image encoder for feature extraction to obtain two-dimensional image features, inputs the multi-view images into the visual basic model for mask division to obtain at least one object area, and uses the text encoder to generate text features corresponding to each object area, associates the two-dimensional image features with the three-dimensional point cloud features according to the preset spatial projection relationship to obtain point-pixel association results, divides the point-pixel association results according to the object area, obtains the two-dimensional region pixel features and three-dimensional region point cloud features corresponding to each object area, performs image-text-point cloud comparison learning according to the text features, two-dimensional region pixel features and three-dimensional region point cloud features, generates semantic region comparison loss values, and adjusts the parameters of the point cloud network based on the semantic region comparison loss values ​​until the trained point cloud network is obtained. The embodiment of the present application introduces multimodal joint training and uses region-level text features to achieve cross-modal alignment. Based on the self-supervised contrastive learning framework of image-text-point cloud, the representation ability of the point cloud network is enhanced, and the local semantic consistency is enhanced during the training process: by shortening the cross-modal feature distance of the same semantic area and pushing the features of different semantic areas apart, fine-grained cross-modal feature alignment is achieved, avoiding global feature blurring, and significantly improving the feature extraction accuracy of 3D point clouds, thereby improving the accuracy of downstream 3D point cloud understanding. There is no need to rely on manually annotated point cloud data, and low-cost and high-efficiency 3D space understanding is achieved.

[0096] In one embodiment, a three-dimensional point cloud data processing method is further provided. The specific steps include: obtaining initial point cloud data, inputting the initial point cloud data into a point cloud network for feature extraction to obtain point cloud feature data, where the point cloud network is trained by the point cloud network training method provided in any of the above embodiments. The point cloud feature data can be used for downstream three-dimensional space processes according to actual requirements.

[0097] An embodiment of the present application further provides a point cloud network training device, which can implement the above point cloud network training method. Referring to Figure 11 , the device includes: Initial feature extraction module 1110: used to obtain lidar point cloud data and corresponding multi-view images, input the lidar point cloud data into a point cloud network for feature extraction to obtain three-dimensional point cloud features, and input the multi-view images into an image encoder for feature extraction to obtain two-dimensional image features.

[0098] Text description generation module 1120: used to input the multi-view images into a vision foundation model for mask division to obtain at least one object region, and use a text encoder to generate text features corresponding to each object region.

[0099] Mapping association module 1130: used to associate the two-dimensional image features and three-dimensional point cloud features according to a preset spatial projection relationship to obtain a point-pixel association result, and divide the point-pixel association result according to the object region to obtain two-dimensional region pixel features and three-dimensional region point cloud features corresponding to each object region.

[0100] Contrast learning module 1140: used to perform image-text-point cloud contrast learning according to the text features, two-dimensional region pixel features, and three-dimensional region point cloud features to generate a semantic region contrast loss value, and adjust the parameters of the point cloud network based on the semantic region contrast loss value until a trained point cloud network is obtained.

[0101] The specific implementation manner of the point cloud network training device in this embodiment is basically the same as that of the above point cloud network training method, and will not be elaborated here.

[0102] An embodiment of the present application further provides an electronic device, including: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes the at least one program to implement the above-mentioned point cloud network training method or three-dimensional point cloud data processing method of the present application. The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.

[0103] Please refer to Figure 12 , Figure 12 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes: A processor 1201, which can be implemented in the form of a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application; A memory 1202, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1202 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1202 and are called by the processor 1201 to execute the point cloud network training method or the three-dimensional point cloud data processing method of the embodiments of the present application; An input / output interface 1203, which is used to implement information input and output; A communication interface 1204, which is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.); A bus 1205, which transmits information between various components of the device (such as the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204); Among them, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 are communicatively connected to each other inside the device through the bus 1205.

[0104] The embodiments of the present application also provide a storage medium, which is a storage medium that stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned point cloud network training method or the three-dimensional point cloud data processing method.

[0105] As a non-transitory storage medium, the memory can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0106] The point cloud network training method, point cloud data processing method and device proposed in the embodiment of the present application obtain laser radar point cloud data and corresponding multi-view images, input the laser radar point cloud data into the point cloud network for feature extraction, obtain three-dimensional point cloud features, and input the multi-view images into the image encoder for feature extraction to obtain two-dimensional image features, input the multi-view images into the visual basic model for mask division to obtain at least one object area, and use the text encoder to generate text features corresponding to each object area, associate the two-dimensional image features with the three-dimensional point cloud features according to the preset spatial projection relationship to obtain point-pixel association results, divide the point-pixel association results according to the object area, obtain the two-dimensional region pixel features and three-dimensional region point cloud features corresponding to each object area, perform image-text-point cloud comparison learning according to the text features, two-dimensional region pixel features and three-dimensional region point cloud features, generate semantic region comparison loss values, and adjust the parameters of the point cloud network based on the semantic region comparison loss values ​​until the trained point cloud network is obtained. The embodiment of the present application introduces multimodal joint training and uses region-level text features to achieve cross-modal alignment. Based on the self-supervised contrastive learning framework of image-text-point cloud, the representation ability of the point cloud network is enhanced, and the local semantic consistency is enhanced during the training process: by shortening the cross-modal feature distance of the same semantic area and pushing the features of different semantic areas apart, fine-grained cross-modal feature alignment is achieved, avoiding global feature blurring, and significantly improving the feature extraction accuracy of 3D point clouds, thereby improving the accuracy of downstream 3D point cloud understanding. There is no need to rely on manually annotated point cloud data, and low-cost and high-efficiency 3D space understanding is achieved.

[0107] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0108] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0110] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0111] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0112] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0113] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0114] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0116] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0117] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A method for training a point cloud network, characterized in that Performed by a self-supervised training system, the self-supervised training system at least includes an image encoder, a text encoder, a vision foundation model, and the point cloud network. Except for the point cloud network, the parameters of other networks are frozen. The method includes: Obtain lidar point cloud data and corresponding multi-view images, input the lidar point cloud data into the point cloud network for feature extraction to obtain three-dimensional point cloud features, and input the multi-view images into the image encoder for feature extraction to obtain two-dimensional image features; Input the multi-view images into the vision foundation model for mask division to obtain at least one object region, and use the text encoder to generate text features corresponding to each object region; Associate the two-dimensional image features and the three-dimensional point cloud features according to a preset spatial projection relationship to obtain a point-pixel association result, and divide the point-pixel association result according to the object region to obtain two-dimensional region pixel features and three-dimensional region point cloud features corresponding to each object region; Perform image-text-point cloud contrast learning according to the text features, the two-dimensional region pixel features, and the three-dimensional region point cloud features to generate a semantic region contrast loss value, and adjust the parameters of the point cloud network based on the semantic region contrast loss value until the trained point cloud network is obtained.

2. The point cloud network training method according to claim 1, wherein The multi-view images include scene images corresponding to at least one view. The step of inputting the multi-view images into the vision foundation model for mask division to obtain at least one object region includes: Input the scene images into the vision foundation model for mask division respectively to obtain mask images corresponding to each scene image, and each mask image includes at least one mask region; Based on the mask regions, perform region division in each scene image to obtain at least one object region corresponding to each scene image.

3. The point cloud network training method according to claim 2, characterized in that, The self-supervised training system further includes an image-to-text foundation model. The step of using the text encoder to generate text features corresponding to each object region includes: For each scene image, input the corresponding object region into the image-to-text foundation model for text generation to obtain a guiding language description corresponding to each object region; Input the guiding language description into the text encoder for feature extraction to obtain the corresponding text features.

4. The point cloud network training method according to claim 1, wherein The step of performing image-text-point cloud contrast learning according to the text features, the two-dimensional region pixel features, and the three-dimensional region point cloud features to generate a semantic region contrast loss value includes: Perform average pooling on the two-dimensional region pixel features to obtain region point features, and perform average pooling on the three-dimensional region point cloud features to obtain region pixel features; Calculate the semantic region contrast loss value according to the region point features, the region pixel features, and the text features.

5. The point cloud network training method according to claim 4, wherein The step of performing contrast learning according to the region point features, the region pixel features, and the text features to calculate the semantic region contrast loss value includes: For each of the said object regions, an indication parameter between it and other said object regions is calculated according to the region pixel features and the text features, and a plurality of negative samples are determined from other said object regions based on the indication parameter; Based on the negative samples, the region point features corresponding to the object region, and the region pixel features, a region loss value is calculated, and the semantic region contrast loss value is obtained according to the mean value of the region loss value.

6. The point cloud network training method according to claim 5, wherein, The calculating the indication parameter between it and other said object regions according to the region pixel features and the text features includes: One by one, a comparison region is selected from other said object regions, and the region pixel features of the comparison region are obtained as comparison pixel features; Calculate a first text similarity between the text features and the region pixel features, and calculate a second text similarity between the text features and the comparison pixel features; Obtain the difference between the first text similarity and the second text similarity. If the difference is greater than or equal to a preset reference value, set the indication parameter to one, which is used to indicate that the comparison region is a negative sample of the object region.

7. The point cloud network training method according to claim 5, wherein The calculating the region loss value based on the negative samples, the region point features corresponding to the object region, and the region pixel features includes: For each of the said object regions, calculate a first semantic similarity between the region point features and the region pixel features; Obtain the region pixel features of each negative sample as sample pixel features, calculate a sample semantic similarity between the region point features and each sample pixel feature, and calculate the sum of the sample semantic similarities as a second semantic similarity; Obtain the sum of the first semantic similarity and the second semantic similarity as an intermediate reference value, and obtain the region loss value according to the quotient of the first semantic similarity and the intermediate reference value.

8. The point cloud network training method according to any one of claims 1 to 7, characterized in that The self-supervised training system further includes a two-dimensional mapping layer and a three-dimensional mapping layer to be trained. Before dividing the two-dimensional image features according to the object regions to obtain two-dimensional region pixel features corresponding to each object region, the method further includes: Input the two-dimensional image features into the two-dimensional mapping layer for channel dimension mapping to update the two-dimensional image features; Input the three-dimensional point cloud features into the three-dimensional mapping layer for channel dimension mapping to update the three-dimensional point cloud features, and the updated three-dimensional point cloud features and the two-dimensional image features have the same channel dimension.

9. A three-dimensional point cloud data processing method, characterized in that including: Obtain initial point cloud data; Input the initial point cloud data into a point cloud network for feature extraction to obtain point cloud feature data, and the point cloud network is trained by the point cloud network training method according to any one of claims 1 to 8.

10. A point cloud network training device, characterized in that, Executed by a self-supervised training system, the self-supervised training system at least includes an image encoder, a text encoder, a visual foundation model, and the point cloud network. Except for the point cloud network, other network parameters are frozen. The device includes: Initial feature extraction module: used to obtain lidar point cloud data and corresponding multi-view images, input the lidar point cloud data into the point cloud network for feature extraction to obtain three-dimensional point cloud features, and input the multi-view images into the image encoder for feature extraction to obtain two-dimensional image features; Text description generation module: used to input the multi-view images into the visual foundation model for mask division to obtain at least one object region, and use the text encoder to generate text features corresponding to each object region; Mapping association module: used to associate the two-dimensional image features and the three-dimensional point cloud features according to a preset spatial projection relationship to obtain a point-pixel association result, and divide the point-pixel association result according to the object region to obtain two-dimensional region pixel features and three-dimensional region point cloud features corresponding to each object region; Contrastive learning module: used to perform image-text-point cloud contrastive learning according to the text features, the two-dimensional region pixel features and the three-dimensional region point cloud features, generate a semantic region contrast loss value, and adjust the parameters of the point cloud network based on the semantic region contrast loss value until the trained point cloud network is obtained.

11. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the point cloud network training method according to any one of claims 1 to 8, or the three-dimensional point cloud data processing method according to claim 9.

12. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the point cloud network training method according to any one of claims 1 to 8, or the three-dimensional point cloud data processing method according to claim 9.

Citation Information

Patent Citations

  • Method for rebuilding point-cloud type three-dimensional surface of nonparallel outline medical image

    CN101625767A