Self-supervised contrast learning method based on BEV network
By adopting a self-supervised comparison learning method based on BEV network in BEV perception, and using the mapping relationship between text labels and BEV grids, the problem that existing methods fail to make full use of multimodal information is solved, and a more accurate 3D object detection effect is achieved.
Patent Information
- Application Number
- CN202510128765.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-10
AI Technical Summary
The existing self-supervised BEV perception method fails to make full use of multimodal information, especially color and texture information in the image, when dealing with complex tasks such as 3D object detection, resulting in poor results when distinguishing objects with similar appearance but different categories.
Using a self-supervised contrast learning method based on BEV network, multimodal contrast pre-training is achieved through pre-generated text labels, the mapping relationship between text and BEV grid is constructed, and a voting strategy based on BEV grid is designed to filter noise and generate high-quality text labels.
Through multimodal comparison pre-training, the semantic understanding ability of the model is significantly enriched, the performance in 3D object detection tasks is improved, and it can more accurately distinguish objects with similar appearance but different categories.
Smart Images

Figure CN120124704A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a self-supervised contrast learning method based on a BEV network. Background Art
[0002] Perception is the cornerstone of an autonomous driving system, and the well-known Bird's Eye View (BEV) provides a unified expression with its natural and intuitive perspective, cleverly avoiding the common occlusion and scale problems in 2D tasks. Past research has deeply explored supervised BEV perception methods, covering LiDAR, cameras, and multi-modal fusion technologies. However, these supervised methods often rely on large labeled datasets for training, which is not only costly but also may be affected by the inherent biases in the annotation process. Therefore, self-supervised BEV perception technology has become an attractive research field, which automatically extracts supervision signals from the data itself through self-supervised methods such as contrast learning, aiming to reduce the dependence on labeled data.
[0003] Previous self-supervised BEV perception methods focused on exploring the potential of a single data modality. BEVContrast focuses on achieving contrast learning by utilizing the point-by-point spatial position information between two BEVs, while BEV-MAE simultaneously utilizes the point cloud spatial position and density information to further learn through masking and reconstruction strategies. Although these methods have achieved certain results in utilizing spatial position information, they have not fully exploited their potential in dealing with complex tasks such as 3D object detection. These tasks not only require spatial position information but also urgently need rich semantic information such as colors and textures in images to more accurately distinguish objects with similar appearances but different categories, so as to achieve a comprehensive understanding of the environment. Summary of the Invention
[0004] The present invention provides a self-supervised contrast learning method based on a BEV network, which realizes multi-modal contrast pre-training by utilizing pre-generated text labels, greatly enriching the semantic understanding ability of the model.
[0005] The present invention provides a self-supervised contrast learning method based on a BEV network, based on a pre-training - fine-tuning architecture, and the method includes:
[0006] Constructing a self-supervised contrast learning model framework, which includes the construction of the mapping relationship between text and BEV, the voting strategy based on BEV grids, and the design of the loss function;
[0007] In the pre-training stage, training the encoder of the self-supervised contrast learning model framework to enhance its feature extraction ability;
[0008] In the fine-tuning stage, a detection head is added to the pre-trained encoder, and the entire self-supervised contrastive learning model is trained end-to-end.
[0009] Furthermore, in the self-supervised contrastive learning model framework, the construction of the text-BEV mapping relationship specifically includes:
[0010] Following the method of CLIP2Scene, the pre-trained CLIP segmentation model MaskCLIP is used to generate dense text-pixel pairs
[0011] The projection matrix between the image and the point cloud is used to generate dense pixel-point pairs
[0012] The point cloud is projected onto the BEV plane to obtain point-grid pairs
[0013] The mapping between the text and the BEV is obtained where t i is the text label, and b i is the coordinate of the BEV grid. The calculation time of the mapping consists of the time for the CLIP segmentation model to infer multiple images, and the time for projecting the points onto multiple images and a BEV respectively; for each frame of the scene, assuming there are k images, p points in the point cloud, and the inference time of the CLIP segmentation model for each image is T, the final time complexity is O(k(T + αp)), where α is a constant.
[0014] Furthermore, in the BEV plane, the equation for defining the coordinates of the projected points on the BEV plane is as follows:
[0015]
[0016] where, respectively represent the x and y coordinates of the i-th point, X low , Y low respectively represent the lower bounds of the x and y axes of the BEV plane, and G x , G y respectively represent the grid sizes in the x and y directions of the BEV plane;
[0017] During the projection process, by setting to filter out the redundant ground points and ensure that the background points and foreground points are roughly balanced.
[0018] Furthermore, in the self-supervised contrastive learning model framework, the voting strategy based on the BEV grid specifically includes:
[0019] Adopt a voting strategy to filter text labels associated with BEV grids: Only retain text labels that appear ≥k times and have a confidence of ≥α as text labels associated with BEV grids; the equation representing this strategy is as follows:
[0020]
[0021] where, represents the text label associated with the j-th BEV grid, T c represents the feature of the c-th text label, represents the number of points of the c-th category projected onto the j-th BEV grid, represents the total number of points projected onto the j-th BEV grid; through the voting strategy, a mapping available for pre-training can finally be obtained
[0022] Furthermore, the calculation time of voting consists of the time for traversing and counting points, and the time for traversing and judging the counting results. For each frame of the scene, there are approximately p paired points and a total of c categories of text labels. The final time complexity is O((c + 1)p);
[0023] Suppose there are m frames of the scene. The total time complexity of projection and voting is O(m(kT+(αk + β(c + 1)p))), where β is a constant. Since the number of images k and the number of text label categories c are both fixed values, therefore, the total time complexity can be simplified to O(m(T + γp)), where γ is a constant, indicating that the calculation time of projection and voting is jointly determined by the scale of the dataset (the number of scene frames m), the inference time T of the CLIP model, and the number of points p within each frame.
[0024] Furthermore, in the self-supervised contrast learning model framework, the design of the loss function specifically includes:
[0025] Combine two loss functions for end-to-end training, and its loss function is:
[0026] L total = αL TB +(1 – α)L BB
[0027] where, α is the balance weight of the loss, L TB is the first loss function, and L BB is the second loss function.
[0028] Furthermore, the first loss function uses text semantic information to select positive and negative samples for contrast learning. The formula of the loss function is as follows:
[0029]
[0030] Among them, C is the number of different text categories, and t i = T c represents t i as the text feature of the c-th category, and r i represents the feature at position b i on the BEV, and τ is the temperature term.
[0031] The second loss function is used for contrastive learning between BEVs, and the loss function formula is:
[0032]
[0033] Among them, M is the number of sample pairs, and r i represents the feature at position b i on the BEV, r' i represents the feature registered to position b i through bilinear interpolation, and τ is the temperature term.
[0034] The present invention also provides a self-supervised contrastive learning device based on a BEV network, based on a pre-training - fine-tuning architecture, and the device includes:
[0035] A construction module for constructing a self-supervised contrastive learning model framework, and the self-supervised contrastive learning model framework includes the construction of the text and BEV mapping relationship, the voting strategy based on BEV grids, and the design of the loss function;
[0036] A pre-training module for training the encoder of the self-supervised contrastive learning model framework during the pre-training stage to enhance its feature extraction ability;
[0037] A fine-tuning module for adding a detection head to the pre-trained encoder and performing end-to-end training on the entire self-supervised contrastive learning model during the fine-tuning stage.
[0038] The present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0039] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the steps of the above method when executed by a processor.
[0040] The beneficial effects of the present invention are:
[0041] The present invention adopts a pre-training and fine-tuning architecture. In the pre-training stage, efforts are made to train the encoder to enhance its feature extraction ability; in the fine-tuning stage, a detection head is integrated into the encoder for end-to-end training to enable it to better adapt to downstream 3D object detection tasks. The present invention proposes a self-supervised contrastive learning framework for BEV perception to address the defects of ignoring multi-modal information and lacking semantics; by constructing a mapping relationship between text labels and BEV grids, positive and negative sample pairs are generated for text-BEV contrastive learning; a voting strategy based on BEV grids is designed to effectively filter out the noise introduced by CLIP and calibration errors in the mapping. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.
[0043] Figure 2 It is a schematic structural diagram of the device according to an embodiment of the present invention.
[0044] Figure 3 It is a schematic internal structure diagram of a computer device according to an embodiment of the present invention.
[0045] The implementation, functional features, and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] The present invention proposes a self-supervised contrastive learning method for BEV perception. Based on the existing BEV contrastive learning framework BEVContrast, text labels are introduced as the supervision signal in the pre-training stage for multi-modal contrastive pre-training. Specifically, first, the zero-shot segmentation ability of the CLIP segmentation model is used to segment multi-view images to obtain pixel-level text labels of the images. Then, in combination with the projection relationship among the image, point cloud, and BEV, a mapping relationship between the text labels and BEV grids is constructed to generate positive and negative sample pairs for contrastive learning. Aiming at the possible noise problems introduced by the CLIP model and calibration errors during the mapping establishment process, a voting mechanism based on BEV grids is further proposed to filter out the noise and ensure high-quality text labels are obtained. Finally, through the use of these pre-generated text labels, multi-modal contrastive pre-training is achieved, greatly enriching the semantic understanding ability of the model.
[0048] As Figure 1 shown, the present invention provides a self-supervised contrastive learning method based on a BEV network. Based on a pre-training and fine-tuning architecture, the method includes:
[0049] S1. Construct a self-supervised contrastive learning model framework, which includes the construction of the text-BEV mapping relationship, the voting strategy based on BEV grids, and the design of the loss function;
[0050] S2. In the pre-training stage, train the encoder of the self-supervised contrastive learning model framework to enhance its feature extraction ability;
[0051] S3. In the fine-tuning stage, add a detection head to the pre-trained encoder and perform end-to-end training on the entire self-supervised contrastive learning model.
[0052] Based on the single-modal contrastive learning method BEVContrast using lidar, a self-supervised contrastive learning method is proposed. Its model framework consists of three components: the construction of the text-BEV mapping relationship, the voting strategy based on BEV grids, and the design of the loss function. For the fine-tuning stage, a detection head is added to the pre-trained encoder, and the entire model is trained end-to-end. The detection head and loss function used are consistent with SECOND.
[0053] Construction of the text-BEV mapping relationship:
[0054] Constructing the mapping between text and BEV is divided into three steps: (1) Use the CLIP segmentation model to generate pixel-level text labels for the image; (2) Use the projection matrix between the image and the point cloud to transfer the text labels to the point cloud; (3) Project the point cloud onto the BEV to construct the mapping between text and BEV.
[0055] Following the method of CLIP2Scene, use the pre-trained CLIP segmentation model MaskCLIP to generate dense text-pixel pairs Use the projection matrix between the image and the point cloud to generate dense pixel-point pairs Project the point cloud onto the BEV plane to obtain point-grid pairs
[0056] Through the above three steps, the mapping between text and BEV is obtained where t i is the text label, and b i is the coordinate of the BEV grid. The calculation time of the mapping consists of the time for the CLIP segmentation model to infer multiple images and the time to project the points onto multiple images and one BEV respectively; for each frame of the scene, assuming there are k images, there are p points in the point cloud, and the inference time of the CLIP segmentation model for each image is T, the final time complexity is O(k(T + αp)), where α is a constant.
[0057] In the BEV plane, define the coordinates of the projected points on the BEV plane The equation is as follows:
[0058]
[0059] Wherein, respectively represent the x and y coordinates of the i-th point, X low , Y low respectively represent the lower bounds of the x and y axes of the BEV plane, G x , G y respectively represent the grid sizes in the x and y directions of the BEV plane;
[0060] Since most points are background points, if all of them are projected onto the BEV plane, it will lead to serious class imbalance, thus reducing the effect of pre-training. Therefore, during the projection process, a simple truncation method is adopted, that is, during the projection process, by setting to filter out redundant ground points and ensure that the background points and foreground points are roughly balanced.
[0061] Voting strategy based on BEV grid:
[0062] There are the following problems with directly using the obtained mapping for pre-training: (1) Since the same point may be associated with pixels on different images, and the segmentation results of MaskCLIP usually contain significant noise, the same point may be associated with different text labels; (2) Points projected onto the same BEV grid may have different text labels.
[0063] Therefore, a voting strategy is adopted to filter the text labels associated with the BEV grid: only retain the text labels that appear ≥k times and have a confidence of ≥α as the text labels associated with the BEV grid; the equation representing this strategy is as follows:
[0064]
[0065] Wherein, represents the text label associated with the j-th BEV grid, T c represents the feature of the c-th text label, represents the number of points of the c-th category projected onto the j-th BEV grid, represents the total number of points projected onto the j-th BEV grid; through the voting strategy, a mapping that can be used for pre-training can finally be obtained
[0066] The calculation time of voting consists of the time for traversing and counting points and the time for traversing and judging the counting results. For each frame of the scene, there are about p paired points and a total of c categories of text labels, and the final time complexity is O((c + 1)p);
[0067] Suppose there are m frames of scenarios in total, and the total time complexity of projection and voting is O(m(kT+(αk+β(c+1)p))), where β is a constant. Since both the number of images k and the number of text label categories c are fixed values, the total time complexity can be simplified to O(m(T+γp)), where γ is a constant, indicating that the calculation time of projection and voting is jointly determined by the scale of the dataset (the number of scenario frames m), the inference time T of the CLIP model, and the number of points p within each frame.
[0068] Design of the loss function:
[0069] The projection and voting steps are crucial for establishing the mapping relationship between the BEV grid and text labels. These steps are only implemented in the data preprocessing stage, aiming to batch generate all the text labels required in the pre-training stage (i.e., the label encodings on the BEV). By restricting the projection and voting to the preprocessing stage, the repetition of these steps in the subsequent pre-training process is effectively avoided, thus significantly improving the efficiency of pre-training. After successfully establishing the mapping between the BEV grid and text labels, two loss functions are combined for end-to-end training, and the loss function is:
[0070] L total =αL TB +(1 - α)L BB
[0071] where α is the balancing weight of the loss, L TB is the first loss function, and L BB is the second loss function.
[0072] The first loss function uses text semantic information to select positive and negative samples for contrastive learning. The formula of the loss function is as follows:
[0073]
[0074] where C is the number of different text categories, t i =T c means that t i is the text feature of the c-th category, r i represents the feature at position b i on the BEV, and τ is the temperature term.
[0075] The second loss function focuses on the framework of text-BEV contrastive pre-training and the establishment of the text-BEV mapping. For the contrastive learning between BEVs, the formula of the loss function is:
[0076]
[0077] where M is the number of sample pairs, and r i represents the feature at position b iThe feature at r′ i represents the feature registered to position b through bilinear interpolation i , and τ is the temperature term.
[0078] The present invention adopts a pre-training - fine-tuning architecture. In the pre-training stage, efforts are made to train the encoder to enhance its feature extraction ability; in the fine-tuning stage, a detection head is integrated into the encoder for end-to-end training to enable it to better adapt to downstream 3D object detection tasks. The present invention proposes a self-supervised contrast learning framework for BEV perception to address the defects of ignoring multi-modal information and lacking semantics; by constructing a mapping relationship between text labels and BEV grids, positive and negative sample pairs are generated for text-BEV contrast learning; a voting strategy based on BEV grids is designed to effectively filter out the noise introduced by CLIP and calibration errors in the mapping.
[0079] As Figure 2 shown, the present invention also provides a self-supervised contrast learning device based on a BEV network. Based on the pre-training - fine-tuning architecture, the device includes:
[0080] A construction module 1 for constructing a self-supervised contrast learning model framework, which includes the construction of the text-BEV mapping relationship, a voting strategy based on BEV grids, and the design of a loss function;
[0081] A pre-training module 2 for training the encoder of the self-supervised contrast learning model framework in the pre-training stage to enhance its feature extraction ability;
[0082] A fine-tuning module 3 for adding a detection head to the pre-trained encoder and performing end-to-end training on the entire self-supervised contrast learning model in the fine-tuning stage.
[0083] In one embodiment, in the construction module 1, the construction of the text-BEV mapping relationship specifically includes:
[0084] Following the method of CLIP2Scene, using the pre-trained CLIP segmentation model MaskCLIP to generate dense text pixel pairs
[0085] Using the projection matrix between the image and the point cloud to generate dense pixel-point pairs
[0086] Projecting the point cloud onto the BEV plane to obtain point-grid pairs
[0087] Obtaining the mapping between text and BEV where t i is the text label, and b iThey are the coordinates of the BEV grid. The calculation time of the mapping consists of the time for the CLIP segmentation model to perform inference on multiple images and the time for projecting points onto multiple images and a BEV respectively. For each frame of the scene, assuming there are k images and p points in the point cloud, and the inference time of the CLIP segmentation model for each image is T, the final time complexity is O(k(T + αp)), where α is a constant.
[0088] In one embodiment, in the BEV plane, the coordinates of the projected points on the BEV plane are defined The equation is as follows:
[0089]
[0090] Where, respectively represent the x and y coordinates of the i-th point, X low , Y low respectively represent the lower bounds of the x and y axes of the BEV plane, G x , G y respectively represent the grid sizes in the x and y directions of the BEV plane;
[0091] During the projection process, by setting to filter out redundant ground points and ensure that the background points and foreground points are roughly balanced.
[0092] In one embodiment, in building block 1, the voting strategy based on the BEV grid specifically includes:
[0093] Adopt a voting strategy to filter text labels associated with the BEV grid: Only retain text labels that appear ≥k times and have a confidence level of ≥α as text labels associated with the BEV grid. The equation representing this strategy is as follows:
[0094]
[0095] Where, represents the text label associated with the j-th BEV grid, T c represents the feature of the c-th text label, represents the number of points of the c-th category projected onto the j-th BEV grid, represents the total number of points projected onto the j-th BEV grid; Through the voting strategy, a mapping available for pre-training can finally be obtained
[0096] In one embodiment, the calculation time of the voting consists of the time for traversing and counting points and the time for traversing and judging the counting results. For each frame of the scene, there are approximately p paired points and a total of c categories of text labels, and the final time complexity is O((c + 1)p);
[0097] Suppose there are m frames of scenes in total, and the total time complexity of projection and voting is O(m(kT+(αk+β(c + 1)p))), where β is a constant. Since both the number of images k and the number of text label categories c are fixed values, the total time complexity can be simplified to O(m(T+γp)), where γ is a constant, indicating that the calculation time of projection and voting is jointly determined by the scale of the dataset (the number of scene frames m), the inference time T of the CLIP model, and the number of points p within each frame.
[0098] In one embodiment, in building block 1, the design of the loss function specifically includes:
[0099] Combining two loss functions for end-to-end training, and the loss function is:
[0100] L total =αL TB +(1 - α)L BB
[0101] where α is the balancing weight of the loss, L TB is the first loss function, and L BB is the second loss function.
[0102] In one embodiment, the first loss function uses text semantic information to select positive and negative samples for contrast learning, and the loss function formula is as follows:
[0103]
[0104] where C is the number of different text categories, t i =T c means that t i is the text feature of the c-th category, r i represents the feature at position b i on the BEV, and τ is the temperature term.
[0105] The second loss function is used for contrast learning between BEVs, and the loss function formula is:
[0106]
[0107] where M is the number of sample pairs, r i represents the feature at position b i on the BEV, r′ i represents the feature registered to position b i through bilinear interpolation, and τ is the temperature term.
[0108] The above-mentioned modules are used to respectively execute the steps in the above-mentioned self-supervised contrast learning method based on the BEV network, and their specific implementation manners refer to the method embodiments described above, and will not be elaborated here.
[0109] As shown Figure 3 in the figure, the present invention also provides a computer device, which may be a server, and its internal structure may be as shown Figure 3 in the figure. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all the data required for the process of the self-supervised contrast learning method based on the BEV network. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the self-supervised contrast learning method based on the BEV network.
[0110] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.
[0111] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above self-supervised contrast learning methods based on the BEV network.
[0112] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0113] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, article, or method including that element.
[0114] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A self-supervised contrastive learning method based on BEV network, characterized in that: Based on the pre-training-fine-tuning architecture, the method includes: Constructing a self-supervised contrastive learning model framework, wherein the self-supervised contrastive learning model framework includes the construction of a mapping relationship between text and BEV, a voting strategy based on a BEV grid, and a loss function design; In the pre-training stage, the encoder of the self-supervised contrastive learning model framework is trained to enhance its feature extraction capability; In the fine-tuning stage, a detection head is added to the pre-trained encoder, and the entire self-supervised contrastive learning model is trained end-to-end.
2. The self-supervised contrastive learning method based on BEV network according to claim 1 is characterized in that: In the self-supervised contrastive learning model framework, the construction of the mapping relationship between text and BEV specifically includes: Following the CLIP2Scene approach, the pre-trained CLIP segmentation model MaskCLIP is used to generate dense text pixel pairs. Generate dense pixel-point pairs using the projection matrix between the image and the point cloud Project the point cloud onto the BEV plane to obtain point-mesh pairs Get the mapping between text and BEV where t i is the text label, b i are the coordinates of the BEV grid. The computational time of the mapping consists of the time it takes the CLIP segmentation model to infer multiple images and the time it takes to project the points onto multiple images and a BEV respectively. For each frame of the scene, assuming there are k images and the point cloud contains p points, the reasoning time of the CLIP segmentation model for each image is T, and the final time complexity is O(k(T+αp)), where α is a constant.
3. The self-supervised contrastive learning method based on BEV network according to claim 2 is characterized in that: In the BEV plane, define the coordinates of the projection point on the BEV plane The equation is as follows: in, represent the x and y coordinates of the ith point, respectively. low , Y low They represent the lower limits of the x and y axes of the BEV plane, respectively, and G x , G y represent the grid sizes in the x and y directions of the BEV plane, respectively; During projection, by setting This can filter out redundant ground points and ensure that the background points and foreground points are roughly balanced.
4. The self-supervised contrastive learning method based on BEV network according to claim 3 is characterized in that: In the self-supervised contrastive learning model framework, the voting strategy based on the BEV grid specifically includes: A voting strategy is used to filter the text labels associated with the BEV grid: only text labels that appear ≥ k times and have a confidence level ≥ α are retained as text labels associated with the BEV grid; the equation representing this strategy is as follows: in, represents the text label associated with the jth BEV grid, T c represents the features of the c-th text label, represents the number of points of the cth category projected onto the jth BEV grid, Represents the total number of points projected onto the jth BEV grid; through the voting strategy, we can finally get a mapping that can be used for pre-training 5. The self-supervised contrastive learning method based on BEV network according to claim 4 is characterized in that: The calculation time of voting consists of the time of traversing and counting points, and the time of traversing and judging the counting results. For each frame scene, there are about p paired points and c types of text labels. The final time complexity is O((c+1)p); Assuming there are m frames of scenes, the total time complexity of projection and voting is O(m(kT+(αk+β(c+1)p))), where β is a constant. Since the number of images k and the number of text label categories c are fixed values, the total time complexity can be simplified to O(m(T+γp)), where γ is a constant, indicating that the calculation time of projection and voting is determined by the scale of the dataset (the number of scene frames m), the inference time T of the CLIP model, and the number of points p in each frame.
6. The self-supervised contrastive learning method based on BEV network according to claim 5 is characterized in that: In the self-supervised contrastive learning model framework, the design of the loss function specifically includes: The two loss functions are combined for end-to-end training, and the loss function is: L total =αL TB +(1―α)L BB Among them, α is the balance weight of the loss, L TB is the first loss function, L BB is the second loss function.
7. The self-supervised contrastive learning method based on BEV network according to claim 6 is characterized in that: The first loss function uses text semantic information to select positive and negative samples for contrastive learning. The loss function formula is as follows: Where C is the number of different text categories, t i =T c Indicates t i is the text feature of the cth category, r i Indicates position b on BEV i The characteristic at , τ is the temperature term. The second loss function is used for contrastive learning between BEVs. The loss function formula is: Where M is the number of sample pairs, r i Indicates position b on BEV i The characteristic at r′ i Indicates registration to position b through bilinear interpolation i The characteristic of , τ is the temperature term.
8. A self-supervised contrastive learning device based on a BEV network, characterized in that: Based on the pre-training-fine-tuning architecture, the device includes: A construction module is used to construct a self-supervised contrastive learning model framework, wherein the self-supervised contrastive learning model framework includes the construction of a mapping relationship between text and BEV, a voting strategy based on a BEV grid, and a loss function design; A pre-training module, used for training the encoder of the self-supervised contrastive learning model framework in a pre-training stage to enhance its feature extraction capability; The fine-tuning module is used to add a detection head to the pre-trained encoder during the fine-tuning phase and perform end-to-end training of the entire self-supervised contrastive learning model.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.