A 3D multi-dataset joint training method, electronic device, and storage medium

By introducing data-level correction operation and semantic-level coupling-recombination module joint training method of 3D multi-data sets in the 3D object detection model, the problem of degradation of detection accuracy between different data sets is solved, and better generalization ability and detection accuracy are achieved.

CN116246250BActive Publication Date: 2025-06-24SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310094360.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-31
Publication Date
2025-06-24
Estimated Expiration
2043-01-31

AI Technical Summary

Technical Problem

When the existing 3D object detection model is deployed between different data sets, the detection accuracy will be severely reduced, and it faces the problems of data distribution differences and semantic category-level distribution differences, resulting in obvious data gaps between data sets.

Method used

A joint training method of 3D multi-datasets is proposed. Through data-level correction operations and semantic-level coupling-recombination modules, the difference in point cloud data distribution of different data sets is aligned and reusable feature expressions are mined, thereby improving the generalization ability of 3D detection models.

Benefits of technology

Through the joint training method, the model can learn more representative feature descriptions from multiple 3D data sets, which significantly improves the generalization ability and detection accuracy of the 3D detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246250B_ABST
    Figure CN116246250B_ABST
Patent Text Reader

Abstract

The present application provides a 3D multi-dataset joint training method, which includes: obtaining a number of 3D datasets; jointly training a 3D object detection model using the number of 3D datasets, where the 3D object detection model includes a data-level correction operation and a semantic-level coupling-recombination module, and the data-level correction operation is used to align the point cloud data distribution differences of the number of 3D datasets, and the semantic-level coupling-recombination module is used to mine reusable feature expressions from different 3D datasets. This solution adopts a data-level correction operation and a semantic-level coupling-recombination module, and can learn more representative feature descriptions from multiple 3D datasets, thereby improving the generalization ability of the 3D detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of object detection, and particularly relates to a 3D multi-dataset joint training method, an electronic device, and a storage medium. Background Art

[0002] 3D object detection technology plays a very crucial role in the field of autonomous driving and can help vehicles perceive the surrounding environment. At the same time, LiDAR (Light Detection and Ranging)-based 3D object detection technology aims to use LiDAR sensors to identify and locate instance objects in a given scene. Thanks to the rapid development of large-scale annotated 3D LiDAR datasets, this technology has recently made great progress. Unfortunately, current mainstream 3D object detection models are designed following the classical single-dataset training-test paradigm, which causes these detection models to be difficult to directly deploy into another dataset with a different data distribution. For example, when a baseline detection model is trained on Waymo and evaluated on another different dataset such as nuScenes, the detection accuracy of the detection model will degrade severely (from 74.60% to 17.31%). Therefore, this single-dataset training-test paradigm cannot perform well on different datasets, further impairing the dataset-level generalization ability of current 3D perception models.

[0003] To reduce the differences between different 3D datasets, some researchers have tried to utilize Unsupervised Domain Adaptation (UDA) technology, which aims to transfer a pre-trained source domain detector to a new domain (or dataset). Although these UDA-based 3D object detection works have achieved good detection accuracy gains on the new target domain, they are still a one-way model transfer process from the source domain to the target domain, rather than a two-way multi-dataset generalization process. And multi-dataset generalization is an important guarantee for future improvement of autonomous driving perception capabilities.

[0004] Existing 3D object detection technologies adopt the single-dataset training-test paradigm. When these 3D object detection models are directly deployed in another dataset, they usually face a serious decline in detection accuracy. Moreover, existing 3D datasets face severe challenges of sensor distribution differences and data distribution differences, resulting in obvious data barriers between datasets. Summary of the Invention

[0005] The purpose of the embodiments of this specification is to provide a 3D multi-dataset joint training method, an electronic device, and a storage medium.

[0006] To solve the above technical problems, the embodiments of the present application are implemented in the following ways:

[0007] In a first aspect, the present application provides a 3D multi-dataset joint training method, which includes:

[0008] Obtain a number of 3D datasets;

[0009] Jointly train a 3D object detection model using a number of 3D datasets, where the 3D object detection model includes a data-level correction operation and a semantic-level coupling-recombination module. The data-level correction operation is used to align the differences in the point cloud data distributions of a number of 3D datasets, and the semantic-level coupling-recombination module is used to mine reusable feature expressions from different 3D datasets.

[0010] In one embodiment, the data-level correction operation includes:

[0011] Adopt the mean-variance distribution specific to the dataset to regularize the feature expressions of each neural network layer:

[0012]

[0013] Among them, represents the feature expression of the j-th network layer of the t-th dataset, is to ensure the stability of numerical calculation; represents the mean of the j-th network layer of the t-th dataset, represents the variance of the j-th network layer of the t-th dataset; represents the point cloud data of the j-th network layer of the t-th dataset;

[0014] Restore the expression ability of the model features with the following learnable transformation process, and the specific calculation method is as follows:

[0015]

[0016] Among them, represents the output feature of the j-th network layer of the t-th dataset, and both represent the learnable parameters of the j-th network layer.

[0017] In one embodiment, the semantic-level coupling-recombination module is used to:

[0018] Couple the output features of a number of datasets and learn the transferable feature expressions between datasets to obtain the feature expressions shared by the datasets;

[0019] Perform an internal feature recombination operation within the dataset according to the feature expressions shared by the datasets to restore the feature expressions of the bird's-eye view scene of the dataset.

[0020] In one embodiment, the output features of several data sets are coupled, and the transferable feature expressions between the data sets are learned to obtain the feature expressions shared by the data sets, including:

[0021] The output features of several data sets are concatenated into a unified feature expression, and the feature expression of the bird's-eye view scene features shared by the data sets is learned to obtain the feature expressions shared by the data sets.

[0022] In one embodiment, the output features of several data sets are concatenated into a unified feature expression, and the feature expression of the bird's-eye view scene features shared by the data sets is learned to obtain the feature expressions shared by the data sets , including:

[0023]

[0024] Among them, represents the concatenation operation at the channel level; and represent the bird's-eye view scene features of the i-th and j-th data sets; represents the spatial saliency enhancement operation; represents the saliency enhancement operation at the data set level; represents the element-wise multiplication operation.

[0025] In one embodiment, according to the feature expressions shared by the data sets, an internal feature recombination operation of the data sets is performed to restore the feature expressions of the bird's-eye view scenes of the data sets, including:

[0026] The feature expressions shared by the data sets are used for channel-level reweighting operations to restore the feature expressions of the bird's-eye view scenes of the data sets.

[0027] In one embodiment, the feature expressions shared by the data sets are used for channel-level reweighting operations to restore the feature expressions of the bird's-eye view scenes of the data sets, including:

[0028]

[0029] Among them, SE represents the squeeze-and-excitation network.

[0030] In one embodiment, the loss function of the 3D object detection model is:

[0031]

[0032] Among them, is the error loss predicted by the k-th data set model, represents the detection head determined by the k-th data set.

[0033] In a second aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the 3D multi-dataset joint training method as in the first aspect.

[0034] In a third aspect, the present application provides a readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the 3D multi-dataset joint training method as in the first aspect.

[0035] As can be seen from the technical solutions provided in the embodiments of this specification above, the data-level correction operation and the semantic-level coupling-recombination module in the 3D multi-dataset joint training method provided by this solution can learn more representative feature descriptions from multiple 3D datasets, thereby improving the generalization ability of the 3D detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0037] Figure 1 It is a schematic diagram of the Ui3D framework provided by the present application;

[0038] Figure 2 It is a schematic flowchart of the 3D multi-dataset joint training method provided by the present application;

[0039] Figure 3 It is a schematic diagram of the semantic-level coupling-recombination module provided by the present application;

[0040] Figure 4 It is a schematic diagram of the structure of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0042] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.

[0043] Without departing from the scope or spirit of the present application, various improvements and changes can be made to the specific embodiments of the specification of the present application, which are obvious to those skilled in the art. Other embodiments obtained from the specification of the present application are obvious to those skilled in the art. The specification and embodiments of the present application are merely exemplary.

[0044] Regarding the terms "comprising", "including", "having", "containing", etc. used herein, they are all open-ended terms, meaning including but not limited to.

[0045] The basic concept of the multi-dataset joint training task

[0046] Assume Define the joint probability distribution, where and respectively represent the input point cloud data and the label space. The multi-dataset joint training task (Multi-Domain Fusion, MDF) defines a series of datasets with different joint probability distributions of datasets , which can be used during the training phase. At the same time, the goal of the MDF task is to train a general model from multiple datasets so that the model can obtain a more general mapping expression .

[0047] Assume that in practical applications, we can access multiple annotated 3D point cloud-based domains or datasets (such as Waymo and nuScenes) simultaneously, but these labeled datasets usually have different label spaces . The present application mainly focuses on the multi-dataset joint training task in the autonomous driving scenario, where model training and evaluation are carried out for the categories of interest related to the autonomous driving scenario, such as categories like vehicles, pedestrians, and bicycles. Note that in many studies on cross-dataset 3D object detection, such as ST3D and ST3D++, the common categories in autonomous driving (such as vehicle, pedestrian, and bicycle categories) are also selected for experiments and evaluations.

[0048] This application mainly designs a general module to be integrated into a known object detection framework to form a 3D object detection joint training paradigm. Therefore, this application first presents an optimization method for the known object detection framework, and then introduces the disadvantages of this method and further improvement solutions.

[0049]

[0050] The above formula shows a common 3D object detection loss function, which consists of three parts: 1) Define a preset candidate box loss function, whose purpose is to correct the pre-initialized 3D candidate boxes; 2) Define the region of interest regression loss function in the second stage, whose purpose is to perform class determination and location regression for each instance-level region; 3) Define the foreground per-keypoint loss, whose purpose is to determine which points in the full-scene point cloud are key points.

[0051] Limitations of 3D object detection based on MDF:

[0052] 1) Data-level distribution differences: Compared with 2D natural images composed of pixels, 3D point clouds usually use different point cloud distribution ranges and different sensor types, which leads to very obvious data-level distribution differences between 3D datasets. In addition, the point clouds from different 3D datasets show a more diverse data distribution because 3D datasets are collected in different cities and countries, and the sizes of foreground instance objects are very different.

[0053] 2) Semantic class-level distribution differences: Different autonomous driving manufacturers usually adopt inconsistent class definitions and manual annotation granularities. For example, for the Waymo manufacturer, all vehicles driving on the road, including cars and trucks, are labeled as a unified class, that is, the "Vehicle" class. For the nuScenes manufacturer, different vehicles use different granularity class definition methods for manual annotation, such as the "Car", "Truck", and "Van" classes. Therefore, the multi-dataset joint training task also considers how to train a unified 3D object detection model under an inconsistent classification label space and effectively reuse knowledge across datasets.

[0054] To solve the above data-level distribution differences and semantic class-level distribution differences, this patent proposes a general 3D multi-dataset joint training paradigm (Unified 3D object detection framework, Ui3D) framework, and the specific framework is as Figure 1As shown in the figure, it uses two technical points to enable the model to be trained from multiple 3D datasets. These include data-level correction operations (i.e., data-level correction in the figure) and semantic-level coupling-recombination modules, which are used to alleviate the inevitable differences between input data level and output semantic level in actual autonomous driving scenarios, and can learn more representative feature descriptions from multiple 3D datasets, thereby improving the generalization ability of the 3D detection model.

[0055] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0056] Reference Figure 2 , which shows a flow chart of a 3D multi-dataset joint training method applicable to an embodiment of the present application.

[0057] like Figure 2 As shown, the 3D multi-dataset joint training method may include:

[0058] S210 , obtaining several 3D data sets.

[0059] Specifically, several 3D datasets can be obtained from other datasets such as the Waymo dataset, the nuScenes dataset, the KITTI dataset, etc., and can also be obtained from different lidars, without limitation here.

[0060] S220. Jointly train a 3D object detection model using several 3D datasets, wherein the 3D object detection model includes a data-level correction operation and a semantic-level coupling-recombination module, the data-level correction operation is used to align the point cloud data distribution differences of several 3D datasets, and the semantic-level coupling-recombination module is used to mine reusable feature expressions from different 3D datasets.

[0061] Data Level Corrective Actions: Assumptions and Define the mean and variance of a network layer. Generally speaking, and It is used to regularize the network features of each layer so that each neural network unit can receive input with a data distribution of zero mean. This application found that this unified mean-variance calculation method is not effective in the setting of joint training of multiple data sets. Because a Batch-Size input data may contain multiple data distribution forms with obvious differences. For this reason, this patent uses the specific mean-variance distribution of the data set under the multi-dataset training paradigm to regularize the feature expression of each neural network layer, as follows:

[0062]

[0063] in, Represents the feature expression of the j-th network layer of the t-th dataset. This is to ensure the stability of numerical calculations. Represents the mean value of the j-th network layer of the t-th dataset. Represents the variance of the j-th network layer of the t-th dataset. Represents the point cloud data of the j-th network layer of the t-th dataset.

[0064] Further similar to the BN (Batch Normalization) operation, the expressive ability of the model features is restored using the following parameter-learnable transformation process, and the specific calculation method is as follows:

[0065]

[0066] Among them, Represents the output feature of the j-th network layer of the t-th dataset. and Both represent the learnable parameters of the j-th network layer.

[0067] In one embodiment, the semantic-level coupling-recombination module is used for:

[0068] Couple the output features of several datasets, learn the transferable feature expressions between datasets, and obtain the feature expressions shared by the datasets;

[0069] Perform an internal feature recombination operation within the dataset based on the feature expressions shared by the datasets to restore the feature expressions of the bird's-eye view scene of the dataset.

[0070] Among them, coupling the output features of several datasets, learning the transferable feature expressions between datasets, and obtaining the feature expressions shared by the datasets includes:

[0071] Concatenate the output features of several datasets into a unified feature expression, and learn the feature expression of the bird's-eye view scene features shared by the datasets to obtain the feature expressions shared by the datasets.

[0072] Specifically, the structure of the semantic-level coupling-recombination module is as Figure 3 shown, obtaining the bird's-eye view scene features (Bird Eye View, BEV) from multiple datasets (exemplarily, Figure 1 in is the bird's-eye view scene feature of dataset one, is the bird's-eye view scene feature of dataset two), then concatenate these multi-dataset features into a unified feature expression, and then learn the BEV feature expression shared by the datasets to obtain the feature expressions shared by the datasets :

[0073]

[0074] Among them, represents a splicing operation at the channel level; and represent the bird's-eye view scene features of the i-th and j-th datasets; represents a spatial saliency enhancement operation; represents a saliency enhancement operation at the dataset level; represents an element-wise multiplication operation, the purpose of which is to enhance the feature expressions that can be reused by multiple datasets.

[0075] It can be understood that the bird's-eye view scene features obtained from multiple datasets are the output features of different datasets output by the above data-level correction operation.

[0076] In one embodiment, according to the feature expressions shared by the datasets, an internal feature recombination operation of the datasets is performed to restore the feature expressions of the bird's-eye view scene of the datasets, including:

[0077] Using a channel-level reweighting operation on the feature expressions shared by the datasets to restore the feature expressions of the bird's-eye view scene of the datasets:

[0078]

[0079] Among them, SE represents a Squeeze-and-Excitation Network (SENet).

[0080] In one embodiment, the loss function of the 3D object detection model is:

[0081]

[0082] Among them, is the error loss predicted by the model of the k-th dataset, represents the detection head determined by the k-th dataset.

[0083] This application proposes a multi-3D dataset joint training strategy, which can enable the current 3D perception model to perform the task of jointly training with multiple 3D datasets, and is a necessary step to improve the performance of the future autonomous driving perception model.

[0084] Existing 3D datasets face the challenges of severe sensor distribution differences and data distribution differences, resulting in obvious data gaps between datasets. The data-level correction operation and semantic-level coupling-recombination module in the 3D multi-dataset joint training method provided in this application can learn more representative feature descriptions from multiple 3D datasets, thereby improving the generalization ability of the 3D detection model.

[0085] According to the characteristics of the point cloud data distribution, this application proposes a data-level correction operation and a semantic-level coupling-recombination module. First, it aligns the differences in the dataset distributions from different datasets and mines the reusable feature representations between different datasets, so as to improve the data reusability of the 3D perception model.

[0086] The 3D multi-dataset joint training method provided in this application is simple and easy to combine with current mainstream 3D object detection baseline models (such as PV-RCNN and Voxel RCNN), enabling them to effectively learn from multiple 3D datasets to obtain more discriminative and generalizable feature representations. We conducted experiments in many dataset merging scenarios, including the Waymo-nuScenes merging operation, nuScenes-KITTI merging operation, Waymo-KITTI merging operation, and Waymo-nuScenes-KITTI merging operation, and achieved significant target domain detection accuracy. Among them, in the KITTI-nuScenes-Waymo dataset merging operation, Uni3D achieved 75.18 mAP and 66.89 mAPH, which were respectively improved by 21.91 and 29.66 compared with the single-dataset training paradigm (53.27 mAP and 37.23 mAPH). Experiments prove that Uni3D provided in this application exceeds the performance of object detectors trained on a single dataset. This work will inspire research on 3D generalization because it will push the limits of perception performance.

[0087] Figure 4 It is a schematic structural diagram of an electronic device provided for an embodiment of the present invention. As Figure 4 shown, it shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of this application.

[0088] As Figure 4As shown, the electronic device 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage section 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0089] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 406 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read therefrom is installed into the storage section 408 as needed.

[0090] In particular, according to an embodiment of the present disclosure, the process described above with reference to Figure 1 can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for performing the above-described 3D multi-dataset joint training method. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 409, and / or installed from the removable medium 411.

[0091] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0092] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. The names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.

[0093] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a mobile phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0094] On the other hand, this application also provides a storage medium. The storage medium can be the storage medium included in the aforementioned device in the above embodiments; it can also exist separately and be not assembled into the device. The storage medium stores one or more programs, and the aforementioned programs are used by one or more processors to execute the 3D multi-dataset joint training method described in this application.

[0095] A storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0096] It should be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0097] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

Claims

1. A 3D multi-dataset joint training method, characterized in that The method includes: Obtaining a plurality of 3D data sets; Jointly training a 3D object detection model using the plurality of 3D data sets, wherein the 3D object detection model includes a data-level correction operation module and a semantic-level coupling-recombination module, the data-level correction operation module is used to align the point cloud data distribution differences of the plurality of 3D data sets, and the semantic-level coupling-recombination module is used to mine reusable feature expressions from different 3D data sets; The data-level correction operation includes: Regularizing the feature expressions of each neural network layer using the mean-variance distribution specific to the data set: Among them, represents the feature expression of the j-th network layer of the t-th data set, is to ensure the stability of numerical calculations; represents the mean of the j-th network layer of the t-th data set, represents the variance of the j-th network layer of the t-th data set; represents the point cloud data of the j-th network layer of the t-th data set; Using the following parameter-learnable transformation process to restore the expression ability of the model features, and the specific calculation method is as follows: Among them, represents the output feature of the j-th network layer of the t-th data set, and both represent the learnable parameters of the j-th network layer; The semantic-level coupling-recombination module is used to: Couple the output features of the plurality of data sets, and learn the transferable feature expressions between the data sets to obtain the feature expressions shared by the data sets; Perform an internal feature recombination operation within the data set according to the feature expressions shared by the data sets to restore the feature expressions of the bird's-eye view scene of the data set.

2. The method according to claim 1, wherein The coupling of the output features of the plurality of data sets, and learning the transferable feature expressions between the data sets to obtain the feature expressions shared by the data sets includes: Concatenating the output features of the plurality of data sets into a unified feature expression, and learning the feature expression of the bird's-eye view scene features shared by the data sets to obtain the feature expressions shared by the data sets.

3. The method according to claim 2, wherein Concatenating the output features of the several data sets into a unified feature representation, and learning a feature representation of the bird's-eye view scene features shared by the data sets to obtain a feature representation shared by the data sets , including: Among them, represents the splicing operation at the channel level; and represent the bird's-eye view scene features of the i-th and j-th data sets; represents the spatial saliency enhancement operation; represents the saliency enhancement operation at the data set level; represents the element-wise multiplication operation; Conv represents the convolution operation of the features on the bird's-eye view scene features.

4. The method according to claim 3, wherein The performing an internal feature recombination operation within the data set according to the feature expressions shared by the data sets to restore the feature expressions of the bird's-eye view scene of the data set includes: Using a channel-level reweighting operation on the feature expressions shared by the data sets to restore the feature expressions of the bird's-eye view scene of the data set.

5. The method according to claim 4, characterized in that The using a channel-level reweighting operation on the feature expressions shared by the data sets to restore the feature expressions of the bird's-eye view scene of the data set includes: Among them, and represent the feature expressions of the bird's-eye view scenes of the i-th and j-th data sets, and SE represents the squeeze-and-excitation network.

6. The method according to any one of claims 1-5, characterized in that, The loss function of the 3D object detection model is as follows: Among them, is the error loss predicted by the k-th dataset model, represents the detection head determined by the k-th dataset, represents the feature expression of the bird's-eye view scene of the k-th dataset.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the 3D multi-data set joint training method described in any one of claims 1-6.

8. A readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the 3D multi-data set joint training method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Image semantic segmentation method based on double category level adversarial network

    CN114612658A

  • Cross-domain target detection model training method, target detection method and storage medium

    CN115115908A