Method, device, equipment, storage medium and program product for model training

By training a recognition model based on point cloud and image features in an autonomous driving system, the problem of misjudgment caused by the lack of accuracy of perception data is solved, and the accuracy of the recognition model and the recognition capability of the autonomous driving system are improved.

CN122116022APending Publication Date: 2026-05-29BEIJING VOYAGER TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING VOYAGER TECH CO LTD
Filing Date
2024-11-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In the field of autonomous driving, the lack of accuracy in perception data leads to misjudgments by recognition models, affecting the decision-making of autonomous driving systems.

Method used

By determining the 3D region based on point cloud data, point cloud features and image features are obtained. These features are then processed using a recognition model to determine the predicted recognition information and occlusion rate. The recognition model is then trained based on the training loss, which includes the difference between the predicted recognition information and the labeled recognition information, as well as the difference between the predicted occlusion rate and the reference occlusion rate.

Benefits of technology

It improved the accuracy of the recognition model and enhanced the recognition capabilities of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116022A_ABST
    Figure CN122116022A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method, device, equipment, storage medium and program product for model training. The method comprises: determining a three-dimensional region based on point cloud data; obtaining point cloud features and image features corresponding to the three-dimensional region, the image features being determined by projecting the three-dimensional region onto a two-dimensional plane corresponding to image data; processing the point cloud features and the image features using a recognition model to determine predicted recognition information of the three-dimensional region and a predicted occlusion rate of the three-dimensional region; determining a training loss based on the predicted recognition information and the predicted occlusion rate, the training loss comprising at least a first part and a second part, the first part indicating a first difference between the predicted recognition information and labeled recognition information, and the second part indicating a second difference between the predicted occlusion rate and a reference occlusion rate; and training the recognition model based on the training loss. According to embodiments of the present disclosure, the recognition capability of the recognition model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of autonomous driving, and particularly to a method, apparatus, device, computer-readable storage medium, and program product for model training. Background Technology

[0002] With the rapid development of technology, perception technologies that utilize multiple sensors working together are widely used in the field of autonomous driving to help recognition models in autonomous driving systems identify the driving environment. For example, perception technologies that use radar and cameras as core sensors.

[0003] However, in some scenarios, due to the lack of accuracy in the perceived data, the recognition model may make misjudgments. Summary of the Invention

[0004] In a first aspect of this disclosure, a method for model training is provided. The method includes: determining a three-dimensional region based on point cloud data; acquiring point cloud features and image features corresponding to the three-dimensional region, the image features being determined by projecting the three-dimensional region onto a two-dimensional plane corresponding to the image data; processing the point cloud features and the image features using a recognition model to determine predicted recognition information and a predicted occlusion rate of the three-dimensional region; determining a training loss based on the predicted recognition information and the predicted occlusion rate, the training loss including at least a first part and a second part, the first part indicating a first difference between the predicted recognition information and labeled recognition information, the second part indicating a second difference between the predicted occlusion rate and a reference occlusion rate; and training the recognition model based on the training loss.

[0005] In a second aspect of this disclosure, an apparatus for model training is provided. The apparatus includes: a first determining module configured to determine a three-dimensional region based on point cloud data; an acquiring module configured to acquire point cloud features and image features corresponding to the three-dimensional region, the image features being determined by projecting the three-dimensional region onto a two-dimensional plane corresponding to the image data; a second determining module configured to process the point cloud features and the image features using a recognition model to determine predicted recognition information and a predicted occlusion rate of the three-dimensional region; a third determining module configured to determine a training loss based on the predicted recognition information and the predicted occlusion rate, the training loss including at least a first part and a second part, the first part indicating a first difference between the predicted recognition information and labeled recognition information, and the second part indicating a second difference between the predicted occlusion rate and a reference occlusion rate; and a training module configured to train the recognition model based on the training loss.

[0006] In a third aspect of this disclosure, a computing device is provided. The device includes at least one processing unit and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of this disclosure, a computer program is provided. The computer program includes computer-executable instructions that, when executed by a processor, implement the method according to a first aspect of this disclosure.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;

[0012] Figure 2 A flowchart illustrating an example process for training a recognition model according to some embodiments of this disclosure is shown;

[0013] Figure 3 A flowchart illustrating an example process for acquiring point cloud features according to some embodiments of the present disclosure is shown;

[0014] Figure 4 A flowchart illustrating an example process for acquiring image features according to some embodiments of the present disclosure is shown;

[0015] Figure 5 A flowchart illustrating an example process for training a recognition model according to some embodiments of the present disclosure is shown;

[0016] Figure 6 A flowchart illustrating an example process for obtaining a reference occlusion rate according to some embodiments of the present disclosure is shown;

[0017] Figure 7A schematic structural block diagram of an example apparatus for model training according to some embodiments of the present disclosure is shown; and

[0018] Figure 8 A block diagram of a computing device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0019] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0020] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0021] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0022] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0023] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0024] As mentioned above, recognition models typically detect the driving environment based on perception data collected by sensors. Image data captured by cameras may contain multiple objects that create occlusion. When a recognition model detects this type of image data, it is susceptible to interference from the image data. This interference can cause the recognition model's results to deviate from the actual driving environment, thus affecting the decision-making of the autonomous driving system.

[0025] Embodiments of this disclosure propose a scheme for model training. The scheme includes: determining a three-dimensional region based on point cloud data; acquiring point cloud features and image features corresponding to the three-dimensional region, wherein the image features are determined by projecting the three-dimensional region onto a two-dimensional plane corresponding to the image data; processing the point cloud features and image features using a recognition model to determine predicted recognition information and predicted occlusion rate of the three-dimensional region; determining a training loss based on the predicted recognition information and predicted occlusion rate, the training loss including at least a first part and a second part, the first part indicating a first difference between the predicted recognition information and the labeled recognition information, and the second part indicating a second difference between the predicted occlusion rate and the reference occlusion rate; and training the recognition model based on the training loss.

[0026] According to embodiments of this disclosure, by taking occlusion rate into account during the training of the recognition model, embodiments of this disclosure can improve the accuracy of the recognition model.

[0027] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0028] Example Environment

[0029] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, the example environment 100 may include a model training system 120 for training the recognition model 110 and a model application system 130 for applying the recognition model 110. The recognition model 110 is configured to detect received perception data 140 to identify objects in the driving environment indicated by the perception data 140. These objects may be vehicles, pedestrians, animals, or obstacles, etc. Figure 1 The upper part shows the training phase of the recognition model 110, and the lower part shows the inference phase of the recognition model 110.

[0030] For training the recognition model 110, training data 150-1, 150-2, ..., 150-N (where N is an integer greater than or equal to 1) can be used to train the recognition model 110. In this paper, training data 150-1, 150-2, ..., 150-N can be collectively referred to as training data 150 or referred to individually.

[0031] In some embodiments, the training data 150 used to train the recognition model 110 can be divided into multiple sets of training data 150. Each set of training data 150 includes a pair of feature data, prediction recognition information 160 corresponding to the feature data, prediction occlusion rate 170, annotation recognition information 180, and reference occlusion rate 190. In one set of training data 150, the pair of feature data consists of point cloud features 152 and image features 154 corresponding to the same object; the prediction recognition information 160 is the negative output sample of one task A of the recognition model 110, and the annotation recognition information 180 is the positive output sample of the recognition model 110 in task A; the prediction occlusion rate 170 is the negative output sample of another task B of the recognition model 110, and the reference occlusion rate 190 is the positive output sample of the recognition model 110 in task B. By training the recognition model 110 using multiple sets of training data 150, the recognition model 110 can be used to detect the perceived data 140 during the inference stage.

[0032] In some embodiments, during the inference phase of the recognition model 110, the recognition model 110 can detect a processed set of perception data 140 and output predicted recognition information 160 to identify objects in the driving environment indicated by the perception data 140. The set of perception data 140 includes point cloud data 142 and image data 144 corresponding to the same object. As an example, the point cloud data 142 and image data 144 can be processed to obtain a processing result that facilitates detection by the recognition module 110.

[0033] exist Figure 1In this embodiment, the model training system 120 and the model application system 130 can be any electronic device with computing capabilities. In some embodiments, the electronic device can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device can also support any type of user-facing interface (such as "wearable" circuitry).

[0034] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0035] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0036] Example process

[0037] Figure 2 A flowchart illustrating an example process 200 for training a recognition model 110 according to some embodiments of the present disclosure is shown. Process 200 can be implemented at a suitable electronic device. Reference is made below. Figure 1 To describe process 200.

[0038] like Figure 2 As shown in box 210, the electronic device determines a three-dimensional region 310 based on point cloud data 142.

[0039] In some embodiments, point cloud data 142 is a dataset representing points in space of a three-dimensional shape or object, which can describe information such as the three-dimensional coordinates, reflectivity, texture, and color of multiple dense points on the object's surface. Point cloud data 142 can be acquired using a three-dimensional scanning device. As an example, point cloud data 142 can be obtained by scanning the environment using a LiDAR scanner.

[0040] In embodiments of this disclosure, point cloud data 142 can indicate the driving environment of an autonomous vehicle. To accurately identify the driving environment, an entity segmentation model can be used to process the point cloud data 142. In some embodiments, by processing the point cloud data 142 using entity segmentation technology, a 3D bounding box corresponding to an object in the driving environment can be determined. In this document, the region corresponding to the 3D bounding box can be referred to as 3D region 310. 3D region 310 represents the presence of entity objects in this space.

[0041] By identifying the point cloud data 142, it can be determined whether there are regions indicating the presence of objects at corresponding locations. When it is determined that the point cloud data 142 reflects the presence of objects in the driving environment, a three-dimensional region 310 is obtained by segmenting the point cloud data 142. The point cloud within this three-dimensional region 310 represents objects in the current driving environment.

[0042] In box 220, the electronic device acquires point cloud features 152 and image features 154 corresponding to the three-dimensional region 310.

[0043] To more accurately detect the driving environment, a method combining point cloud and image data is employed. However, point cloud data, as three-dimensional data, and image data, as two-dimensional data, belong to different dimensions. They operate in different coordinate systems, making it difficult to combine them. In some embodiments, the electronic device can facilitate the identification of the driving environment by acquiring point cloud features 152 and image features 154 that correspond to the three-dimensional region 310 and are located in the same coordinate system.

[0044] Figure 3 A flowchart illustrating an example process 300 for acquiring point cloud features 152 according to some embodiments of the present disclosure is shown. (Refer to...) Figure 3In some embodiments, the electronic device can obtain point cloud features 152 through the following process: First, the electronic device determines a set of points 320 corresponding to the three-dimensional region 310 based on point cloud data 142. The electronic device then projects the set of points 320 onto a preset plane to generate a projection map 330. Then, the electronic device generates point cloud features 152 corresponding to the three-dimensional region 310 based on the projection map 330. The region where the three-dimensional region 310 is located in the point cloud data 142 corresponds to the region where objects are located in the driving environment. The set of points 320 within the three-dimensional region 310 in the point cloud data 142 can reflect objects in the driving environment. By projecting the set of points 320 onto the preset plane, the electronic device can obtain multiple projection points of the set of points 320 in a two-dimensional plane. In other words, the projection map 330 can indicate the distribution of the projection points of the set of points 320 in a two-dimensional plane. Based on this, the electronic device can use image processing technology to segment the multiple projection points of the set of points 320 in the two-dimensional plane from the projection map 330 to obtain point cloud features 152 corresponding to the three-dimensional region 310.

[0045] In some embodiments, point cloud data 142 can be projected onto a preset plane to obtain a projection map 330, and then point cloud features 152 can be generated based on the projection map 330.

[0046] In some embodiments, image features 154 may be determined by projecting a three-dimensional region 310 onto a two-dimensional plane corresponding to image data 144. Image data 144 refers to perception data 140 acquired using a camera.

[0047] In order to associate image feature 154 with point cloud feature 152 for detection by recognition model 110, the object region corresponding to image feature 154 must be consistent with the object region corresponding to point cloud feature 152. Therefore, the electronic device can project the three-dimensional region 310 onto the two-dimensional plane corresponding to image data 144 to determine image feature 154.

[0048] Figure 4 A flowchart illustrating an example process 400 for acquiring image features 154 according to some embodiments of the present disclosure is shown. (Refer to...) Figure 4 In some embodiments, the electronic device can determine a set of vertices 410 corresponding to the three-dimensional region 310, and based on the transformation relationship 420 between the first coordinate system corresponding to the three-dimensional region 310 and the second coordinate system corresponding to the image data 144, determine a set of coordinates 430 in the second coordinate system corresponding to the set of vertices 410. The electronic device then determines a two-dimensional region 440 in the two-dimensional plane based on the set of coordinates 430. Subsequently, the electronic device determines image features 154 corresponding to the two-dimensional region 440 based on the image data 144.

[0049] In some embodiments, as the spatial shape of the three-dimensional region 310 changes, the set of vertices 410 corresponding to the three-dimensional region 310 also changes. For example, the three-dimensional region 310 can be a cuboid, a sphere, or other solid geometric shapes. When the three-dimensional region 310 is a cuboid, there are eight vertices 410 corresponding to it. The coordinates of these eight vertices 410 can be obtained based on point cloud data 142.

[0050] While determining the coordinates of a set of vertices 410 corresponding to the three-dimensional region 310, it is also necessary to determine the first coordinate system in which the three-dimensional region 310 resides and the second coordinate system in which the image data 144 resides, so as to retrieve the transformation relationship 420 between the first and second coordinate systems. Using the transformation relationship 420 between the first and second coordinate systems, the electronic device can transform the coordinates of a set of vertices 410 to the second coordinate system to obtain the corresponding set of coordinates 430. The two-dimensional region 440 formed by the points corresponding to this set of coordinates 430 in the second coordinate system can correspond to the three-dimensional region 310.

[0051] In some embodiments, when an electronic device uses a set of coordinates 430 on a second coordinate system to determine a two-dimensional region 440, the set of coordinates 430 can be connected sequentially to form the two-dimensional region 440, or the circumscribed rectangle corresponding to the set of coordinates 430 can be determined first. Then, the electronic device determines the two-dimensional region 440 based on the circumscribed rectangle.

[0052] Once the two-dimensional region 440 is determined, the electronic device can obtain the image feature 154 by determining the image content corresponding to the two-dimensional region 440 on the image data 144.

[0053] In frame 230, the electronic device uses recognition model 110 to process point cloud features 152 and image features 154 to determine the predicted recognition information 160 of the three-dimensional region 310 and the predicted occlusion rate 170 of the three-dimensional region 310.

[0054] Figure 5 A flowchart illustrating an example process 500 for training a recognition model 110 according to some embodiments of the present disclosure is shown. (Refer to...) Figure 5 The recognition model 110 is used to recognize the content corresponding to the three-dimensional region 310, that is, to recognize the objects represented by the point cloud features 152 and the image features 154. After the point cloud features 152 and the image features 154 are input to the recognition model 110, the recognition model 110 outputs the recognition result of the three-dimensional region 310. As an example, the recognition result can be the predicted recognition information 160 and the predicted occlusion rate 170.

[0055] The predictive recognition information 160 is prediction information related to the content corresponding to the three-dimensional region 310, that is, prediction of information related to the object represented by the point cloud features 152 and image features 154. Examples can be at least one of the object type corresponding to the three-dimensional region 310 and the specific classification associated with that object type. For example, the object corresponding to the three-dimensional region 310 is a cat. The object type corresponding to the three-dimensional region 310 is an animal. The specific classification associated with that object type is a cat. The predictive recognition information 160 can be an animal, can be a cat, or can be both an animal and a cat.

[0056] The predicted occlusion rate 170 is a prediction of the proportion of the object corresponding to the three-dimensional region 310 that is occluded. In some embodiments, the predicted occlusion rate 170 may be the ratio of the number of occluded pixels in the three-dimensional region 310 to the total number of pixels in the three-dimensional region 310. When the object corresponding to the three-dimensional region 310 is occluded, the image data 144 may not reflect the occlusion relationship. This interferes with the recognition model 110. Specifically, the more occluded the region, the greater the interference. In other words, when the object corresponding to the three-dimensional region 310 is occluded, the confidence level of the image data 144 for the recognition model 110 decreases. By predicting the occlusion rate, the confidence level of the image data 144 can be evaluated, thereby correcting the predicted recognition information 160.

[0057] In box 240, the electronic device determines the training loss based on the predicted recognition information 160 and the predicted occlusion rate 170.

[0058] When training the recognition model 110, the predicted occlusion rate 170 can have some impact on the predicted recognition information 160. When the predicted occlusion rate 170 is not accurate enough, the predicted recognition information 160 also lacks accuracy. Therefore, by determining the training loss and then training the recognition model 110 based on the training loss, the electronic device can improve the recognition capability of the recognition model 110.

[0059] In some embodiments, the training loss may include at least a first part and a second part. The first part indicates a first difference between the predicted identification information 160 and the labeled identification information 180. The second part indicates a second difference between the predicted occlusion rate 170 and the reference occlusion rate 190.

[0060] The annotation identification information 180 is real information related to the content corresponding to the three-dimensional region 310. Similar to the predicted identification information 160, the annotation identification information 180 can also be at least one of the type corresponding to the three-dimensional region 310 and the specific classification associated with the type corresponding to the three-dimensional region 310. The annotation identification information 180 can be obtained through manual identification.

[0061] Figure 6A flowchart of an example process 600 for obtaining a reference occlusion rate 190 according to some embodiments of the present disclosure is shown. (Refer to...) Figure 6 The reference occlusion rate 190 is the actual proportion of the object corresponding to the three-dimensional region 310 that is occluded. In some embodiments, the reference occlusion rate 190 can be obtained as follows: First, the electronic device acquires an occlusion map 610, which indicates whether multiple locations in space are occluded. Then, based on the occlusion map 610, the electronic device determines the reference occlusion rate 190 corresponding to the three-dimensional region 310.

[0062] In some embodiments, the occlusion map 610 can be obtained from a texture map generated based on occlusion mapping technology. The occlusion map 610 can be presented in the form of a grayscale image. The higher the grayscale value of a pixel in the occlusion map 610, the higher the occlusion rate of that pixel. For the occlusion map 610, the occlusion rate of a pixel can be used to determine whether the pixel is occluded. Based on this, the electronic device can determine which areas in space are occluded based on the areas formed by the occluded pixels.

[0063] In some embodiments, before determining the reference occlusion rate 190 corresponding to the three-dimensional region 310 based on the occlusion map 610, the electronic device first projects the three-dimensional region 310 onto the viewing angle corresponding to the occlusion map 610 to determine the target region 620. The target region 620 is the region on the occlusion map 610 corresponding to the three-dimensional region 310. Then, based on the grayscale values ​​of multiple pixels in the target region 620 on the occlusion map 610, the electronic device can determine the number of occluded pixels. The electronic device can further use the number of occluded pixels to determine the reference occlusion rate 190 corresponding to the three-dimensional region 310. As an example, the reference occlusion rate 190 corresponding to the three-dimensional region 310 can be the ratio of the number of occluded pixels in the target region 620 to the total number of pixels in the target region 620.

[0064] In some embodiments, the training loss can be determined as follows: First, the electronic device determines a first weight corresponding to the predicted recognition information 160 and a second weight corresponding to the predicted occlusion rate 170. Then, the electronic device applies the first weight and the second weight to a first difference and a second difference, respectively. The first weight and the second weight can be obtained from multiple experimental data.

[0065] In some embodiments, before determining the training loss, the electronic device may first quantize the first difference and the second difference, and then determine the training loss using a formula for calculating the training loss. As an example, the formula for calculating the training loss may be the sum of the product of the quantized value of the first difference and the first weight, and the product of the quantized value of the second difference and the second weight.

[0066] In box 250, the electronic device trains recognition model 110 based on training loss.

[0067] In some embodiments, the electronic device can adjust the parameters of the recognition model 110 by minimizing the training loss until the training converges.

[0068] In some embodiments, during the inference phase of the recognition model 110, the unit in the recognition model 110 used to output the predicted occlusion rate 170 can be disabled. That is, the recognition model 110 can output only the object type corresponding to the 3D region and the specific classification associated with the object type during the inference phase, without outputting the predicted occlusion rate 170. In this way, the embodiments of this disclosure can improve the inference latency of the recognition model 110 and improve the processing efficiency of the model.

[0069] Example devices and equipment

[0070] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 7 A schematic structural block diagram of an example apparatus 700 for model training according to certain embodiments of the present disclosure is shown. Apparatus 700 may be implemented as or included in an electronic device. Various modules / components in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0071] like Figure 7 As shown, the device 700 includes: a first determining module 710 configured to determine a three-dimensional region based on point cloud data; an acquisition module 720 configured to acquire point cloud features and image features corresponding to the three-dimensional region, wherein the image features are determined by projecting the three-dimensional region onto a two-dimensional plane corresponding to the image data; a second determining module 730 configured to process the point cloud features and image features using a recognition model to determine predicted recognition information and predicted occlusion rate of the three-dimensional region; a third determining module 740 configured to determine a training loss based on the predicted recognition information and predicted occlusion rate, wherein the training loss includes at least a first part and a second part, the first part indicating a first difference between the predicted recognition information and the labeled recognition information, and the second part indicating a second difference between the predicted occlusion rate and the reference occlusion rate; and a training module 750 configured to train a recognition model based on the training loss.

[0072] In some embodiments, the method further includes: acquiring an occlusion map, the occlusion map indicating whether multiple locations in the space are occluded; and determining a reference occlusion rate corresponding to the three-dimensional region based on the occlusion map.

[0073] In some embodiments, determining the training loss based on the predicted identification information and the predicted occlusion rate includes: determining a first weight corresponding to the predicted identification information and a second weight corresponding to the predicted occlusion rate; and applying the first weight and the second weight to a first difference and a second difference, respectively, to determine the training loss.

[0074] In some embodiments, the predicted identification information includes at least one of the following: the object type corresponding to the three-dimensional region; and the specific classification associated with the object type.

[0075] In some embodiments, the method further includes: determining a set of vertices corresponding to a three-dimensional region; determining a set of coordinates in the second coordinate system corresponding to the set of vertices based on the transformation relationship between the first coordinate system corresponding to the three-dimensional region and the second coordinate system corresponding to the image data; determining a two-dimensional region in a two-dimensional plane based on the set of coordinates; and determining image features corresponding to the two-dimensional region based on the image data.

[0076] In some embodiments, determining a two-dimensional region in a two-dimensional plane based on a set of coordinates includes: determining the circumscribed rectangle corresponding to the set of coordinates; and determining the two-dimensional region in the two-dimensional plane based on the circumscribed rectangle.

[0077] In some embodiments, the method further includes: determining a set of points corresponding to a three-dimensional region based on point cloud data; projecting the set of points onto a preset plane to generate a projection map; and generating point cloud features corresponding to the three-dimensional region based on the projection map.

[0078] In some embodiments, during the inference phase of the recognition model, the unit in the recognition model used to output the occlusion rate is disabled.

[0079] like Figure 8 As shown, computing device 800 is in the form of a general-purpose electronic device. Components of computing device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage devices 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 800.

[0080] Computing device 800 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to computing device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within computing device 800.

[0081] The computing device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 8 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 820 may include computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0082] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the components of the computing device 800 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0083] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 800 can also communicate as needed with one or more external devices (not shown) via communication unit 840. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 800, or with any device that enables computing device 800 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0084] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0085] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0086] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0087] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0089] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for model training, comprising: Based on point cloud data, determine the three-dimensional region; Obtain point cloud features and image features corresponding to the three-dimensional region, wherein the image features are determined by projecting the three-dimensional region onto a two-dimensional plane corresponding to the image data; The point cloud features and the image features are processed using a recognition model to determine the predicted recognition information of the three-dimensional region and the predicted occlusion rate of the three-dimensional region. Based on the predicted identification information and the predicted occlusion rate, a training loss is determined. The training loss includes at least a first part and a second part. The first part indicates a first difference between the predicted identification information and the labeled identification information, and the second part indicates a second difference between the predicted occlusion rate and the reference occlusion rate. as well as The recognition model is trained based on the training loss.

2. The method according to claim 1, further comprising: Obtain an occlusion map, which indicates whether multiple locations in the space are occluded; as well as Based on the occlusion map, the reference occlusion rate corresponding to the three-dimensional region is determined.

3. The method according to claim 1, wherein determining the training loss based on the predicted recognition information and the predicted occlusion rate includes: Determine a first weight corresponding to the predicted identification information and a second weight corresponding to the predicted occlusion rate; The first weight and the second weight are applied to the first difference and the second difference, respectively, to determine the training loss.

4. The method according to claim 1, wherein the predicted identification information includes at least one of the following: The object type corresponding to the three-dimensional region; The specific category associated with the object type.

5. The method according to claim 1, further comprising: Determine a set of vertices corresponding to the three-dimensional region; Based on the transformation relationship between the first coordinate system corresponding to the three-dimensional region and the second coordinate system corresponding to the image data, a set of coordinates in the second coordinate system corresponding to the set of vertices is determined. Based on the set of coordinates, a two-dimensional region in the two-dimensional plane is determined; as well as Based on the image data, the image features corresponding to the two-dimensional region are determined.

6. The method according to claim 5, wherein determining the two-dimensional region in the two-dimensional plane based on the set of coordinates comprises: Determine the bounding rectangle corresponding to the set of coordinates; as well as Based on the circumscribed rectangle, the two-dimensional region in the two-dimensional plane is determined.

7. The method according to claim 1, further comprising: Based on the point cloud data, determine the set of points corresponding to the three-dimensional region; The set of points is projected onto a preset plane to generate a projection map; as well as Based on the projection map, the point cloud features corresponding to the three-dimensional region are generated.

8. The method according to claim 1, wherein: During the inference phase of the recognition model, the unit in the recognition model used to output the occlusion rate is disabled.

9. An apparatus for model training, comprising: The first determining module is configured to determine the three-dimensional region based on point cloud data; The acquisition module is configured to acquire point cloud features and image features corresponding to the three-dimensional region, wherein the image features are determined by projecting the three-dimensional region onto a two-dimensional plane corresponding to the image data. The second determining module is configured to process the point cloud features and the image features using a recognition model to determine the predicted recognition information of the three-dimensional region and the predicted occlusion rate of the three-dimensional region. The third determining module is configured to determine a training loss based on the predicted identification information and the predicted occlusion rate. The training loss includes at least a first part and a second part. The first part indicates a first difference between the predicted identification information and the labeled identification information, and the second part indicates a second difference between the predicted occlusion rate and the reference occlusion rate. as well as The training module is configured to train the recognition model based on the training loss.

10. A computing device, comprising: At least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the computing device to perform the method according to any one of claims 1 to 8 when executed by the at least one processing unit.

11. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 8.

12. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.