Method, device and storage medium for training object recognition model

By removing ineligible 3D bounding boxes during multi-task model training, the training data is optimized, resolving the data alignment issue between 2D and 3D detection tasks and improving the accuracy and training efficiency of the object recognition model.

CN122135332APending Publication Date: 2026-06-02BEIJING VOYAGER TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING VOYAGER TECH CO LTD
Filing Date
2024-12-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In the training process of existing multi-task models, the data alignment between 2D and 3D detection tasks is poor, which leads to a decrease in model performance. In addition, traditional training methods have problems such as high data annotation costs and small training samples.

Method used

By acquiring the two-dimensional and three-dimensional bounding boxes of the sample images, the projection area is determined, and the three-dimensional bounding boxes that do not meet the conditions are removed according to the positional relationship. The optimized annotation information is used to train the object recognition model and output the two-dimensional and three-dimensional positions.

Benefits of technology

It improves the accuracy and training efficiency of the object recognition model, ensures that the model focuses on contributing data, and enhances the learning effect of multiple detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135332A_ABST
    Figure CN122135332A_ABST
Patent Text Reader

Abstract

Embodiments of this disclosure provide a method, apparatus, device, and storage medium for training an object recognition model. The method includes acquiring annotation information corresponding to a target object in a sample image, the annotation information including two-dimensional bounding boxes and three-dimensional bounding boxes; determining a projection region in the sample image based on the three-dimensional bounding boxes; removing the three-dimensional bounding boxes from the annotation information in response to a preset condition being met by the positional relationship between the projection region and the two-dimensional bounding boxes; and training an object recognition model using the sample image and the removed annotation information, the object recognition model being configured to output the two-dimensional and three-dimensional positions of objects of a preset type. This improves the accuracy and training efficiency of the object recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, computer-readable storage media, and computer program products for training object recognition models. Background Technology

[0002] In the field of autonomous driving, object detection is an essential function for autonomous vehicles to perceive traffic conditions. To achieve stable detection of diverse environments and objects, a multi-task model capable of performing both 2D and 3D detection tasks can be trained simultaneously. This multi-task model can then output both 2D and 3D information about the detected objects. Improving the training efficiency and accuracy of this multi-task model is a key focus. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for training an object recognition model is provided. The method includes: acquiring annotation information corresponding to a target object in a sample image, the annotation information including two-dimensional bounding boxes and three-dimensional bounding boxes; determining a projection region in the sample image based on the three-dimensional bounding boxes; removing the three-dimensional bounding boxes from the annotation information in response to a preset condition being met by the positional relationship between the projection region and the two-dimensional bounding boxes; and training the object recognition model using the sample image and the removed annotation information, the object recognition model being configured to output the two-dimensional and three-dimensional positions of objects of a preset type.

[0004] In a second aspect of this disclosure, an apparatus for vehicle control is provided. The apparatus includes: an acquisition module configured to acquire annotation information corresponding to a target object in a sample image, the annotation information including a two-dimensional bounding box and a three-dimensional bounding box; a first determination module configured to determine a projection region in the sample image based on the three-dimensional bounding box; a removal module configured to remove the three-dimensional bounding box from the annotation information in response to a preset condition being met by the positional relationship between the projection region and the two-dimensional bounding box; and a training module configured to train an object recognition model using the sample image and the removed annotation information, the object recognition model being configured to output the two-dimensional and three-dimensional positions of a preset type of object.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, which, when executed by a processor, implements the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the method of the first aspect.

[0008] It should be understood that the description in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 A flowchart is shown illustrating a process for training an object recognition model according to some embodiments of the present disclosure;

[0012] Figure 3 A schematic structural block diagram of an apparatus for training an object recognition model according to some embodiments of the present disclosure is shown; and

[0013] Figure 4 A block diagram of an electronic device that can implement one or more embodiments of the present disclosure is shown. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0017] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0018] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0019] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably.

[0020] As mentioned earlier, a multi-task model that performs both 2D and 3D detection tasks can be trained simultaneously. From a training perspective, training only the 3D detection task can reduce the training difficulty of this multi-task model. However, in practical applications, downstream modules of autonomous driving often need to use more compact and accurate 2D bounding boxes to complete downstream understanding tasks, such as the analysis of taillights and traffic lights. Therefore, an effective framework for joint 2D and 3D training becomes an essential component.

[0021] Traditional joint training frameworks for 2D and 3D detection tasks fall into several modalities. One approach involves first fixing the network foundation for both 2D and 3D detection tasks. This foundation is pre-trained using semi-supervised or self-supervised methods, or distillation techniques, to extract the model's feature extraction capabilities. Then, each output head of the multi-task model is trained separately to achieve the processing capabilities for both 2D and 3D detection tasks. Another approach directly uses 2D and 3D data from the same source to jointly train the multi-task model. The advantage of this method is strong data consistency, leading to more balanced training and often a higher upper limit of model accuracy.

[0022] However, the first modality cannot use data for end-to-end training, which leads to a lower performance ceiling for each functional model block. The second modality has a high cost of fully labeled data, so the training samples are generally small. Moreover, in practical applications, the difference in timestamps between two-dimensional and three-dimensional data may lead to poor consistency in spatial relationships, which in turn affects model performance.

[0023] In view of this, embodiments of the present disclosure provide an improved scheme for training an object recognition model. In this scheme, annotation information corresponding to a target object in a sample image is obtained; the annotation information includes two-dimensional bounding boxes and three-dimensional bounding boxes; based on the three-dimensional bounding boxes, a projection region in the sample image is determined; in response to the positional relationship between the projection region and the two-dimensional bounding boxes satisfying a preset condition, the three-dimensional bounding boxes are removed from the annotation information; and using the sample image and the removed annotation information, an object recognition model is trained, the object recognition model being configured to output the two-dimensional and three-dimensional positions of objects of a preset type.

[0024] According to embodiments of this disclosure, 3D bounding boxes that meet specific conditions can be removed to optimize training data, ensuring that the object recognition model can focus on data that contributes more to training, effectively improving the accuracy of the object recognition model. Furthermore, this disclosure can integrate 3D and 2D bounding boxes to learn multiple detection tasks for the object recognition model, improving the training efficiency of the object recognition model.

[0025] Example Environment

[0026] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, environment 100 may contain an object recognition model 120 with parameter values ​​before training. The object recognition model 120 may be implemented or included in electronic device 110, or may be implemented or included in other devices.

[0027] exist Figure 1In environment 100, this object recognition model 120 can be a multi-task model, also known as a multi-task detection model. For example, the object recognition model can be used to detect the two-dimensional and three-dimensional locations corresponding to objects of a predetermined type included in an image.

[0028] Before training, the object recognition model 120 can have initial parameter values ​​or pre-trained parameter values ​​obtained through a pre-training process. The object recognition model 120 can be trained via forward and backward propagation, during which the parameter values ​​can be updated and adjusted. After training, a model that can be used for multi-task detection in the model application phase can be obtained.

[0029] During the model training phase, the object recognition model 120 can be trained using a training sample set 112 that includes multiple training samples 112 and a model training system.

[0030] exist Figure 1 In this context, electronic device 110 can include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices can involve any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0031] It should be understood that Figure 1 The components and arrangements shown in environment 100 are merely examples, and a computing system suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. Implementations of this disclosure are not limited in this respect.

[0032] Example process

[0033] Some exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0034] Figure 2 A flowchart of a process 200 for training an object recognition model according to some embodiments of the present disclosure is shown. Process 200 can be implemented at electronic device 110. Reference is made below. Figure 1 Describe the process 200.

[0035] In frame 210, electronic device 110 acquires annotation information corresponding to the target object in the sample image, including two-dimensional annotation frames and three-dimensional annotation frames.

[0036] In some embodiments, the sample images are image data used to train an object recognition model. The target object can be any suitable object contained in the sample images, such as vehicles, pedestrians, traffic signs, animals, etc.

[0037] In some embodiments, a two-dimensional bounding box can be used to indicate the size information of a target object in two-dimensional space (such as a sample image) and its positional information (such as width and height) on a plane. The shape of the two-dimensional bounding box can be set as needed. As an example, the two-dimensional bounding box can be a rectangle, and its coordinates can define the boundary of the target object in the sample image.

[0038] In some embodiments, the three-dimensional bounding box can be used to indicate the size and position information (such as width, height, and depth) of the target object in three-dimensional space, and the depth information can be the depth information of the target object relative to the camera or sensor.

[0039] In order to ensure that the object recognition model can obtain sufficient training samples for both two-dimensional and three-dimensional detection tasks, and to avoid the object recognition model being unable to accurately perform the detection task due to insufficient training samples, the electronic device 110 can sample corresponding sample images from the sample set based on different sampling probabilities.

[0040] Specifically, the electronic device 110 can determine a first number of first-class samples for two-dimensional detection tasks and a second number of second-class samples for three-dimensional detection tasks in the sample set.

[0041] Furthermore, the electronic device 110 can determine a first sampling probability of the first type of samples and a second sampling probability of the second type of samples based on a first number and a second number. In some embodiments, the first sampling probability is proportional to the first number, and the second sampling probability is proportional to the second number.

[0042] As an example, electronic device 110 can determine the first sampling probability and the second sampling probability based on the following formula:

[0043] P_taskA=sampeA / sampleAll+P_designA;

[0044] P_taskB=sampeB / sampleAll+P_designB;

[0045] Wherein, P_taskA is the first sampling probability of the first type of sample used for the 2D detection task; P_taskB is the second sampling probability of the second type of sample used for the 3D detection task; sampleAll is the total number of samples included in the sample set; sampleA is the number of first type samples included in the sample set; sampleB is the number of second type samples included in the sample set; P_designA is a preset probability parameter corresponding to the first type of sample, which represents the importance of the 2D detection task corresponding to the first type of sample to the object recognition model; P_designB is a preset probability parameter corresponding to the second type of sample, which represents the importance of the 3D detection task corresponding to the second type of sample to the object recognition model.

[0046] It should be noted that object recognition models can learn to perform not only 2D and 3D detection tasks, but also other detection tasks. In other words, the sample set can include not only first-class and second-class samples, but also other classes of samples, and the sum of the sampling probabilities for all classes of samples corresponding to the detection tasks is 1, i.e., P0. taskA +P taskB +P taskC +…+P taskN =1, where P taskC P taskN Equations represent the sampling probabilities for samples other than the first and second classes.

[0047] Furthermore, the electronic device 110 can sample multiple sample images corresponding to the target training period from the sample set based on a first sampling probability and a second sampling probability. Specifically, the electronic device 110 can randomly sample multiple sample images corresponding to the target training period from the sample set based on a first sampling frequency and a second sampling frequency, wherein these multiple sample images are sample images corresponding to the first type of sample and / or the second type of sample.

[0048] As an example, electronic device 110 can determine multiple sample images based on the following formula:

[0049] dataset_idx=random.choice((taskA,taskB,...),p=(P_taskA,P_taskB,...))

[0050] Where dataset_idx represents the selected sample images, P_taskA is the first sampling probability of the first type of sample used for the 2D detection task, and P_taskB is the second sampling probability of the second type of sample used for the 3D detection task; taskA represents the 2D detection task, and taskB represents the 3D detection task; random.choice((taskA, taskB, ...), p = (P taskA P taskB ...) represents the random sampling of sample images from the sample set according to the first sampling probability and the second sampling probability.

[0051] In frame 220, electronic device 110 determines the projection area in the sample image based on the three-dimensional annotation frame.

[0052] In some embodiments, the electronic device 110 can project a three-dimensional bounding box onto a 2D plane (also known as a pixel coordinate system) to determine the projection area in the sample image. Specifically, the electronic device 110 can convert the three-dimensional coordinates of the bounding box into two-dimensional image coordinates based on the coordinates of the points in the bounding box × extrinsic parameter matrix × intrinsic parameter matrix. Then, based on the obtained two-dimensional image coordinates corresponding to all points on the bounding box, the minimum bounding rectangle of all these points in the image can be determined. The extrinsic parameter matrix mainly represents the intrinsic properties of the camera or sensor, such as focal length and principal point, while the intrinsic parameter matrix mainly represents the position and orientation of the camera relative to the world coordinate system. Further, the electronic device 110 can use this minimum bounding rectangle to determine the projection area.

[0053] In frame 230, electronic device 110 removes the three-dimensional annotation box from the annotation information in response to the positional relationship between the projection area and the two-dimensional annotation box satisfying a preset condition.

[0054] To avoid misalignment between two-dimensional and three-dimensional data information and improve the accuracy of object recognition models, in some embodiments, the electronic device 110 can determine the center point distance between the projection area and the two-dimensional bounding box. Further, the electronic device 110 can determine the ratio of the center point distance to at least one side length of the two-dimensional bounding box; this ratio can provide information about the relative position of the projection area and the two-dimensional bounding box, particularly whether they are spatially aligned. Further, the electronic device 110 can determine whether the positional relationship between the projection area and the two-dimensional bounding box satisfies a preset condition based on the ratio. As an example, the electronic device 110 can determine whether the positional relationship between the projection area and the two-dimensional bounding box satisfies the preset condition based on a comparison of the ratio with a predetermined first threshold. Specifically, the electronic device 110 can determine that the positional relationship between the projection area and the two-dimensional bounding box does not satisfy the preset condition in response to determining that the ratio is less than or equal to the predetermined first threshold. The electronic device 110 can determine that the positional relationship between the projection area and the two-dimensional bounding box satisfies the preset condition in response to determining that the ratio is greater than the predetermined first threshold.

[0055] In other embodiments, the electronic device 110 can determine the intersection-over-union (IoU) ratio between the projected area and the two-dimensional annotation frame. Further, the electronic device 110 can determine whether the positional relationship between the projected area and the two-dimensional annotation frame satisfies a preset condition based on the IoU ratio. As an example, the electronic device 110 can determine whether the positional relationship between the projected area and the two-dimensional annotation frame satisfies the preset condition based on a comparison of the IoU ratio with a predetermined second threshold. Specifically, the electronic device 110 can determine that the positional relationship between the projected area and the two-dimensional annotation frame satisfies the preset condition in response to determining that the IoU ratio is less than the predetermined second threshold. The electronic device 110 can determine that the positional relationship between the projected area and the two-dimensional annotation frame does not satisfy the preset condition in response to determining that the IoU ratio is greater than or equal to the predetermined second threshold.

[0056] In other embodiments, the electronic device 110 may also determine whether the positional relationship between the projected area and the two-dimensional annotation frame satisfies a preset condition based on the intersection-union ratio and the ratio value. Specifically, the electronic device 110 may determine that the positional relationship between the projected area and the two-dimensional annotation frame does not satisfy the preset condition in response to determining that the intersection-union ratio is greater than or equal to a predetermined second threshold, and / or that the ratio value is less than or equal to a predetermined first threshold. The electronic device 110 may determine that the positional relationship between the projected area and the two-dimensional annotation frame satisfies the preset condition in response to determining that the intersection-union ratio is less than a predetermined second threshold and the ratio value is greater than a predetermined first threshold.

[0057] In some embodiments, if the positional relationship between the projection area and the two-dimensional bounding box meets the preset conditions, it indicates that there is a problem of poor alignment between the two-dimensional data information and the three-dimensional data information. In order to improve the accuracy of the object recognition model, the electronic device 110 can remove the three-dimensional bounding box from the annotation information, that is, the annotation information corresponding to the sample image only includes the two-dimensional bounding box. This sample image with the three-dimensional bounding box removed will only be applied to the object recognition model in the learning process of the two-dimensional detection task, and will not be applied to the object recognition model in the learning process of the three-dimensional detection task.

[0058] This disclosure performs the above-mentioned annotation information filtering operation (i.e., determining whether to remove the 3D annotation box) on each sample image participating in the training of the object recognition model. It can dynamically calculate the filtering information for each individual in each sample image, thereby realizing dynamic adaptive filtering of the full amount of data on a sample image basis.

[0059] In box 240, electronic device 110 uses sample images and removed annotation information to train an object recognition model, which is configured to output the two-dimensional and three-dimensional positions of objects of a preset type.

[0060] In some embodiments, the preset object can be any suitable object, such as an animal, a person, etc.

[0061] In some embodiments, the two-dimensional and three-dimensional positions are output by different output heads of the object recognition model. Different output heads correspond to different detection tasks; for example, the output of the two-dimensional position for a two-dimensional detection task can correspond to the first output head, and the output of the three-dimensional position for a three-dimensional detection task can correspond to the second output head.

[0062] In some embodiments, the electronic device 110 can compare the two-dimensional and three-dimensional positions with the two-dimensional and three-dimensional bounding boxes in the annotation information corresponding to the sample image, respectively. Based on the differences between the two-dimensional position and the two-dimensional bounding box, and the differences between the three-dimensional position and the three-dimensional bounding box, the object recognition model is trained until predetermined training completion conditions are met. The predetermined training completion conditions may be that the number of training iterations reaches a threshold, the training time reaches a threshold, the target loss value determined based on the output of the object recognition model and the differences in the annotation information reaches a minimum, etc., which will not be elaborated here.

[0063] In some embodiments, this object recognition model can be a multi-task detection model, which can learn the detection processes of multiple types of tasks. To improve the training efficiency and accuracy of the object recognition model, this disclosure can train the object recognition model based on a single-step, phased training strategy in the form of a configurable hook.

[0064] Specifically, the electronic device 110 can execute a first training process corresponding to the two-dimensional position output in the first stage of the target training cycle. Further, the electronic device 110 can execute a second training process corresponding to the three-dimensional position output in the second stage of the target training cycle, the second stage being later than the first stage and at least partially overlapping with the first stage. As an example, the start time of the second stage can correspond to the convergence time of the two-dimensional detection task in the first stage, and this convergence time is earlier than the end time of the first stage.

[0065] For example, this target training cycle can correspond to {taskA: 0-20, taskB: 10-30, taskc...}, where taskA can represent a two-dimensional detection task, taskB can represent a three-dimensional detection task, and taskc can represent other types of tasks besides two-dimensional and three-dimensional detection tasks. 0-20 corresponds to the first stage, and 10-30 corresponds to the second stage. The electronic device 110 can place the training of the two-dimensional detection task in the early warm-up stage (i.e., stage 0-20), while the training of the three-dimensional detection task is placed in the stage after the convergence of the two-dimensional detection task (corresponding to stage 10-30), where 10 is the convergence time corresponding to the two-dimensional detection task.

[0066] It should be noted that the training phase for each type of detection task can be set according to requirements. Specifically, the training phase for each task can be set based on training convergence experience and training information relationships.

[0067] It should be noted that the data sampling algorithm based on data distribution and the configurable hook-based single-step phased training strategy disclosed herein are applicable to any multi-task training framework in practice and can be used as components individually or in combination.

[0068] According to embodiments of this disclosure, 3D bounding boxes that meet specific conditions can be removed to optimize training data, ensuring that the object recognition model can focus on data that contributes more to training, effectively improving the accuracy of the object recognition model. Furthermore, this disclosure can integrate 3D and 2D bounding boxes to learn multiple detection tasks for the object recognition model, improving the training efficiency of the object recognition model.

[0069] Example devices, vehicles, and equipment

[0070] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 3A schematic structural block diagram of an apparatus 300 for training an object recognition model according to some embodiments of the present disclosure is shown. The apparatus 300 may be implemented in or included in an electronic device 110. Various modules / components in the apparatus 300 may be implemented by hardware, software, firmware, or any combination thereof.

[0071] like Figure 3 As shown, the device 300 includes an acquisition module 310 configured to acquire annotation information corresponding to a target object in a sample image, the annotation information including a two-dimensional annotation box and a three-dimensional annotation box; a first determination module 320 configured to determine a projection region in the sample image based on the three-dimensional annotation box; a removal module 330 configured to remove the three-dimensional annotation box from the annotation information in response to a preset condition being met by the positional relationship between the projection region and the two-dimensional annotation box; and a training module 340 configured to train an object recognition model using the sample image and the removed annotation information, the object recognition model being configured to output the two-dimensional and three-dimensional positions of a preset type of object.

[0072] In some embodiments, the device 300 further includes a second determining module configured to: determine the center point distance between the projection area and the two-dimensional annotation box; a third determining module configured to: determine the ratio of the center point distance to at least one side length of the two-dimensional annotation box; and a fourth determining module configured to: determine, based on the ratio, whether the positional relationship between the projection area and the two-dimensional annotation box satisfies a preset condition.

[0073] In some embodiments, the device 300 further includes a fifth determining module configured to: determine the intersection-union ratio between the projection area and the two-dimensional annotation box; and a sixth determining module configured to: determine, based on the intersection-union ratio, whether the positional relationship between the projection area and the two-dimensional annotation box meets a preset condition.

[0074] In some embodiments, the apparatus 300 further includes a seventh determining module configured to: determine a first number of first-class samples for a two-dimensional detection task and a second number of second-class samples for a three-dimensional detection task in the sample set; an eighth determining module configured to: determine a first sampling probability of the first-class samples and a second sampling probability of the second-class samples based on the first number and the second number; and a sampling module configured to: sample multiple sample images corresponding to the target training period from the sample set based on the first sampling probability and the second sampling probability.

[0075] In some embodiments, the first sampling probability is proportional to the first number, and the second sampling probability is proportional to the second number.

[0076] In some embodiments, training the object recognition model includes: performing a first training process corresponding to a two-dimensional position output in a first phase of a target training cycle; and performing a second training process corresponding to a three-dimensional position output in a second phase of the target training cycle, wherein the second phase is later than the first phase and the second phase at least partially overlaps with the first phase.

[0077] In some embodiments, the two-dimensional position and the three-dimensional position are output by different output heads of the object recognition model.

[0078] The units and / or modules included in device 300 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 300 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0079] Figure 4 A block diagram of an electronic device 400 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 4 The electronic device 400 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.

[0080] like Figure 4 As shown, electronic device 400 is in the form of a general-purpose computing device. Components of electronic device 400 may include, but are not limited to, one or more processors or processing units 410, memory 420, storage device 430, one or more communication units 440, one or more input devices 450, and one or more output devices 460. Processing unit 410 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 420. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 400.

[0081] Electronic device 400 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 400, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 420 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 430 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 400.

[0082] Electronic device 400 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 4 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 420 may include computer program product 425 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0083] Communication unit 440 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 400 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 400 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0084] Input device 450 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 460 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 400 can also communicate with one or more external devices (not shown) via communication unit 440 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 400, or with any device that enables electronic device 400 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0085] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transient computer-readable medium and includes computer-executable instructions that are executed by a processor to implement the methods described above.

[0086] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0087] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0088] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0090] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the implementations disclosed herein.

Claims

1. A method for training an object recognition model, comprising: Obtain annotation information corresponding to the target object in the sample image, wherein the annotation information includes two-dimensional annotation boxes and three-dimensional annotation boxes; Based on the three-dimensional bounding box, the projection region in the sample image is determined; In response to the positional relationship between the projection area and the two-dimensional annotation box satisfying a preset condition, the three-dimensional annotation box is removed from the annotation information; as well as Using the sample images and the removed annotation information, an object recognition model is trained. The object recognition model is configured to output the two-dimensional and three-dimensional positions of objects of a preset type.

2. The method according to claim 1, further comprising: Determine the center distance between the projection area and the two-dimensional annotation frame; Determine the ratio of the center point distance to at least one side length of the two-dimensional annotation box; as well as Based on the ratio, determine whether the positional relationship between the projection area and the two-dimensional annotation box satisfies the preset condition.

3. The method according to claim 1, further comprising: Determine the intersection-union ratio between the projected area and the two-dimensional annotation frame; as well as Based on the intersection-union ratio, determine whether the positional relationship between the projection area and the two-dimensional annotation box satisfies the preset condition.

4. The method according to claim 1, further comprising: Determine the first number of first-class samples used for two-dimensional detection tasks and the second number of second-class samples used for three-dimensional detection tasks in the sample set; Based on the first number and the second number, determine the first sampling probability of the first type of sample and the second sampling probability of the second type of sample; as well as Based on the first sampling probability and the second sampling probability, multiple sample images corresponding to the target training period are sampled from the sample set.

5. The method of claim 4, wherein the first sampling probability is proportional to the first number, and the second sampling probability is proportional to the second number.

6. The method according to claim 4, wherein training the object recognition model comprises: In the first stage of the target training cycle, a first training process corresponding to the two-dimensional position output is executed; as well as In the second phase of the target training cycle, a second training process corresponding to the three-dimensional position output is performed. The second phase is later than the first phase and at least partially overlaps with the first phase.

7. The method according to claim 1, wherein the two-dimensional position and the three-dimensional position are output by different output heads of the object recognition model.

8. An apparatus for training an object recognition model, comprising: The acquisition module is configured to acquire annotation information corresponding to the target object in the sample image, the annotation information including two-dimensional annotation boxes and three-dimensional annotation boxes; The first determining module is configured to determine the projection region in the sample image based on the three-dimensional annotation box; The removal module is configured to remove the three-dimensional annotation box from the annotation information in response to the positional relationship between the projection area and the two-dimensional annotation box satisfying a preset condition; as well as The training module is configured to train an object recognition model using the sample images and the removed annotation information, the object recognition model being configured to output the two-dimensional and three-dimensional positions of objects of a preset type.

9. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 7.