Method, apparatus, device, storage medium and program product for object recognition

By combining point cloud and image feature detection models, the problem of low object recognition accuracy in existing technologies has been solved, achieving more efficient and accurate object recognition.

CN122223673APending Publication Date: 2026-06-16BEIJING VOYAGER TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING VOYAGER TECH CO LTD
Filing Date
2024-12-13
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing object recognition methods have low accuracy in autonomous and assisted driving, making it difficult to guarantee the accuracy and efficiency of object recognition results generated from perception data.

Method used

The first detection model is used to process point cloud features to determine the first set of candidate regions, and the second detection model is used to process image features to determine the second set of candidate regions. Query features are then initialized based on these candidate regions, and finally, object recognition results are generated.

Benefits of technology

By combining point cloud and image feature detection models, the accuracy and efficiency of object recognition are improved, generating more accurate object recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223673A_ABST
    Figure CN122223673A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method, apparatus, device, storage medium and program product for object recognition. The method comprises: processing point cloud features by using a first detection model to determine a first set of candidate regions, and processing image features by using a second detection model to determine a second set of candidate regions; initializing a first set of query features based on the first set of candidate regions, and initializing a second set of query features based on the second set of candidate regions, wherein the object categories corresponding to the first set of query features are determined based on the first detection model, and the object categories corresponding to the second set of query features are determined based on the second detection model; and generating an object recognition result based on at least the first set of query features and the second set of query features. In this way, embodiments of the present disclosure can improve the accuracy and efficiency of recognizing objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and particularly to methods, apparatus, devices, computer-readable storage media, and computer program products for object identification. Background Technology

[0002] With the development of computer technology, autonomous driving and driver assistance technologies have emerged to reduce the demands on drivers and free them from certain driving responsibilities. Typically, vehicles can assist drivers by using perception data collected by onboard sensors and other sensing devices, or they can use control units to process the perception data and directly control the vehicle. Therefore, ensuring the accuracy and efficiency of object recognition results generated based on perception data is a crucial concern. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for object recognition is provided. The method includes: processing point cloud features using a first detection model to determine a first set of candidate regions, and processing image features using a second detection model to determine a second set of candidate regions; initializing a first set of query features based on the first set of candidate regions, and initializing a second set of query features based on the second set of candidate regions, wherein the object category corresponding to the first set of query features is determined based on the first detection model, and the object category corresponding to the second set of query features is determined based on the second detection model; and generating an object recognition result based at least on the first set of query features and the second set of query features.

[0004] In a second aspect of this disclosure, an apparatus for object recognition is provided. The apparatus includes: a region determination module configured to process point cloud features using a first detection model to determine a first set of candidate regions, and to process image features using a second detection model to determine a second set of candidate regions; a feature determination module configured to initialize a first set of query features based on the first set of candidate regions, and to initialize a second set of query features based on the second set of candidate regions, wherein the object category corresponding to the first set of query features is determined based on the first detection model, and the object category corresponding to the second set of query features is determined based on the second detection model; and a generation module configured to generate an object recognition result based at least on the first set of query features and the second set of query features.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.

[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram is shown in which an example system in which embodiments of the present disclosure may be implemented;

[0011] Figure 2 A schematic diagram illustrating an example process for object recognition according to some embodiments of the present disclosure is shown;

[0012] Figure 3 A schematic structural block diagram of an example apparatus for object recognition according to some embodiments of the present disclosure is shown; and

[0013] Figure 4 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0017] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0018] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.

[0019] As mentioned above, with the development of computer technology, autonomous driving and driver assistance technologies have emerged to reduce the demands on drivers and free them from certain driving responsibilities. Typically, vehicles can assist drivers by using perception data collected by onboard sensors or other sensing devices, or they can use control units to process the perception data and directly control the vehicle. Therefore, ensuring the accuracy and efficiency of object recognition results generated from perception data is crucial. However, existing object recognition methods currently suffer from low accuracy.

[0020] The embodiments of this disclosure propose an object recognition scheme. This scheme utilizes a first detection model to process point cloud features to determine a first set of candidate regions, and a second detection model to process image features to determine a second set of candidate regions; it initializes a first set of query features based on the first set of candidate regions, and initializes a second set of query features based on the second set of candidate regions, wherein the object category corresponding to the first set of query features is determined based on the first detection model, and the object category corresponding to the second set of query features is determined based on the second detection model; and it generates an object recognition result based at least on the first set of query features and the second set of query features.

[0021] In this manner, embodiments of the present disclosure can utilize a detection model to determine corresponding candidate regions for point cloud features and image features respectively, and initialize corresponding query features based on the candidate regions. Furthermore, embodiments of the present disclosure can generate object recognition results based on the initialized query features. Thus, embodiments of the present disclosure can initialize query features based on a detection model during the object recognition process, thereby improving the accuracy and efficiency of object recognition.

[0022] Example object identification

[0023] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0024] First see Figure 1 The illustration shows an example system 100 for training a recognition model that may be implemented according to embodiments of the present disclosure.

[0025] like Figure 1 As shown, system 100 may include a point cloud feature processing unit 110, an image feature processing unit 120, a query feature generation unit 130, and a detection unit 140. System 100 may be configured in a vehicle (e.g., an autonomous vehicle with autonomous driving capabilities) for processing road condition information.

[0026] System 100 can use configured sensors such as cameras, radar (e.g., lidar, millimeter-wave radar) to acquire sensing information about the physical environment in which the vehicle is located.

[0027] For example, system 100 can acquire image information (e.g., image data) of the physical environment in which the vehicle is located based on cameras. For example, system 100 can acquire multi-view images of the physical environment in which the vehicle is located based on multiple cameras.

[0028] System 100 can acquire point cloud information (e.g., point cloud data) of the physical environment in which the vehicle is located based on radar. For example, System 100 can acquire multi-view point cloud information of the physical environment in which the vehicle is located based on multiple radars. As an example, point cloud data is a data structure that represents a set of points in three-dimensional space. Each point contains its coordinates in three-dimensional space (usually X, Y, and Z coordinates), and sometimes additional information such as color, intensity, and normal. Point cloud data can be used to represent entities in three-dimensional space such as the surface shape of objects, terrain, and buildings.

[0029] The point cloud feature processing unit 110 can determine the feature information (e.g., point cloud feature map) corresponding to the received point cloud data (e.g., multi-scale point cloud data). As an example, the feature information can be multi-scale bird's-eye view point cloud features (e.g., point cloud features projected onto a bird's-eye view plane). As another example, the point cloud feature processing unit 110 can extract the point cloud feature map corresponding to the point cloud data based on a trained backbone network (e.g., pointpollars, voxelnet, etc.) and / or neck network used for extracting point cloud features.

[0030] As an example, the point cloud feature processing unit 110 can utilize a first detection model to determine a first group of candidate regions based on the feature information corresponding to the point cloud data. The point cloud feature processing unit 110 can determine the center point location, depth, and object category corresponding to each candidate region in the first group of candidate regions. As an example, the first detection model can include a dense Bird's Eye View (BEV) detection head. For example, a dense BEV detection head can be implemented as a centerpoint model. As an example, the first detection model can be implemented as a lightweight model with few parameters.

[0031] The image feature processing unit 120 can determine feature information (e.g., image feature map) corresponding to the received image data (e.g., multi-scale multi-camera image data). As an example, the image feature extraction unit 110 can extract the image feature map corresponding to the image data based on a trained backbone network (e.g., ResNet, Deep Aggregator Network (DLA), etc.) and / or a neck network for extracting image features. Exemplarily, the neck network can be implemented as a feature pyramid network, etc.

[0032] As an example, the image feature processing unit 120 can utilize a second detection model to determine a second set of candidate regions based on feature information corresponding to the image data. The image feature processing unit 120 can determine the center point position, depth, and object category corresponding to each candidate region in the second set of candidate regions. As an example, the second detection model can include a dense two-dimensional detection head. For example, a dense two-dimensional detection head can be implemented as a series of "You Only LookOnce" (YOLO) algorithms (e.g., YOLOv2, YOLOv3, YOLOv4, etc.). As an example, the second detection model can be implemented as a lightweight model with few parameters.

[0033] The query feature generation unit 130 can initialize point cloud query features (or the first set of query features) based on the first set of candidate regions determined by the point cloud feature processing unit 110. The query feature generation unit 130 can initialize image query features (or the second set of query features) based on the second set of candidate regions determined by the image feature processing unit 120.

[0034] Alternatively, the query feature generation unit 130 may also randomly initialize at least one additional query feature.

[0035] The detection unit 140 can generate object recognition results based on point cloud query features, image query features and / or at least one additional query feature.

[0036] The following will detail the specific implementation of the point cloud feature processing unit 110, image feature processing unit 120, query feature generation unit 130, and detection unit 140 in system 100.

[0037] Figure 2 A flowchart of an example process 200 for object identification according to some embodiments of the present disclosure is shown. Process 200 can be implemented in system 100. Reference is made below. Figure 1 Describe the process 200.

[0038] like Figure 2 As shown in box 210, the point cloud feature processing unit 110 processes point cloud features using a first detection model to determine a first set of candidate regions. The image feature processing unit 120 processes image features using a second detection model to determine a second set of candidate regions.

[0039] As an example, the first detection model may include a dense bird's eye view (BEV) detection head. For example, a dense BEV detection head may be implemented as a centerpoint model. The point cloud feature processing unit 110 may use the first detection model to determine the centerpoint location, depth, and object category corresponding to each candidate region in the first set of candidate regions.

[0040] As an example, the second detection model may include a dense two-dimensional detection head. For instance, the dense two-dimensional detection head may be implemented as a series of algorithms called "You Only Look Once" (YOLO) (e.g., YOLOv2, YOLOv3, YOLOv4, etc.). The point cloud feature processing unit 110 may use the second detection model to determine the center point location, depth, and object category corresponding to each candidate region in the second set of candidate regions.

[0041] As an example, the object categories corresponding to the candidate regions may include, but are not limited to: vehicles (e.g., cars, trucks, etc.), pedestrians, lane lines, fences (or barriers), road obstacles (e.g., long-tail objects such as plastic bags and small animals), and road signs (e.g., signs, traffic lights, etc.).

[0042] In some embodiments, the first set of candidate regions or the second set of candidate regions includes at least one two-dimensional region. As an example, the first set of candidate regions may include at least one two-dimensional region corresponding to a bird's-eye view plane. As an example, the second set of candidate regions may include at least one two-dimensional region corresponding to a camera two-dimensional plane.

[0043] As an example, the first set of candidate regions and / or the second set of candidate regions can also be referred to as a set of target proposals. Each proposal in a set of target proposals only needs to include the center point location and depth of the target (or object), as well as the target category, and does not need to regress specific 3D bounding box information. As an example, a set of target proposals can be determined based on a lightweight model with few parameters (e.g., fewer than a threshold).

[0044] In box 220, the query feature generation unit 130 initializes a first set of query features based on a first set of candidate regions, and initializes a second set of query features based on a second set of candidate regions. The object category corresponding to the first set of query features is determined based on a first detection model. The object category corresponding to the second set of query features is determined based on a second detection model.

[0045] As an example, the query feature generation unit 130 can determine the first set of features corresponding to the first set of candidate regions from the point cloud features (or BEV feature map). Taking the first candidate region in the first set of candidate regions as an example, the query feature generation unit 130 can interpolate a set of feature values ​​related to the first center point position based on the first center point position corresponding to the first candidate region in the BEV feature map to determine the feature vector of the first feature corresponding to the first candidate region. Further, the query feature generation unit 130 can construct the first set of query features based on the first set of features. As an example, the reference point position of the first set of query features can be the center point position of the first set of features.

[0046] As an example, the query feature generation unit 130 can determine a second set of features corresponding to a second set of candidate regions from image features (or a two-dimensional feature map). Taking a second candidate region within the second set of candidate regions as an example, the query feature generation unit 130 can interpolate a set of feature values ​​related to the second center point position based on the second center point position corresponding to the second candidate region in the two-dimensional feature map to determine the feature vector of the second feature corresponding to the second candidate region. Further, the query feature generation unit 130 can construct a second set of query features based on the second set of features. As an example, the reference point position of the second set of query features can be a three-dimensional position calculated based on the center point position and depth of the second set of features.

[0047] Alternatively or additionally, the first set of query features or the second set of query features may include the target query features.

[0048] As an example, the query feature generation unit 130 can determine the size of the 3D anchor box (e.g., 3D anchor) corresponding to the target query feature based on the object category corresponding to the target query feature. As an example, the 3D anchor box can serve as a predefined (e.g., based on manual settings or clustering based on training set annotations) reference box in 3D object detection for object localization and identification. The position, size, and / or orientation of the object are predicted by matching the 3D anchor box with the true 3D bounding box (e.g., calculating the intersection-union ratio), thereby improving the accuracy and efficiency of detection. As an example, the size of the 3D anchor box corresponding to different object categories can be different. For example, the size of the 3D anchor box corresponding to the object category of vehicle can be larger than the size of the 3D anchor box corresponding to the object category of pedestrian.

[0049] As an example, the query feature generation unit 130 can apply self-attention and / or deformable cross-attention mechanisms to the first set of query features and / or the second set of query features to update the first set of query features and / or the second set of query features.

[0050] As an example, the query feature generation unit 130 can determine the number of sampling points (also known as the number of feature sampling points) corresponding to the target query feature based on the object category corresponding to the target query feature. As an example, the number of sampling points can be the number of feature sampling points in a deformable cross-attention mechanism. As an example, the number of sampling points corresponding to different object categories can be different. As an example, for object categories with larger sizes and / or complex shapes, the query feature generation unit 130 can set more sampling points. As an example, for object categories with smaller sizes and / or simpler shapes, the query feature generation unit 130 can set fewer sampling points. For example, the number of sampling points corresponding to the object category "vehicle" can be greater than the number of sampling points corresponding to the object category "pedestrian".

[0051] As an example, the query feature generation unit 130 can determine different numbers of the first set of query features and / or the second set of query features based on different object categories. In this way, using different numbers of the first set of query features and / or the second set of query features for object recognition for different object categories can improve recognition accuracy and efficiency. For example, when the object category is lane lines, a higher number of the second set of query features associated with image features than the number of the first set of query features associated with point cloud features is more conducive to successful recognition.

[0052] In some embodiments, the query feature generation unit 130 may determine a first number of first group query features based on preset configuration information. For example, the query feature generation unit 130 may determine the first number of first group query features based on the object category corresponding to the first group of query features. For instance, the preset configuration information may indicate the first number of first group query features corresponding to an object category. For instance, different object categories may correspond to different first numbers of first group query features. For instance, if the first group of query features is associated with object category A (e.g., vehicles), the query feature generation unit 130 may determine the first number of first group query features as number A based on the preset configuration information. For instance, if the first group of query features is associated with object category B (e.g., pedestrians), the query feature generation unit 130 may determine the first number of first group query features as number B based on the preset configuration information. Number A may be different from number B. For instance, if the first group of query features is associated with object category B (e.g., pedestrians), the query feature generation unit 130 may determine k (e.g., number B) query features from the first group of candidate query features corresponding to the first group of candidate regions based on a top k algorithm. Similarly, the query feature generation unit 130 can determine the second number of the second set of query features based on preset configuration information. The second number may be different from the first number.

[0053] Alternatively, the query feature generation unit 130 may also randomly initialize at least one additional query feature. As an example, randomly initialized at least one additional query feature can improve model recall.

[0054] For example, the query feature generation unit 130 can assign at least one additional query feature to multiple object categories based on the number of multiple object categories in the first set of query features and / or the second set of query features. For example, the multiple object categories may include object category A and object category B. The query feature generation unit 130 can assign at least one additional query feature to object category A and object category B in a 1:3 ratio based on the ratio of the number of object category A to the number of object category B being 1:3.

[0055] As an example, the distribution of at least one randomly initialized additional query feature can be determined based on prior information associated with a preset object category. For example, the prior information can be determined based on point cloud features and / or image features. For example, the prior information may include map information. The map information may, for example, indicate that a first region is associated with a lane and a second region is associated with a sidewalk. Further, the query feature generation unit 130 may set the number of additional query features associated with the vehicle category to be greater than the number of additional query features associated with the pedestrian category in the first region. Conversely, the query feature generation unit 130 may set the number of additional query features associated with the vehicle category to be less than the number of additional query features associated with the pedestrian category in the second region.

[0056] As an example, the query feature generation unit 130 can set different perception ranges for additional query features based on different object categories. For example, for object categories such as vehicles or roadside buildings, the query feature generation unit 130 can set a first perception range for the additional query features. For object categories such as traffic lights, the query feature generation unit 130 can set a second perception range for the additional query features. For example, since vehicles or roadside buildings are more important for maintaining safe vehicle distances, while traffic lights rely more on the details provided by image data, the first perception range can be larger than the second perception range.

[0057] In box 230, detection unit 140 generates object recognition results based on at least the first set of query features and the second set of query features.

[0058] Alternatively or additionally, the detection unit 140 may also generate object recognition results based on at least one additional query feature.

[0059] As an example, detection unit 140 can use a detection head to generate object recognition results based on a first set of query features, a second set of query features, at least one additional query feature, a point cloud feature map, and / or an image feature map. As an example, the detection head can be, for example, a sparse detection head. For example, detection unit 140 can generate object recognition results based on a Transformer decoder. As an example, this process can be implemented based on at least one of DETR3D, PETR, StreamPERT, FUTR3D, or SparseLIF.

[0060] For example, object recognition results can indicate a three-dimensional region (e.g., a bounding box) of at least one object. The bounding box may include, for example, the coordinates of the object's center point, width, and / or height.

[0061] As an example, object recognition results may include type information of objects in the target scene. Type information may indicate object categories, for example. Object categories may include, but are not limited to: vehicles (e.g., cars, trucks, etc.), pedestrians, lane lines, fences (or barriers), road obstacles (e.g., long-tailed objects such as plastic bags and small animals), and road signs (e.g., signs, traffic lights, etc.).

[0062] Based on the process described above, in this manner, embodiments of the present disclosure can utilize a detection model to determine corresponding candidate regions for point cloud features and image features respectively, and initialize corresponding query features based on the candidate regions. Furthermore, embodiments of the present disclosure can generate object recognition results based on the initialized query features and additional query features. Thus, in the process of object recognition, embodiments of the present disclosure can determine object recognition results based on additional query features and initial query features generated based on the detection model, thereby improving the accuracy and efficiency of object recognition.

[0063] Example devices and equipment

[0064] Figure 3 A schematic structural block diagram of an example apparatus 300 for training a recognition model according to some embodiments of the present disclosure is shown. Apparatus 300 may be implemented as or included in system 100. Various modules / components in apparatus 300 may be implemented by hardware, software, firmware, or any combination thereof.

[0065] like Figure 3As shown, the device 300 includes a region determination module 310, configured to process point cloud features using a first detection model to determine a first set of candidate regions, and process image features using a second detection model to determine a second set of candidate regions; a feature determination module 320, configured to initialize a first set of query features based on the first set of candidate regions, and initialize a second set of query features based on the second set of candidate regions, wherein the object category corresponding to the first set of query features is determined based on the first detection model, and the object category corresponding to the second set of query features is determined based on the second detection model; and a generation module 330, configured to generate an object recognition result based at least on the first set of query features and the second set of query features.

[0066] In some embodiments, the feature determination module 320 is further configured to: determine a first set of features corresponding to a first set of candidate regions from point cloud features; and construct a first set of query features based on the first set of features.

[0067] In some embodiments, the feature determination module 320 is further configured to: initialize a second set of query features based on a second set of candidate regions, including: determining a second set of features corresponding to the second set of candidate regions from image features; and constructing a second set of query features based on the second set of features.

[0068] In some embodiments, the first set of candidate regions or the second set of candidate regions includes at least one two-dimensional region.

[0069] In some embodiments, the first set of query features or the second set of query features includes target query features, and the size and / or number of sampling points of the three-dimensional anchor box corresponding to the target query features are determined based on the object category corresponding to the target query features.

[0070] In some embodiments, the first number of the first set of query features and / or the second number of the second set of query features are determined based on preset configuration information.

[0071] In some embodiments, the object identification result is also determined based on at least one additional query feature that is randomly initialized.

[0072] In some embodiments, the distribution of at least one additional query feature is determined based on prior information associated with a preset object category.

[0073] In some embodiments, the object identification result indicates at least one of the following: a three-dimensional region of at least one object; type information of at least one object.

[0074] The modules included in device 300 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the modules in device 300 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0075] Figure 4 A block diagram of an electronic device 400 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 4 The electronic device 400 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 400 can be used to implement... Figure 1 System 100.

[0076] like Figure 4 As shown, electronic device 400 is in the form of a general-purpose electronic device. Components of electronic device 400 may include, but are not limited to, one or more processors or processing units 410, memory 420, storage device 430, one or more communication units 440, one or more input devices 450, and one or more output devices 460. Processing unit 410 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 420. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 400.

[0077] Electronic device 400 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 400, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 420 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 430 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 400.

[0078] Electronic device 400 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 4 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 420 may include computer program product 425 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0079] Communication unit 440 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 400 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 400 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0080] Input device 450 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 450 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 400 can also communicate with one or more external devices (not shown) via communication unit 440 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 400, or with any device that enables electronic device 400 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0081] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the methods described above.

[0082] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0083] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0084] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0086] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the implementations disclosed herein.

Claims

1. A method for object recognition, comprising: The first detection model is used to process point cloud features to determine the first set of candidate regions, and the second detection model is used to process image features to determine the second set of candidate regions. A first set of query features is initialized based on the first set of candidate regions, and a second set of query features is initialized based on the second set of candidate regions. The object category corresponding to the first set of query features is determined based on the first detection model, and the object category corresponding to the second set of query features is determined based on the second detection model. as well as Based at least on the first set of query features and the second set of query features, an object recognition result is generated.

2. The method according to claim 1, wherein initializing the first set of query features based on the first set of candidate regions includes: From the point cloud features, determine the first set of features corresponding to the first set of candidate regions; as well as Based on the first set of features, construct the first set of query features.

3. The method according to claim 1, wherein initializing the second set of query features based on the second set of candidate regions includes: From the image features, determine a second set of features corresponding to the second set of candidate regions; as well as Based on the second set of features, construct the second set of query features.

4. The method according to claim 1, wherein the first group of candidate regions or the second group of candidate regions includes at least one two-dimensional region.

5. The method according to claim 1, wherein the first set of query features or the second set of query features includes target query features, and the size and / or number of sampling points of the three-dimensional anchor frame corresponding to the target query features is determined based on the object category corresponding to the target query features.

6. The method according to claim 1, wherein the first number of the first group of query features and / or the second number of the second group of query features are determined based on preset configuration information.

7. The method of claim 1, wherein the object identification result is further determined based on at least one additional query feature randomly initialized.

8. The method of claim 7, wherein the distribution of the at least one additional query feature is determined based on prior information associated with a preset object category.

9. The method of claim 1, wherein the object identification result indicates at least one of the following: At least one three-dimensional region of an object; The type information of at least one object.

10. An apparatus for object recognition, comprising: The region determination module is configured to process point cloud features using a first detection model to determine a first set of candidate regions, and to process image features using a second detection model to determine a second set of candidate regions. The feature determination module is configured to initialize a first set of query features based on the first set of candidate regions and initialize a second set of query features based on the second set of candidate regions. The object category corresponding to the first set of query features is determined based on the first detection model, and the object category corresponding to the second set of query features is determined based on the second detection model. as well as The generation module is configured to generate object recognition results based at least on the first set of query features and the second set of query features.

11. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.