Methods, apparatus, equipment, storage media and program products for object detection
By identifying the target area in the perceived image and projecting point cloud data, the problem of poor detection of small environmental obstacles in traditional technologies is solved, achieving accurate identification of the vehicle's surrounding environment and improving safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING VOYAGER TECH CO LTD
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-26
AI Technical Summary
Traditional technologies are less effective at detecting small, irregularly shaped, and small environmental obstacles around vehicles, making it difficult to guarantee vehicle safety.
By identifying the target region from the perceived image, projecting point cloud data into a two-dimensional space, determining the point set based on the point cloud distribution and depth information, and recognizing the shape and position of the object in three-dimensional space.
It improves the accuracy of detecting smaller environmental obstacles, ensuring the safety of vehicle operation.
Smart Images

Figure CN122090410A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, computer-readable storage media, and computer program products for object detection. Background Technology
[0002] With the continuous advancement of computer technology, computers can now replace or assist human drivers in perceiving the vehicle's surroundings, planning its trajectory, controlling its journey to a designated destination, and controlling its speed. For example, computers can analyze perceived environmental information to prevent problems such as collisions. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for object detection is provided. The method includes: determining a target region corresponding to an object of a target type from a perceived image; projecting point cloud data onto a two-dimensional space corresponding to the perceived image to determine a point cloud distribution associated with the target region; determining depth information corresponding to the target region based on the point cloud distribution; determining a set of points corresponding to the target region from the point cloud data based on the depth information and a search range corresponding to the target type; and determining an object recognition result based on the point set, the recognition result indicating the shape and / or position of the object in three-dimensional space.
[0004] In a second aspect of this disclosure, an apparatus for object detection is provided. The apparatus includes: a first determining module configured to determine a target region corresponding to an object of a target type from a perceived image; a projection module configured to project point cloud data onto a two-dimensional space corresponding to the perceived image to determine a point cloud distribution associated with the target region; a second determining module configured to determine depth information corresponding to the target region based on the point cloud distribution; a third determining module configured to determine a set of points corresponding to the target region from the point cloud data based on the depth information and a search range corresponding to the target type; and a fourth determining module configured to determine an object recognition result based on the point set, the recognition result indicating the shape and / or position of the object in three-dimensional space.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.
[0008] It should be understood that the content described in this summary section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0011] Figure 2 A schematic diagram of an example process for object detection according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A schematic diagram illustrating an example process of two-dimensional image processing according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic diagram illustrating an example multi-camera view fusion process according to some embodiments of the present disclosure is shown;
[0014] Figure 5 A schematic diagram illustrating an example process of two-dimensional image processing according to some embodiments of the present disclosure is shown;
[0015] Figure 6 A schematic structural block diagram of an example apparatus for object detection according to certain embodiments of the present disclosure is shown; and
[0016] Figure 7 A block diagram of an apparatus capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0020] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0021] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0022] As briefly mentioned earlier, computers can sense the vehicle's surrounding environment and obtain environmental information for analysis. This environmental information includes, for example, perceived images, point cloud data, and high-precision map data. Based on the analysis results, computers can provide guidance on vehicle operation, such as deceleration and braking, thereby preventing collisions or other events that could affect vehicle safety between the vehicle and objects / vehicles / pedestrians in the environment.
[0023] However, traditional technologies are less effective at detecting smaller environmental obstacles such as traffic cones, small stones on the road, small branches, and puddles, due to their irregular three-dimensional shapes, small size, and random angles. In such cases, vehicle safety cannot be adequately guaranteed.
[0024] Based on this, embodiments of this disclosure propose an object detection scheme. According to this scheme, a target region corresponding to an object of a target type can be determined from a perceived image; further, point cloud data can be projected onto a two-dimensional space corresponding to the perceived image to determine the point cloud distribution associated with the target region; further, depth information corresponding to the target region can be determined based on the point cloud distribution; further, a set of points corresponding to the target region can be determined from the point cloud data based on the depth information and a search range corresponding to the target type; further, the object recognition result can be determined based on the point set, the recognition result indicating the shape and / or position of the object in three-dimensional space.
[0025] In this way, embodiments of this disclosure can determine the target region corresponding to the target type of object, thereby determining the point cloud distribution associated with the target region based on point cloud data, reducing the interference of a large amount of point cloud data on the detection of the target type of object; furthermore, embodiments of this disclosure obtain the depth information corresponding to the target region through the point cloud distribution, ensuring the accuracy of the object's position in three-dimensional space; thus, based on such depth information and the search range corresponding to the target type, this disclosure searches for the point set corresponding to the target region, ensuring the accuracy of the point set associated with the object, thereby improving the accuracy of the object recognition result determined based on the point set.
[0026] Therefore, embodiments of this disclosure can obtain a more accurate position and / or shape of a target type object in three-dimensional space based on point cloud data and depth information, according to the target region. Furthermore, by utilizing the three-dimensional shape and / or position of such objects, embodiments of this disclosure can help improve the safety of vehicle operation.
[0027] Example Environment
[0028] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. As shown, environment 100 may include vehicle 110. Environment 100 can be applied to autonomous driving scenarios, assisted driving scenarios, intelligent transportation scenarios, etc. Vehicle 110 may be an autonomous vehicle, etc. An autonomous vehicle is a vehicle with autonomous driving capability (or driverless driving capability), also known as a driverless car, autonomous driving vehicle, etc. For ease of understanding, this disclosure uses vehicle 110 as an example of an autonomous vehicle for explanation and illustration, but is not limited thereto.
[0029] In some scenarios, vehicle 110 can be assigned to provide travel services to users. For example, users can obtain travel services provided by vehicle 110 through a travel application. In some scenarios, vehicle 110 may also be called a driverless taxi or robotaxi. During the process of vehicle 110 providing travel services to users, vehicle 110 may be equipped with a safety operator. The safety operator can, for example, take over vehicle 110 in case of an emergency. Alternatively, vehicle 110 may also be in an unmanned state.
[0030] During the operation of vehicle 110, vehicle 110 can perceive some surrounding environmental data 120. For example, some cameras deployed in vehicle 110 can perceive some images, and some radars deployed in vehicle 110 can perceive some point cloud data, etc. In some scenarios, vehicle 110 can also be equipped with one or more display devices inside the vehicle to provide human-machine interaction functions.
[0031] In some scenarios, such environmental data 120 can be acquired by electronic device 130. Such electronic device 130 may include a terminal and / or a server.
[0032] Such a terminal can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 130 may also support any type of user-facing interface (such as "wearable" circuitry).
[0033] Such servers can be standalone physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Servers can include, for example, computing systems / servers such as mainframes, edge computing nodes, and computing devices in cloud environments, etc.
[0034] The electronic device 130 can be installed in the vehicle 110, or deployed in any electronic unit of the vehicle 110 (such as sensors, control units, etc.), or it can be independent of the vehicle 110.
[0035] When the electronic device 130 is independent of the vehicle 110, a communication connection can be established between the vehicle 110 and the electronic device 130. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus, and Wi-Fi connections; the embodiments of this disclosure are not limited in this respect. Based on this, the vehicle 110 and the electronic device 130 can achieve signaling interaction through their communication connection.
[0036] Therefore, the electronic device 130 acquires the environmental data 120 of the vehicle 110, and can detect objects in the vicinity of the electronic device 130 based on the data processing of the environmental data 120. Such objects can be larger objects in the traffic environment, such as traffic signs, plants (e.g., Figure 1 Object 141), other vehicles, pedestrians, etc. Such objects can also be smaller objects in the traffic environment, such as traffic cones (e.g., Figure 1 (Object 142) Small stones, small branches, water accumulation on the road surface, etc.
[0037] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0038] Example process
[0039] The following will be referenced Figure 2 This document describes an example process for object detection according to some embodiments of the present disclosure. Figure 2 A schematic diagram of an example process for object detection according to some embodiments of the present disclosure is shown. Process 200 may, for example, be performed by... Figure 1 The electronic device 130 shown is provided. See below for reference. Figure 1 To describe process 200.
[0040] At box 210, electronic device 130 determines the target region corresponding to an object of the target type from the perceived image.
[0041] As examples, the perceived image can be one type of data in the environmental data 120. For instance, the perceived image could include images captured by a camera mounted on a vehicle. The target type object is a small object, such as object 142, or other smaller objects.
[0042] In some embodiments, an object of the target type may satisfy one or more of the following conditions: size less than or equal to a size threshold, volume less than or equal to a volume threshold, proportion in the perceived image less than or equal to a proportion threshold, and object pixels less than or equal to a pixel threshold.
[0043] As examples, a target type object occupies a number of pixels in a perceptual image, and the region formed by these pixels can be considered the target region corresponding to the target type object. Alternatively, the circumscribed quadrilateral or circumscribed polygon of the region formed by these pixels can also be considered the target region corresponding to the target type object. In practical applications, this can be set according to requirements.
[0044] In some embodiments, the electronic device 130 can determine the target region corresponding to an object of a target type from a perceived image based on various image processing methods. To improve the accuracy of the target region, embodiments of this disclosure can utilize image processing models for target region detection.
[0045] Specifically, the electronic device 110 can use an image processing model to process the perceived image, determine feature information corresponding to multiple scales, determine multiple sets of candidate regions corresponding to multiple scales based on the feature information, and determine the target region corresponding to the object based on the multiple sets of candidate regions.
[0046] As some examples, an image processing model can be a neural network model capable of performing functions such as image feature extraction, feature fusion, image detection, and / or image classification. For example, an image processing model can be a neural network model based on the YOLOv8 (You Only Look Once version 8) model.
[0047] Furthermore, the image processing model may include, for example, a feature extraction network (Backbone), a feature fusion network (Neck), a detection head network (Head), etc., to achieve the above-mentioned image processing functions.
[0048] The following is in conjunction with the appendix Figure 3 This document describes an example process for two-dimensional image processing according to some embodiments of the present disclosure. Figure 3 A schematic diagram of an example process for two-dimensional image processing according to some embodiments of the present disclosure is shown. Process 300 can be implemented by electronic device 130 using an image processing model. The image processing model can be deployed in electronic device 130 or in other electronic devices, so that electronic device 130 implements process 300 through an interface.
[0049] As examples, process 300 can be an image processing procedure for the perceived image 310. Electronic device 130 can extract image features at multiple scales from the perceived image 310, with different scales corresponding to different resolutions / depths of semantic information.
[0050] For example, the image features 321 extracted from layer P5, 322 extracted from layer P4, 323 extracted from layer P3, and 324 extracted from layer P2 have progressively lower depth semantics and progressively higher resolution. Image features 324 extracted from the higher-resolution layer P2 can better capture the detailed information of smaller target objects, that is, the detailed information of target-type objects.
[0051] In some embodiments, the Feature Pyramid Network (FPN) and Path Aggregation Network (PAN) in the Feature Fusion Network (Neck) can be used for multi-scale image feature fusion. Therefore, embodiments of this disclosure enable the image features 324 extracted by the higher resolution layer (P2 layer) to participate more fully in the detection process at different scales, further improving the comprehensiveness of information about the target type of object.
[0052] In some embodiments, the detection head network can independently and in parallel perform object detection on fused image features at different scales (e.g., fused features 331, fused features 332, fused features 333, and fused features 334). Further, by fusing the results of different object detections (e.g., detection results 341, 342, 343, and 344), the electronic device 130 can obtain the target region 350 in the perceived image 310 corresponding to an object of the target type. Thus, the electronic device 130 can improve the accuracy of the target region corresponding to the target type of object determined from the perceived image, and increase the detection probability of the target type of object.
[0053] In some embodiments, the image processing model can be pre-trained to further improve the model's object detection performance. Specifically, the image processing model can be trained using a set of sample images, which includes at least one of the following: multiple sample images corresponding to different viewpoints; a sample image determined by stitching together multiple sample images from different viewpoints; or a sample image obtained by combining an image region corresponding to an object of the target type with a preset background.
[0054] As examples, perspectives include camera view, webcam view, virtual camera view, etc. Sample images from different perspectives can capture environmental information about the vehicle from different angles. This increases the proportion of positive samples in the sample images, optimizing the positive-to-negative sample ratio, thus mitigating the imbalance between positive and negative samples and improving the stability of model training. Furthermore, training the image processing model with sample images from multiple perspectives can also improve the model's generalization ability.
[0055] As examples, image stitching can be achieved by stitching multiple images in a certain order or in a certain combination. Stitching multiple sample images from different perspectives to determine the aforementioned sample images can further increase the co-occurrence probability of background and target types of objects, thereby further improving the training stability of the image processing model.
[0056] As examples, a preset background can include a preset blank background image, an image selected from sample images, and so on. The image region corresponding to the target type object can be the pixel region occupied by the target type object in the sample image. Mixing multiple image regions or combining one or more image regions into the preset background can further increase the proportion of positive samples, thereby further improving sample diversity and robustness of the image processing model.
[0057] Therefore, embodiments of this disclosure can more accurately determine the target region corresponding to an object of the target type from a perceived image.
[0058] In some embodiments, the electronic device 130 can determine a target region corresponding to an object of a target type from multiple perceived images to further improve the accuracy of the target region. As some examples, the electronic device 130 can determine each set of target regions from each perceived image separately, and combine the sets of target regions in each set of perceived images to determine the target region corresponding to the object of the target type.
[0059] In some embodiments, multiple sensing images may have overlapping content. Based on this, the electronic device 130 can fuse the regions where the target type of object is located in the multiple sensing images to reduce the repetition of the target region and thereby improve the accuracy of the target region.
[0060] Specifically, the perceived image may include at least a first perceived image and a second perceived image, wherein the first perceived image is captured by a first camera mounted on the autonomous vehicle, and the second perceived image is captured by a second camera mounted on the autonomous vehicle. Based on this, the electronic device 130 can determine a first set of candidate regions corresponding to an object of the target type from the first perceived image; determine a second set of candidate regions corresponding to an object of the target type from the second perceived image; and determine a target region corresponding to the object by fusing the first set of candidate regions and the second set of candidate regions.
[0061] As examples, the first camera and the second camera can be different cameras mounted on the vehicle 110. The first camera and the second camera can correspond to different viewpoints, or they can correspond to different focal lengths within the same viewpoint. Both the first camera and the second camera can capture perceived images containing objects of the target type.
[0062] Furthermore, the electronic device 130 can determine one or more regions corresponding to the target type of object from the first perceived image, i.e., a first set of candidate regions. The electronic device 130 can determine the first set of candidate regions using the image processing model described above, or it can determine the first set of candidate regions using other suitable methods. Similarly, the electronic device 130 can also determine a second set of candidate regions from the second perceived image in such a manner.
[0063] Based on this, the electronic device 130 can fuse the first set of candidate regions and the second set of candidate regions to determine the target region corresponding to the target type of object (i.e., the smaller-sized object). As some examples, fusion can be achieved in various ways, such as using neural network models, or other reasonable methods, such as deduplication and recombination.
[0064] In some embodiments, the electronic device 130 may preferentially retain candidate regions in the perceived image with higher perception accuracy to improve the detection accuracy of objects for smaller targets. Specifically, the electronic device 130 may, in response to a first perception range of a first camera being covered by a second perception range of a second camera, compare the perception accuracy of the first camera and the second camera; in response to a first camera having a higher perception accuracy than the second camera, add a first set of candidate regions determined from the first perceived image to a target region set; and, based on a comparison between the second set of candidate regions and the first perception range, determine at least one candidate region from the second set of candidate regions; and add at least one candidate region to the target region set.
[0065] As examples, the sensing range can include the extent of the environment perceived in the sensing image; the larger the sensing range, the larger the area of the environment indicated by the sensing image. Sensing precision can include resolution, focal length, focal length, etc.; the higher the sensing precision, the clearer the details in the sensing image.
[0066] Based on this, when the first sensing area of the first camera is covered by the second sensing range of the second camera, and the sensing accuracy of the first camera is higher than that of the second camera, the electronic device 130 can add all the first group of candidate regions in the first sensing image to the target region set.
[0067] Furthermore, based on the comparison between the second set of candidate regions and the first sensing range, the electronic device 130 can identify candidate regions that do not overlap with or partially overlap with the first sensing range and add them to the target region set. Thus, the electronic device 130 can achieve the fusion of object regions in multiple perceived images, avoiding the omission of object regions and thereby improving the accuracy of the target region.
[0068] In some embodiments, the electronic device 130 may determine at least one candidate region to be added to the target region set from the second set of candidate regions based on overlapping regions, thereby improving the accuracy of the target region while reducing the probability of repeated object detection. Specifically, the electronic device 130 may determine the overlapping regions between each candidate region in the second set of candidate regions and the first sensing range; and determine at least one candidate region from the second set of candidate regions based on the overlapping regions, wherein the ratio of the overlapping region corresponding to the at least one candidate region to the corresponding candidate region is less than a threshold.
[0069] As examples, electronic device 130 can determine overlapping regions by projecting a first sensing range onto a second sensing image to calculate the intersection of each candidate region of the second sensing image with the projected region. The threshold can be set based on actual needs.
[0070] Taking a threshold of 70% as an example, the following is combined with the appendix Figure 4 This describes an example multi-camera perspective fusion process according to some embodiments of the present disclosure. Figure 4 A schematic diagram of an example multi-camera view fusion process according to some embodiments of the present disclosure is shown. Process 400 can be implemented by electronic device 130.
[0071] like Figure 4 As shown, the electronic device 130 projects a first sensing range onto a second sensing image 410 to obtain a projection area 420. The second sensing image contains candidate areas 431 and 432 of an object. The ratio of the overlap between candidate area 431 and the projection area 420 to candidate area 431 is 25%, and the ratio of the overlap between candidate area 432 and the projection area 420 to candidate area 431 is 75%.
[0072] Furthermore, the electronic device 130 compares each ratio with a threshold. Based on the comparison of the overlapping area of each candidate region with the projection region 420 (i.e., the first projection range) and the threshold, the electronic device 130 can add the candidate region 431 to the target region set.
[0073] Therefore, the embodiments of this disclosure improve the recall capability of target type objects through such fusion.
[0074] Back Figure 2 At box 220, electronic device 130 projects point cloud data into a two-dimensional space corresponding to the perceived image to determine the point cloud distribution associated with the target region.
[0075] As examples, point cloud data can be generated by the sensors of vehicle 110 or by other reasonable means. Electronic device 130 projects the point cloud data into a two-dimensional space corresponding to the perceived image, thereby obtaining points in that two-dimensional space that fall within the range of the perceived image. The point cloud distribution can indicate the distribution of points falling in that two-dimensional space, such as number, density, etc.
[0076] Specifically, the electronic device 130 can project point cloud data into a two-dimensional space to determine a set of projection points associated with a target area. Based on the number of these projection points, the electronic device 130 can determine the distribution of the point cloud.
[0077] At frame 230, electronic device 130 determines the depth information corresponding to the target area based on the point cloud distribution.
[0078] As examples, depth information includes the depth of an object of the target type relative to the camera corresponding to the perceived image. Electronic device 130 can determine depth information using projection points corresponding to point cloud data, or it can use a high-precision map to obtain the depth of the ground associated with the target area as the object's depth information.
[0079] Furthermore, when there are many projection points, the electronic device 130 can obtain depth information using point cloud data; when there are few projection points, the electronic device 130 can obtain depth information using a high-precision map. Based on this, the electronic device 130 can more accurately detect objects of the target type.
[0080] In some embodiments, the electronic device 130 may determine depth information corresponding to the target region based on a first reference depth of a set of projection points in response to the number of a set of projection points reaching a threshold.
[0081] As examples, the first reference depth of a set of projection points can be the depth of the set of projection points relative to the camera position corresponding to the perceived image. Thus, the electronic device 130 can more accurately locate objects of the target type.
[0082] In some embodiments, the electronic device 130 may, in response to a set of projection points being less than or equal to a threshold, determine a reference equidistant line corresponding to the perceived image based on a high-precision map; determine a second reference depth of the road surface corresponding to the target area in the perceived image based on the reference equidistant line; and determine depth information corresponding to the target area based on the second reference depth.
[0083] As examples, reference equidistant lines can be drawn in a perceived image based on a high-precision map. Objects on the same reference equidistant line can have the same depth relative to the corresponding camera position in the perceived image.
[0084] As examples, electronic device 130 can determine the road surface corresponding to a target area in a perceived image based on a set of projection points. Based on the reference equidistant line where the road surface corresponding to the target area is located, the depth of the road surface relative to the camera position can be determined, thereby determining the depth information corresponding to the target area.
[0085] Based on different point cloud distributions, the electronic device 130 can determine the depth information of the target area in different ways. Therefore, the electronic device 130 can determine more accurate depth information.
[0086] At box 240, electronic device 130 determines the set of points corresponding to the target region from point cloud data based on depth information and the search range corresponding to the target type.
[0087] As examples, different object types can have different object size ranges, thus allowing different object types to correspond to different search ranges, thereby ensuring the accuracy of object search.
[0088] Based on this, and using depth information, the electronic device 130 can determine the search starting point. Based on the search range corresponding to the target type, the electronic device 130 can determine the furthest search range of objects of that target type. Thus, the electronic device can start searching from the point cloud data based on the search starting point until it reaches the search range, thereby determining the points falling within that search range, which is the set of points corresponding to the target area.
[0089] In some embodiments, the electronic device 130 can use a point cloud statistical map to search for a set of points to improve the accuracy of the point set. Specifically, the electronic device 130 can project point cloud data onto a preset plane to generate a point cloud statistical map associated with the preset plane, the point cloud statistical map indicating the distribution of multiple projected points in multiple grids; determine a set of grids to be searched among the multiple grids based on depth information and search range; determine at least one target grid based on the number of projected points in the set of grids; and determine the set of points corresponding to the at least one grid from the point cloud data.
[0090] As examples, the preset plane can be the plane where the perceived image is located, or it can be a bird's-eye view plane. Based on the projection of point cloud data onto the preset plane, the electronic device 130 can generate a point cloud statistical map under the preset plane to achieve the generation of the three-dimensional shape of the target type object.
[0091] Furthermore, the electronic device 130 can determine the center point based on depth information, and search for the corresponding point cloud grid on the point cloud statistical map from the center point until the search range is reached.
[0092] Therefore, the electronic device 130 can determine whether a point cloud grid corresponds to an object based on the number of projected points in a set of point grids. In this way, the electronic device 130 can determine the target grid corresponding to the object, and based on this target grid, the electronic device 130 can determine the set of points from the point cloud data.
[0093] At frame 250, electronic device 130 determines the object recognition result based on a set of points, the recognition result indicating the object's shape and / or position in three-dimensional space.
[0094] As examples, each point in the point set has corresponding three-dimensional position information and depth information. Based on the three-dimensional position information and depth information of each point in the point set, the shape and / or position of the point object in three-dimensional space can be determined.
[0095] In this way, the electronic device 130 can accurately detect target-type objects in the environment surrounding the vehicle 110, thereby avoiding situations such as collisions between the electronic device 130 and the object, thus improving vehicle safety.
[0096] The following is in conjunction with the appendix Figure 5 Example procedures for three-dimensional image processing according to some embodiments of the present disclosure are described exemplarily. Figure 5 A schematic diagram of an example process for two-dimensional image processing according to some embodiments of the present disclosure is shown. Process 500 can be implemented by electronic device 130.
[0097] like Figure 5 As shown, the electronic device 130 inputs the perceived image 510 (e.g., the perceived image 310) into the image processing model 520 to obtain the target region 530 corresponding to the perceived image 510.
[0098] Furthermore, the electronic device 130 performs a depth estimation of the target area (540). Specifically, the electronic device 130 can use a high-precision map 551 and / or laser point cloud data 552 (i.e., point cloud data) to obtain the depth information 560 corresponding to the target area.
[0099] Furthermore, the electronic device 130 can determine (570) the three-dimensional shape and / or position based on the laser point cloud data 552 and depth information 560 through the above-mentioned point cloud statistical map, so as to obtain the polygon 580 of the target type object, the polygon representing the shape and / or position of the target type object in three-dimensional space.
[0100] Based on this approach, embodiments of this disclosure can determine the target region corresponding to the target type of object, thereby determining the point cloud distribution associated with the target region based on point cloud data, reducing the interference of a large amount of point cloud data on the detection of the target type of object; furthermore, embodiments of this disclosure obtain the depth information corresponding to the target region through the point cloud distribution, ensuring the accuracy of the object's position in three-dimensional space; thus, based on such depth information and the search range corresponding to the target type, this disclosure searches for the point set corresponding to the target region, ensuring the accuracy of the point set associated with the object, thereby improving the accuracy of the object recognition result determined based on the point set.
[0101] Therefore, embodiments of this disclosure can obtain a more accurate position and / or shape of a target type object in three-dimensional space based on point cloud data and depth information, according to the target region. Furthermore, by utilizing the three-dimensional shape and / or position of such objects, embodiments of this disclosure can help improve the safety of vehicle operation.
[0102] Example devices and equipment
[0103] Figure 6 A schematic structural block diagram of an object detection apparatus 600 according to certain embodiments of the present disclosure is shown. Apparatus 600 may be implemented as or included in electronic device 130. Various modules / components in apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0104] As shown in the figure, the device 600 includes a first determining module 610 configured to determine a target region corresponding to an object of a target type from a perceived image; a projection module 620 configured to project point cloud data onto a two-dimensional space corresponding to the perceived image to determine the point cloud distribution associated with the target region; a second determining module 630 configured to determine depth information corresponding to the target region based on the point cloud distribution; a third determining module 640 configured to determine a set of points corresponding to the target region from the point cloud data based on the depth information and a search range corresponding to the target type; and a fourth determining module 650 configured to determine the object recognition result based on the point set, the recognition result indicating the shape and / or position of the object in three-dimensional space.
[0105] In some embodiments, the perceived image includes at least a first perceived image and a second perceived image, the first perceived image being captured by a first camera mounted on the autonomous vehicle, the second perceived image being captured by a second camera mounted on the autonomous vehicle, and the first determining module 610 is further configured to: determine a first set of candidate regions corresponding to an object of the target type from the first perceived image; determine a second set of candidate regions corresponding to an object of the target type from the second perceived image; and determine a target region corresponding to the object by fusing the first set of candidate regions and the second set of candidate regions.
[0106] In some embodiments, the first determining module 610 is further configured to: in response to a first sensing range of the first camera being covered by a second sensing range of the second camera, compare the sensing accuracy of the first camera and the second camera; in response to a sensing accuracy of the first camera being higher than that of the second camera, add a first set of candidate regions determined from the first sensing image to a target region set; and based on a comparison between the second set of candidate regions and the first sensing range, determine at least one candidate region from the second set of candidate regions; and add at least one candidate region to the target region set.
[0107] In some embodiments, the first determining module 610 is further configured to: determine the overlapping region between each candidate region in the second group of candidate regions and the first sensing range; and
[0108] Based on the overlapping region, at least one candidate region is determined from the second group of candidate regions, wherein the ratio of the overlapping region corresponding to the at least one candidate region to the corresponding candidate region is less than a threshold.
[0109] In some embodiments, the projection module 620 is further configured to: project point cloud data into a two-dimensional space to determine a set of projection points associated with a target region; and determine the point cloud distribution based on the number of projection points in the set.
[0110] In some embodiments, the projection module 620 is further configured to: in response to a threshold number of a set of projection points, determine depth information corresponding to the target region based on a first reference depth of the set of projection points.
[0111] In some embodiments, the projection module 620 is further configured to: in response to a set of projection points being less than or equal to a threshold, determine a reference equidistant line corresponding to the perceived image based on a high-precision map; determine a second reference depth of the road surface corresponding to the target area in the perceived image based on the reference equidistant line; and determine depth information corresponding to the target area based on the second reference depth.
[0112] In some embodiments, the third determining module 640 is further configured to: project point cloud data onto a preset plane to generate a point cloud statistical map associated with the preset plane, the point cloud statistical map indicating the distribution of multiple projected points in multiple grids; determine a set of grids to be searched in the multiple grids based on depth information and search range; determine at least one target grid based on the number of projected points in the set of grids; and determine a set of points corresponding to at least one grid from the point cloud data.
[0113] In some embodiments, the first determining module 610 is further configured to: process a perceived image using an image processing model to determine feature information corresponding to multiple scales; determine multiple sets of candidate regions corresponding to multiple scales based on the feature information; and determine a target region corresponding to an object based on the multiple sets of candidate regions.
[0114] In some embodiments, the image processing model is trained using a set of sample images, which includes at least one of the following: multiple sample images corresponding to different viewpoints; a sample image determined by stitching together multiple sample images from different viewpoints; or a sample image obtained by combining an image region corresponding to an object of a target type with a preset background.
[0115] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The electronic device 700 shown can be used to achieve Figure 1 Electronic devices 130 or Figure 6 Device 600.
[0116] like Figure 7 As shown, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.
[0117] Electronic device 700 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 700.
[0118] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0119] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0120] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0121] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0122] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0123] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0124] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0126] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for object detection, comprising: Identify the target region corresponding to the target type from the perceived image; The point cloud data is projected onto a two-dimensional space corresponding to the perceived image to determine the point cloud distribution associated with the target region; Based on the point cloud distribution, determine the depth information corresponding to the target region; Based on the depth information and the search range corresponding to the target type, a set of points corresponding to the target region is determined from the point cloud data; as well as Based on the set of points, the recognition result of the object is determined, and the recognition result indicates the shape and / or position of the object in three-dimensional space.
2. The method according to claim 1, wherein the perceived image includes at least a first perceived image and a second perceived image, the first perceived image being captured by a first camera mounted on the autonomous vehicle, the second perceived image being captured by a second camera mounted on the autonomous vehicle, and determining the target region corresponding to the object from the perceived image includes: Determine a first group of candidate regions corresponding to the object of the target type from the first perceived image; Determine a second set of candidate regions corresponding to the object of the target type from the second perceived image; as well as The target region corresponding to the object is determined by fusing the first group of candidate regions and the second group of candidate regions.
3. The method according to claim 2, wherein determining the target region corresponding to the object by fusing the first group of candidate regions and the second group of candidate regions comprises: In response to the first sensing range of the first camera being covered by the second sensing range of the second camera, the sensing accuracy of the first camera and the second camera are compared. In response to the fact that the perception accuracy of the first camera is higher than that of the second camera, the first set of candidate regions determined from the first perceived image is added to the target region set; as well as Based on the comparison between the second group of candidate regions and the first sensing range, at least one candidate region is determined from the second group of candidate regions; as well as Add the at least one candidate region to the target region set.
4. The method of claim 3, wherein determining at least one candidate region from the second group of candidate regions based on a comparison between the second group of candidate regions and the first sensing range comprises: Determine the overlapping area between each candidate region in the second group of candidate regions and the first sensing range; as well as Based on the overlapping region, at least one candidate region is determined from the second group of candidate regions, wherein the ratio of the overlapping region corresponding to the at least one candidate region to the corresponding candidate region is less than a threshold.
5. The method of claim 1, wherein projecting the point cloud data onto the two-dimensional space to determine the point cloud distribution comprises: The point cloud data is projected onto the two-dimensional space to determine a set of projection points associated with the target region; as well as The point cloud distribution is determined based on the number of the set of projection points.
6. The method according to claim 5, wherein determining the depth information corresponding to the target region based on the point cloud distribution includes: In response to the number of the set of projection points reaching a threshold, the depth information corresponding to the target region is determined based on a first reference depth of the set of projection points.
7. The method according to claim 5, wherein determining the depth information corresponding to the target region based on the point cloud distribution includes: In response to the fact that the number of the set of projection points is less than or equal to a threshold, a reference equidistant line corresponding to the perceived image is determined based on a high-precision map; Based on the reference equidistant line, a second reference depth of the road surface corresponding to the target area in the perceived image is determined; as well as Based on the second reference depth, the depth information corresponding to the target area is determined.
8. The method of claim 1, wherein determining the point set from the point cloud data based on the depth information and the search range comprises: The point cloud data is projected onto a preset plane to generate a point cloud statistical map associated with the preset plane, the point cloud statistical map indicating the distribution of multiple projected points in multiple grids; Based on the depth information and the search range, a set of grids to be searched among the multiple grids is determined; Based on the number of projection points in the set of grids, at least one target grid is determined; as well as The point set corresponding to the at least one grid is determined from the point cloud data.
9. The method of claim 1, wherein determining the target region corresponding to the object from the perceived image comprises: The perceived image is processed using an image processing model to determine feature information corresponding to multiple scales; Based on the feature information, multiple sets of candidate regions corresponding to the multiple scales are determined; as well as Based on the multiple candidate regions, the target region corresponding to the object is determined.
10. The method of claim 9, wherein the image processing model is trained using a set of sample images, the set of sample images comprising at least one of the following: Multiple sample images corresponding to different viewpoints; The sample image is determined by stitching together multiple sample images from different perspectives; A sample image is obtained by combining an image region corresponding to an object of the target type with a preset background.
11. An apparatus for object detection, comprising: The first determining module is configured to determine the target region corresponding to an object of the target type from the perceived image; The projection module is configured to project point cloud data onto a two-dimensional space corresponding to the perceived image to determine the point cloud distribution associated with the target region; The second determining module is configured to determine the depth information corresponding to the target region based on the point cloud distribution; The third determining module is configured to determine a set of points corresponding to the target region from the point cloud data based on the depth information and the search range corresponding to the target type. as well as The fourth determining module is configured to determine the recognition result of the object based on the set of points, the recognition result indicating the shape and / or position of the object in three-dimensional space.
12. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.
14. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.