Object Detection Method and Device
The monocular camera obtains 2D information and builds a 3D frame. Combined with the distance transformation relationship, the monocular camera cannot perceive the vehicle's 3D information, and realizes efficient 3D distance detection, reducing costs and improving ADAS performance and security.
Patent Information
- Application Number
- CN202010806077.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-08-12
AI Technical Summary
In the prior art, monocular cameras cannot sense the 3D information of the vehicle, limiting the performance and safety of the autonomous driving assistance system (ADAS).
By acquiring the 2D information in the image collected by the monocular camera, a 3D frame of the vehicle is constructed, and the distance transformation relationship is used to calculate the distance between the vehicle and other objects, thereby realizing the acquisition of 3D information.
It is realized that the 3D distance information of the vehicle and other objects is obtained without the need for a multi-object camera or a depth camera, which reduces the cost of the target detection device and improves the performance and safety of ADAS.
Smart Images

Figure CN114078247B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to a target detection method and apparatus in the field of autonomous driving or intelligent transportation. Background Art
[0002] With the rapid development of the economy, the market share of automobiles has increased year by year, and the frequency of traffic accidents has also increased sharply. To reduce the frequency of traffic accidents, an Advanced Driver Assistance System (ADAS) has emerged. The ADAS can sense the three-dimensional (3D) information of external vehicles through a sensing module, such as information about the position and size of the vehicle, remind the driver of possible dangers during vehicle driving, and perform path planning and driving strategy adjustment based on the sensed information. The more accurate the 3D information of the vehicle sensed by the sensing module, the better the performance of the ADAS and the higher the safety of the vehicle.
[0003] Currently, the cost of a monocular camera is relatively low, and a monocular camera is usually configured in the sensing module. However, a monocular camera senses information with a single camera and cannot sense the 3D information of the vehicle. Summary of the Invention
[0004] This application provides a target detection method and apparatus, which can obtain the 3D information of a vehicle based on an image collected by a monocular camera.
[0005] To achieve the above object, the embodiments of this application adopt the following technical solutions:
[0006] First aspect, an embodiment of the present application provides a target detection method and apparatus. The method includes: obtaining 2D information of a second object in an image obtained by a first object; the 2D information of the second object includes: coordinate information of endpoints of a first dividing line of the second object and type information of the second object, where the first dividing line is a dividing line between a first surface and a second surface of the second object, and at least one of the first surface and the second surface is included in a first 2D frame of the second object, and the first 2D frame is a polygon that includes the second object and is included in the image; obtaining 3D information of the second object according to the 2D information of the second object; the 3D information of the second object includes coordinate information of endpoints of a second dividing line, where the second dividing line is a dividing line between the first surface and the second surface in a 3D frame of the second object, and the 3D frame is a 3D model of the second object, and the length of the boundary of the 3D frame corresponds to the type information of the second object; obtaining a first transformation relationship according to the coordinate information of the endpoints of the first dividing line and the coordinate information of the endpoints of the second dividing line; the first transformation relationship is a transformation relationship between the three-dimensional coordinate system corresponding to the 3D frame and the three-dimensional coordinate system of the first object when the distance between the second object and the first object is K, and K>0; obtaining the distance between the second object and the first object according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the first transformation relationship.
[0007] The method provided in the above first aspect can construct a 3D frame based on an image collected by a monocular camera, and use the transformation relationship between the three-dimensional coordinate system corresponding to the 3D frame and the three-dimensional coordinate system of the first object when the distance between the second object and the first object is K to obtain the distance between the second object and the first object, and the computing cost is relatively low. In addition, through the method provided in the first aspect, it is not necessary to use a multi-camera or a depth camera to obtain the distance between the second object and the first object, so the cost of the target detection device can be reduced.
[0008] In a possible implementation, the 2D information of the second object further includes: coordinate information of the first 2D bounding box of the second object; obtaining the 3D information of the second object based on the 2D information of the second object, including: constructing the 3D bounding box according to the type information of the second object to obtain the coordinate information of the 3D bounding box; determining the first plane and the second plane according to the coordinate information of the first 2D bounding box and the coordinate information of the endpoints of the first dividing line; obtaining the coordinate information of the endpoints of the second dividing line according to the first plane, the second plane and the coordinate information of the 3D bounding box. Based on the above method, when the 2D information of the second object includes the coordinate information of the endpoints of the first dividing line of the second object, the coordinate information of the first 2D bounding box of the second object, and the type information of the second object, a 3D bounding box can be constructed, and the coordinate information of the endpoints of the second dividing line in the 3D bounding box can be determined according to the coordinate information of the endpoints of the first dividing line and the coordinate information of the first 2D bounding box, so as to subsequently obtain the transformation relationship between the three-dimensional coordinate system corresponding to the 3D bounding box and the three-dimensional coordinate system of the first object according to the points on the second dividing line and the mapping points of these points on the first dividing line.
[0009] In a possible implementation, the 2D information of the second object further includes: plane information of the second object, and the plane information of the second object is used to indicate the first plane and the second plane; obtaining the 3D information of the second object based on the 2D information of the second object, including: constructing the 3D bounding box according to the type information of the second object to obtain the coordinate information of the 3D bounding box; obtaining the coordinate information of the endpoints of the second dividing line according to the plane information of the second object and the coordinate information of the 3D bounding box. Based on the above method, when the 2D information of the second object includes the coordinate information of the endpoints of the first dividing line of the second object, the plane information of the second object, and the type information of the second object, a 3D bounding box can be constructed, and the coordinate information of the endpoints of the second dividing line in the 3D bounding box can be determined according to the plane information of the second object, so as to subsequently obtain the transformation relationship between the three-dimensional coordinate system corresponding to the 3D bounding box and the three-dimensional coordinate system of the first object according to the points on the second dividing line and the mapping points of these points on the first dividing line.
[0010] In a possible implementation, obtaining the distance between the second object and the first object according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the first transformation relationship includes: performing coordinate transformation on the first coordinate through the first transformation relationship, the intrinsic matrix, and the extrinsic matrix to obtain a second coordinate. The first coordinate is on the second dividing line, the distance between the second coordinate and the first object is K, the second coordinate, the third coordinate, and the first object are on a straight line, the third coordinate is the coordinate corresponding to the first coordinate on the first dividing line, the intrinsic matrix is the intrinsic matrix of the device that captures the image, and the extrinsic matrix is the extrinsic matrix of the device; performing a composite operation on the second coordinate, the third coordinate, and the K to obtain the distance between the second object and the first object. Based on the above method, the distance between the second object and the first object can be obtained according to the transformation relationship between the three-dimensional coordinate system corresponding to the 3D box and the three-dimensional coordinate system of the first object, and the property of similar triangles.
[0011] In a possible implementation, the method further includes: obtaining a second transformation relationship according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the distance between the second object and the first object; the second transformation relationship is the transformation relationship between the three-dimensional coordinate system corresponding to the 3D box and the three-dimensional coordinate system of the first object. Based on the above method, the transformation relationship between the three-dimensional coordinate system corresponding to the 3D box and the three-dimensional coordinate system of the first object can be obtained, so as to determine the orientation angle of the second object or calibrate the length and / or width of the 3D box according to the transformation relationship.
[0012] In a possible implementation, the method further includes: rotating the 3D box N times with the second dividing line as the center to obtain the corresponding second 2D box for the 3D box after each rotation. The second 2D box is obtained by performing coordinate transformation on the rotated 3D box, and N is a positive integer; obtaining the second 2D box with the shortest distance between its boundary and the boundary of the first 2D box among the N second 2D boxes corresponding to the 3D boxes after N rotations according to the N second 2D boxes corresponding to the 3D boxes after N rotations and the first 2D box; determining the orientation angle of the second object according to the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance. Based on the above method, the second 2D box corresponding to the rotated 3D box can be compared with the first 2D box to obtain the second 2D box closest to the first 2D box, thereby obtaining the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance, and determining the orientation angle of the second object according to the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance. In this way, when planning a path, in addition to referring to the distance between the second object and the first object, the orientation angle of the second object can also be referred to, making the reference of the planned path higher.
[0013] In a possible implementation, the second 2D box is obtained by performing coordinate transformation on the rotated 3D box, including: the second 2D box is obtained by performing coordinate transformation on the rotated 3D box through the second transformation relationship, the intrinsic matrix, and the extrinsic matrix. The intrinsic matrix is the intrinsic matrix of the device that captures the image, and the extrinsic matrix is the extrinsic matrix of the device. Based on the above method, the second 2D box can be obtained by performing coordinate transformation on the rotated 3D box through the second transformation relationship, the intrinsic matrix, and the extrinsic matrix, so as to obtain the orientation angle of the second object according to the second 2D box subsequently.
[0014] In a possible implementation, the method further includes: adjusting the length of the boundary of the 3D box corresponding to the second 2D box with the shortest distance M times, obtaining the third 2D boxes corresponding to the 3D boxes after each adjustment. The third 2D box is obtained by performing coordinate transformation on the adjusted 3D box, and M is a positive integer; obtaining, according to the M third 2D boxes corresponding to the 3D boxes after M adjustments and the first 2D box, the third 2D box with the shortest distance between the boundary and the boundary of the first 2D box among the M third 2D boxes; determining the length of the boundary of the second object according to the length of the boundary of the 3D box corresponding to the third 2D box with the shortest distance. Based on the above method, the third 2D box corresponding to the adjusted 3D box can be compared with the first 2D box to obtain the third 2D box closest to the first 2D box, so as to obtain the length of the boundary of the second object. In this way, a more accurate size of the second object can be obtained.
[0015] In a possible implementation, the third 2D box is obtained by performing coordinate transformation on the adjusted 3D box, including: the third 2D box is obtained by performing coordinate transformation on the adjusted 3D box through the second transformation relationship, the intrinsic matrix, and the extrinsic matrix. The intrinsic matrix is the intrinsic matrix of the device that captures the image, and the extrinsic matrix is the extrinsic matrix of the device. Based on the above method, the third 2D box can be obtained by performing coordinate transformation on the adjusted 3D box through the second transformation relationship, the intrinsic matrix, and the extrinsic matrix, so as to obtain a more accurate size of the second object according to the third 2D box subsequently.
[0016] In a possible implementation, obtaining the 2D information of the second object in the image captured by the first object includes: inputting the image into a neural network model to obtain the 2D information of the second object. Based on the above method, the 2D information of the second object can be obtained through the neural network model, so as to obtain the distance between the second object and the first object according to the 2D information of the second object.
[0017] Second aspect, embodiments of the present application provide an object detection device, which can implement the method in the above first aspect or any possible implementation manner of the first aspect. The device includes corresponding units or components for executing the above method. The units included in the device can be implemented in software and / or hardware. The device can be, for example, an ADAS, or a chip, a chip system, or a processor that supports the ADAS to implement the above method.
[0018] Third aspect, embodiments of the present application provide an object detection device, including: a processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the device implements the method described in the above first aspect or any possible implementation manner of the first aspect.
[0019] Fourth aspect, embodiments of the present application provide an object detection device, which is used to implement the method described in the above first aspect or any possible implementation manner of the first aspect.
[0020] Fifth aspect, embodiments of the present application provide a computer-readable medium, on which computer programs or instructions are stored. When the computer programs or instructions are executed, the computer executes the method described in the above first aspect or any possible implementation manner of the first aspect.
[0021] Sixth aspect, embodiments of the present application provide a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer executes the method described in the above first aspect or any possible implementation manner of the first aspect.
[0022] Seventh aspect, embodiments of the present application provide a chip, including: a processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the chip implements the method described in the above first aspect or any possible implementation manner of the first aspect.
[0023] It can be understood that any of the above provided object detection devices, chips, computer-readable media, or computer program products are used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods, and will not be elaborated here. Description of the Drawings
[0024] Figure 1 It is a schematic diagram of the 3D box provided by the embodiments of the present application;
[0025] Figure 2A It is a schematic diagram of the system architecture provided by the embodiments of the present application Figure 1 ;
[0026] Figure 2B It is the second schematic diagram of the system architecture provided by the embodiment of the present application;
[0027] Figure 2C It is the schematic diagram of the system architecture provided by the embodiment of the present application Figure 3 ;
[0028] Figure 2D It is the schematic diagram of the system architecture provided by the embodiment of the present application Figure 4 ;
[0029] Figure 2E It is the schematic diagram of the system architecture provided by the embodiment of the present application Figure 5 ;
[0030] Figure 2F It is the schematic diagram of the system architecture provided by the embodiment of the present application Figure 6 ;
[0031] Figure 3 It is the schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application;
[0032] Figure 4 It is the schematic flow chart of the object detection method provided by the embodiment of the present application Figure 1 ;
[0033] Figure 5 It is the schematic diagram of the image captured by the perception module provided by the embodiment of the present application;
[0034] Figure 6 It is the schematic diagram of the second coordinate and the third coordinate provided by the embodiment of the present application;
[0035] Figure 7 It is the second schematic flow chart of the object detection method provided by the embodiment of the present application;
[0036] Figure 8 It is the schematic diagram of any second 2D box and the first 2D box provided by the embodiment of the present application;
[0037] Figure 9 It is the schematic flow chart of the object detection method provided by the embodiment of the present application Figure 3 ;
[0038] Figure 10 It is the schematic diagram of the structure of the object detection device provided by the embodiment of the present application;
[0039] Figure 11 It is the schematic diagram of the structure of the chip provided by the embodiment of the present application. Detailed implementation manners
[0040] To facilitate the understanding of the solution of the embodiment of the present application, first, various coordinate systems involved in the embodiment of the present application are introduced:
[0041] 1. Coordinate system of the first object
[0042] The coordinate system of the first object is a coordinate system with the first object as a reference, and it is a three-dimensional coordinate system. Exemplarily, taking the first object as a vehicle, the coordinate system of the first object is a coordinate system with the center of mass of the vehicle as the origin. Among them, the center of mass of the vehicle is the center of the vehicle's mass. Further, in the coordinate system of the first object, the coordinates of the vehicle can be represented by the coordinates of the center of mass of the vehicle.
[0043] 2. Coordinate system of the second object
[0044] The coordinate system of the second object is the coordinate system corresponding to the 3D box, and it is a three-dimensional coordinate system. The 3D box is a 3D model of the second object established by the target detection device according to the type information of the second object. The origin of the coordinate system of the second object can be any point on the 3D box. For example, as Figure 1 shown, the origin of the coordinate system of the second object can be the lower demarcation point 102 on the demarcation line between the left side and the rear side of the 3D box 101.
[0045] In some embodiments, there is a mapping relationship between the points in the coordinate system of the first object and the points in the coordinate system of the second object. Exemplarily, the points (x, y, z) in the coordinate system of the first object and the points (X, Y, Z) in the coordinate system of the second object satisfy the following formula: where T δ is the translation vector from the coordinate system of the second object to the coordinate system of the first object. T δ can be a three-dimensional vector. For example, T δ =(δ x , δ y , δ z ).
[0046] 3. Coordinate system of the camera
[0047] The coordinate system of the camera is a coordinate system with the optical center of the camera as the origin, and it is a three-dimensional coordinate system. The camera can be a module in the target detection device or may not be included in the target detection device.
[0048] In some embodiments, there is a mapping relationship between the points in the coordinate system of the camera and the points in the coordinate system of the second object. There is a mapping relationship between the points in the coordinate system of the camera and the points in the coordinate system of the first object.
[0049] Exemplarily, the points (B, C, D) in the coordinate system of the camera and the points (X, Y, Z) in the coordinate system of the second object satisfy the following formula: The points (B, C, D) in the coordinate system of the camera and the points (x, y, z) in the coordinate system of the first object satisfy the following formula:
[0050] Among them, [R|T] is the external parameter matrix of the camera. For the specific introduction of the external parameter matrix of the camera and the method for obtaining the external parameter matrix of the camera, reference can be made to the explanations and descriptions in the conventional technology. T δ The introduction of can be referred to the introduction in the coordinate system of the second object above.
[0051] 4. Image coordinate system
[0052] The image coordinate system is the coordinate system corresponding to the image captured by the camera and is a two-dimensional coordinate system. For example, the image coordinate system is a coordinate system established on the image with the center of the image captured by the camera as the origin.
[0053] In some embodiments, there is a mapping relationship between the points in the image coordinate system and the points in the coordinate system of the second object. There is a mapping relationship between the points in the image coordinate system and the points in the coordinate system of the first object. There is a mapping relationship between the points in the image coordinate system and the points in the coordinate system of the camera.
[0054] Exemplarily, the point (a, b) in the image coordinate system and the point (X, Y, Z) in the coordinate system of the second object satisfy the following formula: The point (a, b) in the image coordinate system and the point (x, y, z) in the coordinate system of the first object satisfy the following formula: The point (a, b) in the image coordinate system and the point (B, C, D) in the coordinate system of the camera satisfy the following formula:
[0055] Among them, s is the scale factor. A is the internal parameter matrix of the camera. [R|T] is the external parameter matrix of the camera. For the specific introduction of the scale factor, the internal parameter matrix of the camera, the external parameter matrix of the camera, and the method for obtaining the internal and external parameter matrices of the camera, reference can be made to the explanations and descriptions in the conventional technology. T δ The introduction of can be referred to the introduction in the coordinate system of the second object above.
[0056] Next, the implementation manners of the embodiments of the present application will be described in detail with reference to the accompanying drawings.
[0057] The object detection method and device provided by the embodiments of the present application can be applied to any scenario that requires detecting the 3D information of the target object. For example, the object detection method and device can be applied to ADAS of vehicles or drones, etc. Through the object detection method and device provided by the embodiments of the present application, the 3D information of the target object can be obtained, and the calculation result is accurate and the operation cost is low.
[0058] First, the system architecture applicable to the embodiments of the present application will be described.
[0059] In a possible implementation, the system architecture to which the embodiments of the present application can be applied includes a target detection device. Among them, a sensing module is deployed in the target detection device. The sensing module may include a camera. For example, a monocular camera, a binocular camera, a trinocular camera, or a multiocular camera, etc. The specific introductions of the monocular camera, binocular camera, trinocular camera, or multiocular camera can refer to the explanations in the conventional technology, and the embodiments of the present application will not elaborate. Exemplarily, the system architecture may be as Figure 2A shown. Figure 2A The shown system architecture includes a target detection device 201. A sensing module 2011 is deployed in the target detection device 201.
[0060] The above-mentioned target detection device can capture an image including a second object through the sensing module. The target detection device can also obtain the two-dimensional (2 dimensions, 2D) information of the second object in the image; obtain the 3D information of the second object according to the 2D information of the second object; obtain a first transformation relationship according to the coordinate information of the endpoints of the first dividing line and the coordinate information of the endpoints of the second dividing line; and obtain the distance between the second object and the first object according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the first transformation relationship. Specifically, reference can be made to the following Figure 4 shown method.
[0061] Optionally, the above-mentioned system architecture further includes a detection module. A neural network model is deployed in the detection module. By inputting the above-mentioned image into the neural network model, the 2D information of the second object can be obtained. For example, the neural network model includes a picture preprocessing module and a network inference module, etc. Among them, the picture preprocessing module is used to perform a normalization operation on the collected image to obtain a model with stronger generalization ability. The network inference module is used to obtain the 2D information of the second object in the image according to the image after the normalization operation, such as the coordinate information of the first 2D box, the coordinate information of the endpoints of the first dividing line, the type information of the second object, and the surface information of the second object, etc. The introductions of the coordinate information of the first 2D box, the coordinate information of the endpoints of the first dividing line, the type information of the second object, and the surface information of the second object can be referred to in the following Figure 4 shown method.
[0062] It can be understood that the above-mentioned detection module may be included in the target detection device or may be independent of the target detection device. When the detection module is independent of the target detection device, it can communicate with the target detection device through wire or wireless. Exemplarily, the system architecture may be as Figure 2B or Figure 2C shown. Figure 2BThe system architecture shown includes a target detection device 201. A sensing module 2011 and a detection module 2012 are deployed in the target detection device 201. Figure 2C The system architecture shown includes a target detection device 201 and a detection module 202. A sensing module 2011 is deployed in the target detection device 201.
[0063] It can be understood that when the detection module is included in the target detection device, the target detection device can capture an image including a second object through the sensing module and detect the 2D information of the second object in the image through the detection module. When the detection module is independent of the target detection device, the target detection device can capture an image including a second object through the sensing module and send the image to the detection module. After receiving the image, the detection module detects the 2D information of the second object in the image and sends the 2D information of the second object to the target detection device.
[0064] In another possible implementation, the system architecture applicable to the embodiments of the present application includes a target detection device and a sensing module. Among them, the sensing module and the target detection device are independent of each other. The sensing module can communicate with the target detection device in a wired or wireless manner. The sensing module may include a camera. Exemplarily, the system architecture may be as Figure 2D shown. Figure 2D The system architecture shown includes a target detection device 203 and a sensing module 204.
[0065] The above sensing module can capture an image including a second object and send the image to the target detection device. After receiving the image, the target detection device can obtain the 2D information of the second object in the image; obtain the 3D information of the second object according to the 2D information of the second object; obtain a first transformation relationship according to the coordinate information of the endpoints of the first dividing line and the coordinate information of the endpoints of the second dividing line; obtain the distance between the second object and the first object according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the first transformation relationship. Specifically, reference can be made to the Figure 4 method shown below.
[0066] Optionally, the above system architecture further includes a detection module. A neural network model is deployed in the detection module. By inputting the above image into the neural network model, the 2D information of the second object can be obtained.
[0067] It can be understood that the above detection module may be included in the target detection device or may be independent of the target detection device. When the detection module is independent of the target detection device, it can communicate with the target detection device in a wired or wireless manner. Exemplarily, the system architecture may be as Figure 2E or Figure 2F shown. Figure 2EThe system architecture shown includes a target detection device 203 and a perception module 204. A detection module 2031 is deployed in the target detection device 203. Figure 2F The system architecture shown includes a target detection device 203, a perception module 204, and a detection module 205.
[0068] It can be understood that when the detection module is included in the target detection device, the perception module can capture an image including a second object and send the image to the target detection device. After receiving the image, the target detection device can detect the 2D information of the second object in the image through the detection module. When the detection module is independent of the target detection device, the perception module can capture an image including a second object and send the image to the detection module. After receiving the image, the detection module can detect the 2D information of the second object in the image and send the 2D information of the second object to the target detection device.
[0069] It can be understood that the first object in this application can be a vehicle, a drone, or an intelligent device (for example, robots in various application scenarios, such as home robots, industrial scenario robots, etc.). A perception module, and / or a target detection device, and / or a detection module can be deployed on the first object. For example, a perception module, and / or a target detection device, and / or a detection module is deployed in the ADAS of the first object.
[0070] It can be understood that the second object in this application can be a vehicle, a guardrail, a road pile, a building, etc.
[0071] It should be noted that Figures 2A - 2F The system architecture shown is only for illustration and is not used to limit the technical solutions of this application. Those skilled in the art should understand that in the specific implementation process, the system architecture may further include other devices, and at the same time, the number of the perception module, the target detection device, or the detection module can also be determined according to specific needs.
[0072] Optionally, in the embodiments of this application Figures 2A - 2F the target detection device can be a functional module within a device. It can be understood that the above functions can be either electronic components in a hardware device, such as a chip of an ADAS, or software functions running on dedicated hardware, or virtualized functions instantiated on a platform (for example, a cloud platform).
[0073] For example, Figures 2A - 2F the target detection devices in Figure 3 can all be implemented by the electronic device 300 in Figure 3 The figure shows a schematic hardware structure diagram of an electronic device applicable to the embodiments of this application. The electronic device 300 includes at least one processor 301, a communication line 302, and a memory 303.
[0074] The processor 301 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present application.
[0075] The communication line 302 may include a path for transmitting information between the above components, such as a bus.
[0076] The memory 303 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may exist independently and be connected to the processor through the communication line 302. The memory may also be integrated with the processor. The memory provided in the embodiments of the present application generally has non-volatility. Among them, the memory 303 is used to store the computer execution instructions involved in executing the solution of the present application, and is controlled by the processor 301 to execute. The processor 301 is used to execute the computer execution instructions stored in the memory 303, so as to implement the method provided in the embodiments of the present application.
[0077] Optionally, the computer execution instructions in the embodiments of the present application may also be referred to as application code, and the embodiments of the present application do not make specific limitations thereto.
[0078] Optionally, the electronic device 300 further includes a communication interface 304. The communication interface 304 may use any transceiver-like device for communicating with other devices or communication networks, such as an Ethernet interface, a radio access network (RAN) interface, a wireless local area networks (WLAN) interface, etc.
[0079] Optionally, the electronic device 300 further includes a sensing module ( Figure 3 not shown in the figure). The sensing module may include a monocular camera, a binocular camera, a trinocular camera, or a multiocular camera. The sensing module may be used to capture an image including a second object.
[0080] Optionally, the electronic device 300 further includes a detection module ( Figure 3 not shown in the figure). A neural network model is deployed in the detection module. By inputting the captured image into the neural network model, 2D information of the second object can be obtained.
[0081] In a specific implementation, as an embodiment, the processor 301 may include one or more CPUs, such as Figure 3 CPU0 and CPU1 shown in the figure.
[0082] In a specific implementation, as an embodiment, the electronic device 300 may include multiple processors, such as Figure 3 processor 301 and processor 307 shown in the figure. Each of these processors may be a single-CPU processor or a multi-CPU processor. Here, the processor may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0083] In a specific implementation, as an embodiment, the electronic device 300 may further include an output device 305 and an input device 306. The output device 305 communicates with the processor 301 and can display information in various ways. For example, the output device 305 may be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 306 communicates with the processor 301 and can receive user input in various ways. For example, the input device 306 may be a mouse, a keyboard, a touch screen device, or a sensing device, etc.
[0084] Those skilled in the art can understand that Figure 3 the hardware structure shown in the figure does not constitute a limitation on the target detection device. The target detection device may include more or fewer components than shown in the figure, or combine certain components, or have a different component layout.
[0085] Next, in combination with Figures 1 - 3 , taking the second object as a vehicle as an example, the target detection method provided by the embodiments of the present application will be specifically described.
[0086] It should be noted that the object detection method provided by the embodiments of the present application can be applied to multiple fields, such as: the field of driverless, the field of autonomous driving, the field of assisted driving, the field of intelligent driving, the field of connected driving, the field of intelligent connected driving, the field of car sharing, etc.
[0087] It should be noted that the information names or the names of the parameters in the information in the following embodiments of the present application are only examples, and in specific implementations, they can also be other names. The embodiments of the present application do not make specific limitations on this.
[0088] It should be noted that in the description of the present application, words such as "first" or "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order. The "first 2D box" and other 2D boxes with different numbers in the present application are only for the convenience of writing in the context. The different sequence numbers themselves do not have specific technical meanings. For example, the first 2D box, the second 2D box, etc. can be understood as one or any one of a series of 2D boxes.
[0089] It should be noted that in the following embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the following embodiments of the present application should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way.
[0090] It can be understood that in the embodiments of the present application, the same step or steps with the same function or messages can be mutually referenced and learned between different embodiments.
[0091] It can be understood that in the embodiments of the present application, the object detection device can execute some or all of the steps in the embodiments of the present application. These steps are only examples, and the embodiments of the present application can also execute other steps or various deformations of the steps. In addition, each step can be executed in a different order presented in the embodiments of the present application, and it is possible not to execute all the steps in the embodiments of the present application.
[0092] In the embodiments of the present application, the specific structure of the execution subject of the object detection method is not particularly limited in the embodiments of the present application. As long as it can run a program recording the code of the object detection method of the embodiments of the present application to perform corresponding operations according to the object detection method of the embodiments of the present application. For example, the execution subject of the object detection method provided by the embodiments of the present application can be an object detection device, or a component applied to the object detection device, such as a chip. The present application does not make limitations on this.
[0093] Such as Figure 4As shown, a target detection method provided by an embodiment of the present application, the target detection method includes steps 401 - step 404.
[0094] Step 401: The target detection device acquires 2D information of a second object in the image acquired by a first object.
[0095] Wherein, the target detection device may be Figures 2A - 2F any of the target detection devices shown in. The target detection device may be deployed in the first object. It should be understood that if the target detection device is Figure 2D , Figure 2E or Figure 2F the target detection device in, a sensing module is also deployed in the first object, for example, a camera.
[0096] Wherein, the image is acquired by the sensing module. The sensing module may include a monocular camera, a binocular camera, a trinocular camera or a multi - camera. The image in step 401 may be taken by the sensing module in the first object. The sensing module may be Figures 2A - 2F the sensing module in. For example, if the target detection device is Figure 2A the target detection device 201 in, then the sensing module is Figure 2A the sensing module 2011 in. If the target detection device is Figure 2E the target detection device 203 in, then the sensing module is Figure 2E the sensing module 204 in.
[0097] Wherein, the image acquired by the first object includes one or more second objects. For example, the image may be as Figure 5 shown, the image includes multiple second objects.
[0098] Wherein, the 2D information of the second object includes the coordinate information of the endpoints of the first dividing line of the second object and the type information of the second object. Further, the 2D information of the second object also includes the coordinate information of the first 2D box of the second object and the face information of the second object.
[0099] Wherein, the coordinate information of the first 2D box is used to indicate the coordinates of the second object in the image coordinate system corresponding to the image. As mentioned above, the image coordinate system is a two - dimensional coordinate system. Therefore, the first 2D box is a planar figure, such as a rectangle or a polygon, etc. The first 2D box is a polygon that encloses the second object in the image. Exemplarily, the first 2D box may be Figure 5The 2D box 501 in []. The coordinate information of the first 2D box may include the coordinates of each corner in the first 2D box. For example, if the first 2D box is a rectangle, the coordinate information of the first 2D box includes the coordinates of the four corners of the rectangle. It should be understood that in addition to including the coordinates of each corner in the first 2D box, the coordinate information of the first 2D box may also indicate the coordinates of the second object in the image coordinate system in other ways, without limitation.
[0100] Among them, the first dividing line may be the dividing line between the first side and the second side of the second object. At least one of the first side and the second side is included in the first 2D box. Exemplarily, the first dividing line is the dividing line between the rear and the left side of the second object, or the dividing line between the rear and the right side of the second object, or the dividing line between the front and the left side of the second object, or the dividing line between the front and the right side of the second object, etc. For example, the first dividing line may be Figure 5 the dividing line 502 in []. The dividing line 502 is the dividing line between the rear and the left side of the second object.
[0101] Among them, the coordinate information of the endpoints of the first dividing line is used to indicate the coordinates of the first dividing line in the image coordinate system corresponding to the image. Exemplarily, the coordinate information of the endpoints of the first dividing line includes: the coordinates of the lower dividing point of the first dividing line, and / or, the coordinates of the upper dividing point of the first dividing line. Among them, the lower dividing point of the first dividing line is the intersection point of the first dividing line and the lower boundary of the first 2D box. The upper dividing point of the first dividing line is the intersection point of the first dividing line and the upper boundary of the first 2D box.
[0102] It can be understood that in practical applications, the coordinates of other points on the first dividing line may also be used to indicate the coordinates of the first dividing line in the image coordinate system corresponding to the image. For example, the coordinates of the midpoint of the first dividing line. The midpoint of the first dividing line is the midpoint of the line connecting the upper dividing point and the lower dividing point of the first dividing line.
[0103] Among them, the type information of the second object is used to indicate the type to which the second object belongs. The types of the second object include one or more of the following types: hatchback car, sedan, minicar, sport utility vehicle (SUV), pickup truck, minivan, small truck (open type), small truck (closed type), light truck, heavy truck, engineering vehicle, medium bus, large bus or double-decker bus. It should be understood that the above types are only examples of the types of the second object, and in practical applications, the types of the second object also include other types, without limitation.
[0104] The type information of the second object may include the identifier of the type to which the second object belongs. Exemplarily, taking the second object as a hatchback car and the identifier of the hatchback car as ID 2, the type information of the second object includes ID 2.
[0105] Among them, the surface information of the second object is used to indicate the above-mentioned first surface and / or second surface. Exemplarily, the surface information of the second object may include the identifier of the first surface and / or the identifier of the second surface. For example, taking the dividing line 502 in Figure 5 as the dividing line, and the identifier of the left surface of the second object being ID 1 and the identifier of the rear surface of the second object being ID 2 as an example, the surface information of the second object includes ID 1 and ID 2.
[0106] It should be noted that when the included angle between the target detection device and the traveling direction of the second object is less than or equal to the first threshold, in the image, one surface of the second object that can be seen. As Figure 5 shown, the second object 503 can be seen from the rear in the Figure 5 image shown, but the side surface cannot be seen. In this case, the surface information of the second object is used to indicate this one surface, and the first dividing line is the left or right boundary of the first 2D box.
[0107] It can be understood that when there are multiple second objects in the image, the target detection device can obtain the 2D information of one or more second objects among the multiple second objects in the image. Further, when the target detection device obtains the 2D information of multiple second objects in the image, the target detection device can obtain the 2D information of the multiple second objects simultaneously, or can obtain the 2D information of each second object among the multiple second objects one by one.
[0108] Exemplarily, the target detection device can obtain the 2D information of the second object in the image in the following two ways.
[0109] One possible implementation manner is that the target detection device obtains the 2D information of the second object in the image according to the user's input.
[0110] Exemplarily, taking the Figure 2A shown system architecture as an example, the target detection device 201 captures an image including the second object through the sensing module 2011. The target detection device 201 displays the image for the user through the human-computer interaction interface and receives the 2D information of the second object input by the user.
[0111] Exemplarily, taking the Figure 2D shown system architecture as an example, the sensing module 204 captures an image including the second object and sends the image to the target detection device 203. After receiving the image from the sensing module 204, the target detection device 203 displays the image for the user through the human-computer interaction interface and receives the 2D information of the second object input by the user.
[0112] In another possible implementation, the target detection device inputs the image into a neural network model to obtain the 2D information of the second object. For example, the target detection device inputs the image into the neural network model in the detection module to obtain the 2D information of the second object. Herein, the detection module can be the Figures 2A - 2F detection module therein. For example, if the target detection device is the Figure 2B target detection device 201 therein, then this detection module is the Figure 2B detection module 2012 therein. If the target detection device is the Figure 2F target detection device 203 therein, then this detection module is the Figure 2F detection module 205 therein. The introduction of the neural network model can refer to the corresponding description in the above introduction of the system architecture.
[0113] Exemplarily, taking the Figure 2B shown system architecture as an example, the target detection device 201 captures an image including the second object through the sensing module 2011. The target detection device 201 inputs the image into the neural network model in the detection module 2012 to obtain the 2D information of the second object.
[0114] Exemplarily, taking the Figure 2C shown system architecture as an example, the target detection device 201 captures an image including the second object through the sensing module 2011 and sends the image to the detection module 202. After receiving the image, the detection module 202 inputs the image into the neural network model to obtain the 2D information of the second object and sends the 2D information of the second object to the target detection device 201.
[0115] Exemplarily, taking the Figure 2E shown system architecture as an example, the sensing module 204 captures an image including the second object and sends the image to the target detection device 203. After receiving the image from the sensing module 204, the target detection device 203 inputs the image into the neural network model to obtain the 2D information of the second object.
[0116] Exemplarily, taking the Figure 2F shown system architecture as an example, the sensing module 204 captures an image including the second object and sends the image to the detection module 205. After receiving the image, the detection module 205 inputs the image into the neural network model to obtain the 2D information of the second object and sends the 2D information of the second object to the target detection device 203.
[0117] Step 402: The target detection device obtains the 3D information of the second object according to the 2D information of the second object.
[0118] Among them, the 3D information of the second object includes the coordinate information of the endpoints of the second demarcation line. Further, the 3D information of the second object further includes the coordinate information of the 3D bounding box of the second object.
[0119] Among them, the 3D bounding box is a 3D model of the second object established according to the type information of the second object. Exemplarily, the 3D bounding box can be a three-dimensional figure, such as a cuboid. For example, the 3D bounding box can be as Figure 1 shown. The coordinate information of the 3D bounding box is used to indicate the coordinates of the 3D bounding box in the coordinate system corresponding to the 3D bounding box. The coordinate system corresponding to the 3D bounding box can also be referred to as the coordinate system of the second object. The coordinate information of the 3D bounding box can include the coordinates of each corner in the 3D bounding box. For example, if the 3D bounding box is a cuboid, the coordinate information of the 3D bounding box includes the coordinates of the eight corners of the cuboid. It should be understood that in addition to including the coordinates of each corner in the 3D bounding box, the coordinate information of the 3D bounding box can also indicate the coordinates of the 3D bounding box in the coordinate system corresponding to the 3D bounding box in other ways, without limitation.
[0120] Among them, the second demarcation line corresponds to the first demarcation line. That is to say, the second demarcation line is the demarcation line between the first face and the second face in the 3D bounding box of the second object. For example, if the first demarcation line is the demarcation line between the back and the left side of the second object in the image, then the second demarcation line is the demarcation line between the back and the left side of the 3D bounding box; if the first demarcation line is the demarcation line between the back and the right side of the second object in the image, then the second demarcation line is the demarcation line between the back and the right side of the 3D bounding box; if the first demarcation line is the demarcation line between the front and the left side of the second object in the image, then the second demarcation line is the demarcation line between the front and the left side of the 3D bounding box; if the first demarcation line is the demarcation line between the front and the right side of the second object in the image, then the second demarcation line is the demarcation line between the front and the right side of the 3D bounding box. For example, the second demarcation line can be Figure 1 the demarcation line 103 in. The demarcation line 103 is the demarcation line between the back and the left side of the 3D bounding box. Among them, the front or back of the 3D bounding box is the face composed of the height and the width of the 3D bounding box, and the left or right side of the 3D bounding box is the face composed of the height and the length of the 3D bounding box.
[0121] Among them, the coordinate information of the endpoints of the second demarcation line is used to indicate the coordinates of the second demarcation line in the coordinate system corresponding to the 3D bounding box. Exemplarily, the coordinate information of the endpoints of the second demarcation line includes: the coordinates of the lower demarcation point of the second demarcation line, and / or, the coordinates of the upper demarcation point of the second demarcation line. Among them, the lower demarcation point of the second demarcation line is the intersection point of the second demarcation line and the lower plane of the 3D bounding box. The upper demarcation point of the second demarcation line is the intersection point of the second demarcation line and the upper plane of the 3D bounding box.
[0122] It can be understood that in practical applications, the coordinates of other points on the second dividing line can also be used to indicate the coordinates of the second dividing line in the coordinate system corresponding to the 3D box. For example, the coordinates of the midpoint of the second dividing line. The midpoint of the second dividing line is the midpoint of the line connecting the upper dividing point and the lower dividing point of the second dividing line.
[0123] The target detection device can obtain the 3D information of the second object through the following two exemplary methods.
[0124] Method 1: The 2D information of the second object includes the coordinate information of the endpoints of the first dividing line, the coordinate information of the first 2D box, and the type information of the second object. The target detection device constructs a 3D box for the second object according to the type information of the second object to obtain the coordinate information of the 3D box; the target detection device determines the first plane and the second plane according to the coordinate information of the first 2D box and the coordinate information of the endpoints of the first dividing line; the target detection device obtains the coordinate information of the endpoints of the second dividing line according to the first plane, the second plane, and the coordinate information of the 3D box.
[0125] In a possible implementation, the length of the boundary of the 3D box corresponds to the type information of the second object. Further, the length, width, and height of the 3D box correspond to the type information of the second object.
[0126] Exemplarily, taking the types of the second object including hatchback cars, sedan cars, and SUVs as an example, the corresponding relationship between the length, width, and height of the 3D box and the type information of the second object can be shown in Table 1. As shown in Table 1, if the type of the second object is a hatchback car, the length of the 3D box is L 1 , the width of the 3D box is W 1 , and the height of the 3D box is H 1 . If the type of the second object is a sedan car, the length of the 3D box is L 2 , the width of the 3D box is W 2 , and the height of the 3D box is H 2 . If the type of the second object is an SUV, the length of the 3D box is L 3 , the width of the 3D box is W 3 , and the height of the 3D box is H 3 .
[0127] Table 1
[0128] Type of the second object Length of the 3D box Width of the 3D box Height of the 3D box Two - box car <![CDATA[L 1 > <![CDATA[W 1 > <![CDATA[H 1 > Three - box car <![CDATA[L 2 > <![CDATA[W 2 > <![CDATA[H 2 > SUV <![CDATA[L 3 > <![CDATA[W 2 > <![CDATA[H 3 >
[0129] It can be understood that the above Table 1 is only an example of the corresponding relationship between the length, width, and height of the 3D box and the type information of the second object. In practical applications, the corresponding relationship between the length, width, and height of the 3D box and the type information of the second object can also be in other forms, without limitation.
[0130] It is understandable that if the distance between the first dividing line and the left or right boundary of the first 2D frame is greater than or equal to the second threshold, at least two faces of the second object can be seen in the image. If the distance between the first dividing line and the left or right boundary of the first 2D frame is less than the second threshold, one face of the second object can be seen in the image.
[0131] Exemplarily, taking the case where the distance between the first dividing line and the left or right boundary of the first 2D frame is greater than or equal to the second threshold as an example, if the second object has the same driving direction as the target detection device and the first 2D frame is on the right side of the image, the target detection device determines that the first face is the left face and the second face is the rear face, or the first face is the rear face and the second face is the left face. Subsequently, the target detection device determines the dividing line between the left face and the rear face of the 3D frame as the second dividing line and obtains the coordinate information of the endpoints of the second dividing line. If the second object has the opposite driving direction to the target detection device and the first 2D frame is on the right side of the image, the target detection device determines that the first face is the right face and the second face is the front face, or the first face is the front face and the second face is the right face. Subsequently, the target detection device determines the dividing line between the right face and the front face of the 3D frame as the second dividing line and obtains the coordinate information of the endpoints of the second dividing line.
[0132] Exemplarily, taking the case where the distance between the first dividing line and the left or right boundary of the first 2D frame is less than the second threshold as an example, if the second object has the same driving direction as the target detection device, the target detection device determines that the face shown in the image is the rear face of the second object. If the distance between the first dividing line and the left boundary of the first 2D frame is less than the second threshold, the target detection device determines the dividing line between the rear face and the left face of the 3D frame as the second dividing line and obtains the coordinate information of the endpoints of the second dividing line; if the distance between the first dividing line and the right boundary of the first 2D frame is less than the second threshold, the target detection device determines the dividing line between the rear face and the right face of the 3D frame as the second dividing line and obtains the coordinate information of the endpoints of the second dividing line.
[0133] Method 2: The 2D information of the second object includes the coordinate information of the endpoints of the first dividing line, the type information of the second object, and the face information of the second object. The target detection device constructs the 3D frame of the second object according to the type information of the second object to obtain the coordinate information of the 3D frame; the target detection device obtains the coordinate information of the endpoints of the second dividing line according to the face information of the second object and the coordinate information of the 3D frame.
[0134] Among them, the specific process of the target detection device constructing the 3D frame of the second object according to the type information of the second object to obtain the coordinate information of the 3D frame can refer to that described in the above Method 1 and will not be elaborated here.
[0135] In a possible implementation, the target detection device determines a second demarcation line based on the first surface and / or the second surface indicated in the surface information of the second object, and obtains the coordinate information of the endpoints of the second demarcation line according to the coordinate information of the 3D box.
[0136] Exemplarily, taking the surface information of the second object including the identifier of the left surface and the identifier of the rear surface of the second object as an example, the target detection device determines the demarcation line between the left surface and the rear surface of the 3D box as the second demarcation line, and obtains the coordinate information of the endpoints of the second demarcation line according to the coordinate information of the 3D box.
[0137] Exemplarily, taking the surface information of the second object including the identifier of the rear surface of the second object as an example, if the first demarcation line is the left boundary of the first 2D box, the target detection device determines the demarcation line between the left surface and the rear surface of the 3D box as the second demarcation line, and obtains the coordinate information of the endpoints of the second demarcation line according to the coordinate information of the 3D box. If the first demarcation line is the right boundary of the first 2D box, the target detection device determines the demarcation line between the right surface and the rear surface of the 3D box as the second demarcation line, and obtains the coordinate information of the endpoints of the second demarcation line according to the coordinate information of the 3D box.
[0138] Step 403: The target detection device obtains a first transformation relationship according to the coordinate information of the endpoints of the first demarcation line and the coordinate information of the endpoints of the second demarcation line.
[0139] Wherein, the first transformation relationship is the transformation relationship between the coordinate system corresponding to the 3D box and the coordinate system of the first object when the distance between the second object and the first object is K. K > 0. For example, the first transformation relationship is the transformation relationship between the coordinate system corresponding to the 3D box and the coordinate system of the first object when the distance between the second object and the optical center of the camera of the first object is K.
[0140] In a possible implementation, any point on the second demarcation line and the mapped point of this point in the image coordinate system satisfy Formula 1: In the case where the distance between the second object and the camera is K, any point on the above-mentioned second demarcation line and the mapped point of this point in the coordinate system of the camera satisfy Formula 2: The target detection device combines Formula 1 and Formula 2 above to solve and can obtain s and the first transformation relationship, that is, s and T δ .
[0141] Wherein, (X, Y, Z) is any point on the second demarcation line, for example, the lower demarcation point of the second demarcation line. (a, b) is the mapped point of any point on the second demarcation line in the image coordinate system. If (X, Y, Z) is the lower demarcation point of the second demarcation line, then (a, b) is the lower demarcation point of the first demarcation line. A and [R|T] are known quantities. - indicates that when the target detection device solves s and the first transformation relationship, this value can be ignored.
[0142] Step 404: The target detection device obtains the distance between the second object and the first object according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the first transformation relationship.
[0143] In a possible implementation, the target detection device transforms the first coordinate through the first transformation relationship, the internal parameter matrix and the external parameter matrix to obtain the second coordinate; the target detection device performs a check operation on the second coordinate, the third coordinate and K to obtain the distance between the second object and the first object. Among them, the internal parameter matrix is the internal parameter matrix of the camera, that is, the internal parameter matrix is A. The external parameter matrix is the external parameter matrix of the camera, that is, the external parameter matrix is [R|T].
[0144] The first coordinate is on the second dividing line. Because the first transformation relationship is the transformation relationship between the coordinate system corresponding to the 3D frame and the coordinate system of the first object when the distance between the second object and the first object is K, the distance between the second coordinate and the camera is K. The second coordinate, the third coordinate and the camera are on a straight line. The third coordinate is the coordinate on the first dividing line corresponding to the first coordinate.
[0145] Furthermore, when the distance between the second object and the first object is K, the first coordinate (X, Y, Z) and the second coordinate (x, y) satisfy the following formula: Among them, s and T δ is the value calculated in the above step 403. The second coordinate (x, y) can be obtained by the above formula.
[0146] For example, the distance between the optical center of the second object and the first object is K, and the first coordinate is the coordinate of the lower dividing point of the second dividing line. Figure 6 As shown, point A on the first object 601 is the position of the optical center of the camera, the coordinate of point C is the second coordinate, and the coordinate of point E is the third coordinate, which is also the coordinate of the lower dividing point of the first dividing line. Point A, point C and point E are on a straight line. Point D is the upper dividing point of the first dividing line. Triangle ABC and triangle ADE are constructed based on point A, point C, point E and point D. Triangle ABC and triangle ADE are similar triangles. Therefore, the target detection device performs a composite operation on the second coordinate, the third coordinate and K according to the properties of similar triangles, and can obtain the distance between the second object and the first object.
[0147] For example, if the coordinates of point D are (x 1 ,y 1 ), the coordinates of point E are (x 1 ,y 2 ), the coordinates of point C are (x 2 ,y 3 ), we can infer that the coordinates of point B are (x 2 ,y1 )。It can be known from the properties of similar triangles that: Therefore, From this, the P value can be obtained. Subsequently, the target detection device can use the P value as the distance between the second object and the first object. Further, the target detection device can also obtain the distance between the centroid of the second object and the camera based on the P value and the coordinate information of the 3D box, and use this distance as the distance between the second object and the first object.
[0148] Optionally, after step 404, the target detection device can perform path planning based on the distance between the second object and the first object, and control the first object to travel along the planned path, so as to effectively avoid obstacles and increase the comfort and safety of autonomous driving.
[0149] Based on Figure 4 the method shown, the target detection device can obtain the 2D information of the second object in the image obtained by the first object, for example, the coordinate information of the endpoints of the first dividing line of the second object and the type information of the second object. The target detection device can also obtain the 3D information of the second object, for example, the coordinate information of the endpoints of the second dividing line of the second object. The target detection device can also obtain the first transformation relationship between the coordinate system corresponding to the 3D box and the coordinate system of the first object when the distance between the second object and the first object is K based on the coordinate information of the endpoints of the first dividing line and the coordinate information of the endpoints of the second dividing line, and obtain the distance between the second object and the first object based on the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the first transformation relationship. In this way, the target detection device constructs a 3D box based on the image collected by the monocular camera, and can obtain the distance between the second object and the first object by using the transformation relationship between the coordinate system corresponding to the 3D box and the coordinate system of the first object when the distance between the second object and the first object is K, and the calculation cost is relatively low. In addition, through Figure 4 the method shown, the distance between the second object and the first object can be obtained without using a multi-camera or a depth camera, so the cost of the target detection device can be reduced.
[0150] Optionally, the target detection device can also obtain the orientation angle of the second object, that is, the included angle between the second object and the driving direction of the first object. Specifically, as Figure 7 shown, Figure 4 the method shown also includes steps 701 - step 703.
[0151] Step 701: The target detection device rotates the 3D box N times with the second dividing line as the center, and obtains the corresponding second 2D box of the 3D box after each rotation.
[0152] Among them, N is a positive integer.
[0153] In a possible implementation, the target detection device rotates the 3D box centered on the second dividing line, and obtains the corresponding second 2D box of the rotated 3D box every rotation angle α. Where 0 ≤ α ≤ 360°.
[0154] Exemplarily, taking α as 60° as an example, the target detection device needs to obtain the corresponding second 2D boxes of the rotated 3D boxes when the 3D box is rotated by 60°, 120°, 180°, 240°, 300° and 360° respectively.
[0155] In another possible implementation, the N is predefined or randomly determined by the target detection device. Among the N rotation angles obtained by rotating the 3D box N times, the difference between adjacent two rotation angles can be the same or different.
[0156] Exemplarily, taking N as 7, and the 7 rotation angles being 0°, 30°, 90°, 150°, 200°, 260° and 310° respectively as an example, the target detection device needs to obtain the corresponding second 2D boxes of the rotated 3D boxes when the 3D box is rotated by 0°, 30°, 90°, 150°, 200°, 260° and 310° respectively.
[0157] In a possible implementation, the second 2D box is obtained by performing coordinate transformation on the rotated 3D box. Further, the second 2D box is obtained by performing coordinate transformation on the rotated 3D box through the second transformation relationship, the internal parameter matrix and the external parameter matrix.
[0158] Wherein, the second transformation relationship is the transformation relationship between the coordinate system corresponding to the 3D box and the coordinate system of the first object. The coordinate information of the second 2D box is used to indicate the coordinates of the second 2D box in the image coordinate system.
[0159] Further, any point (x, y) on the second 2D box and its mapped point (X, Y, Z) on the rotated 3D box satisfy the following formula: Wherein, (X, Y, Z), s, A, [R|T] and T δ ' are known quantities, and thus (x, y) can be obtained.
[0160] The above T δ ' is the second transformation relationship, and the process for the target detection device to obtain the second transformation relationship is as follows:
[0161] In a possible implementation, the second transformation relationship is obtained based on the coordinate information of the endpoints of the first demarcation line, the coordinate information of the endpoints of the second demarcation line, and the distance between the second object and the first object. Further, the target detection device obtains a system of equations with the second transformation relationship as the unknown based on the coordinate information of the endpoints of the first demarcation line, the coordinate information of the endpoints of the second demarcation line, and the distance between the second object and the first object, and solves the system of equations to obtain the second transformation relationship.
[0162] Exemplarily, any point on the second demarcation line and its mapped point in the image coordinate system satisfy Equation 3: Any point on the above-mentioned second demarcation line and its mapped point in the coordinate system of the camera satisfy Equation 4: The target detection device can solve for s and the second transformation relationship by combining Equation 3 and Equation 4 above, that is, s and T δ '.
[0163] Where (X, Y, Z) is any point on the second demarcation line, for example, the lower demarcation point of the second demarcation line. (a, b) is the mapped point of any point on the second demarcation line in the image coordinate system. If (X, Y, Z) is the lower demarcation point of the second demarcation line, then (a, b) is the lower demarcation point of the first demarcation line. A and [R|T] are known quantities. P is the distance between the second object and the first object obtained in step 404. - indicates that the target detection device can ignore this value when solving for s and the second transformation relationship.
[0164] Step 702: The target detection device obtains, according to the N second 2D frames corresponding to the 3D frame after N rotations and the first 2D frame, the second 2D frame with the shortest distance between its boundary and the boundary of the first 2D frame among the N second 2D frames.
[0165] In a possible implementation, among the N second 2D frames, the sum of the first distance and the second distance corresponding to the second 2D frame with the shortest distance is the smallest. Wherein, the first distance is the distance between the left boundary of the second 2D frame with the shortest distance and the left boundary of the first 2D frame. The second distance is the distance between the right boundary of the second 2D frame with the shortest distance and the right boundary of the first 2D frame.
[0166] Please refer to Figure 8 , Figure 8 for the schematic diagram of any second 2D frame and the first 2D frame. Figure 8Among them, the distance between the left boundaries of the second 2D box 801 and the first 2D box 802 is Δa, and the distance between the right boundaries of the second 2D box 801 and the first 2D box 802 is Δb. Among the N second 2D boxes, the Δa + Δb corresponding to the second 2D box with the shortest distance is the smallest. That is, the rotation angle α of the 3D box corresponding to the second 2D box with the shortest distance satisfies the following formula: g = argmin[Δa(α) + Δb(α)]. Here, argmin represents the value of α when [Δa(α) + Δb(α)] reaches the minimum value.
[0167] Exemplarily, taking Figure 8 the 3D box corresponding to the second 2D box 801 shown as Figure 1 the 3D box 101 shown as an example, if the coordinates of P 1 in the coordinate system corresponding to the 3D box are (X 1 , Y 1 , Z 1 ), and the coordinates of P 2 in the coordinate system corresponding to the 3D box are (X 2 , Y 2 , Z 2 ), in the second 2D box 801, the coordinates of the point p 1 corresponding to P 1 in the image coordinate system are (x 1 , y 1 ), the coordinates of the point p 2 corresponding to P 2 in the image coordinate system are (x 2 , y 2 ), in the first 2D box 802, the coordinates of Q 1 in the image coordinate system are (x 3 , y 3 ), and the coordinates of Q 2 in the image coordinate system are (x 4 , y 4 ), then (X 1 , Y 1 , Z 1 ) and (x 1 , y 1 ) satisfy the following formula: Among them, s, A, [R|T], T δ ', and (X 1 , Y 1 , Z 1 ) are known quantities, and (x 1 , y 1 ) can be obtained. Similarly, (X 2 , Y 2 , Z 2 ) and (x2 , y 2 ) satisfies the following formula: where s, A, [R|T], T δ ', and (X 2 , Y 2 , Z 2 ) are known quantities, and (x 2 , y 2 ) can be obtained. Then Δa = |x 1 - x 3 |, and Δb = |x 2 - x 4 |.
[0168] It can be understood that the larger the N value, the smaller the error when the target detection device calculates the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance, and the more accurate the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance obtained.
[0169] Step 703: The target detection device determines the orientation angle of the second object according to the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance.
[0170] A possible implementation is that the target detection device determines the orientation angle of the second object as the sum of the orientation angle of the 3D box corresponding to step 402 and the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance.
[0171] For example, if the orientation angle of the 3D box corresponding to step 402 is 0°, then the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance is the orientation angle of the second object. If the orientation angle of the 3D box corresponding to step 402 is 30° and the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance is 30°, then the orientation angle of the second object is 60°.
[0172] Optionally, after step 703, the target detection device can perform path planning according to the distance between the second object and the first object and the orientation angle of the second object, and control the first object to travel along the planned path, so as to effectively avoid obstacles and increase the comfort and safety of autonomous driving.
[0173] Based on Figure 7 the method shown, the target detection device can compare the second 2D box corresponding to the rotated 3D box with the first 2D box to obtain the second 2D box closest to the first 2D box, thereby obtaining the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance, and determining the orientation angle of the second object according to the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance. In this way, when the target detection device plans a path, in addition to referring to the distance between the second object and the first object, it can also refer to the orientation angle of the second object, making the reference of the planned path higher.
[0174] Understandably, the length, width, and height of the 3D box in step 402 above are obtained according to the type of the second object. In practical applications, for vehicles of the same type, the sizes of vehicles of different brands may be different. Therefore, the length, width, or height of the 3D box obtained according to the type of the second object may not be accurate.
[0175] Optionally, the target detection device can also calibrate the length and / or width of the 3D box. Specifically, as Figure 9 shown, Figure 7 the method shown also includes steps 901 - 903.
[0176] Step 901: The target detection device makes M adjustments to the length of the boundary of the 3D box corresponding to the second 2D box with the shortest distance, and obtains the third 2D box corresponding to the 3D box after each adjustment.
[0177] Where M is a positive integer. M is predefined or determined by the target detection device. The length of the boundary of the 3D box includes the length of the 3D box and / or the width of the 3D box.
[0178] Understandably, the purpose of the target detection device making M adjustments to the length of the boundary of the 3D box corresponding to the second 2D box with the shortest distance is to minimize the difference between the boundary of the third 2D box corresponding to the adjusted 3D box and the boundary of the first 2D box. In this way, a more accurate length of the boundary of the 3D box can be obtained. That is, a more accurate length of the second object and a more accurate width of the second object can be obtained.
[0179] Exemplarily, the target detection device makes M adjustments to the length and / or width of the 3D box corresponding to the second 2D box with the shortest distance. For example, the target detection device increases the length and / or width of the 3D box corresponding to the second 2D box with the shortest distance by Δj each time. Each increased Δj can be the same or different. Also for example, the target detection device decreases the length and / or width of the 3D box corresponding to the second 2D box with the shortest distance by Δj each time. Each decreased Δj can be the same or different. Δj is predefined or determined by the target detection device.
[0180] Understandably, when the target detection device adjusts the length and width of the 3D box corresponding to the second 2D box with the shortest distance each time, the adjustment values of the length and width of the 3D box corresponding to the second 2D box with the shortest distance can be the same or different. For example, in one adjustment, the target detection device can increase the length of the 3D box corresponding to the second 2D box with the shortest distance by Δj and decrease the width of the 3D box corresponding to the second 2D box with the shortest distance by Δr.
[0181] In a possible implementation, the third 2D box is obtained by performing a coordinate transformation on the adjusted 3D box. Further, the third 2D box is obtained by performing a coordinate transformation on the adjusted 3D box through a second transformation relationship, an intrinsic matrix, and an extrinsic matrix.
[0182] Further, for any point (x, y) on the third 2D box and its corresponding mapped point (X, Y, Z) on the 3D box corresponding to the third 2D box, the following formula is satisfied: where (X, Y, Z), s, A, [R|T], and T δ ' are known quantities, from which (x, y) can be obtained.
[0183] Step 902: The target detection device obtains, from the third 2D boxes corresponding to the M adjusted 3D boxes and the first 2D box, the third 2D box with the shortest distance between its boundary and the boundary of the first 2D box among the M third 2D boxes.
[0184] For example, among the N third 2D boxes, the sum of the third distance and the fourth distance corresponding to the third 2D box with the shortest distance is the smallest. The third distance is the distance between the left boundary of the third 2D box with the shortest distance and the left boundary of the first 2D box. The fourth distance is the distance between the right boundary of the third 2D box with the shortest distance and the right boundary of the first 2D box.
[0185] For another example, among the N third 2D boxes, the sum of the third distance, the fourth distance, and the fifth distance corresponding to the third 2D box with the shortest distance is the smallest. The fifth distance is the distance between the demarcation line on the third 2D box with the shortest distance and the first demarcation line.
[0186] Exemplarily, taking the lengths of the boundaries of the 3D box including the length and width of the 3D box as an example, for the above M third 2D boxes, the length l and width w of the 3D box corresponding to the third 2D box with the shortest distance satisfy the following formula: h = argmin[Δc(l, w) + Δd(l, w) + Δe(l, w)]. Where Δc is the distance between the left boundary of any third 2D box and the left boundary of the first 2D box. Δd is the distance between the right boundary of any third 2D box and the right boundary of the first 2D box. Δe is the distance between the demarcation line of any third 2D box and the first demarcation line. argmin represents the values of l and w when [Δc(l, w) + Δd(l, w) + Δe(l, w)] reaches the minimum.
[0187] Step 903: The target detection device determines the length of the boundary of the second object according to the length of the boundary of the 3D box corresponding to the third 2D box with the shortest distance.
[0188] In a possible implementation, the target detection device determines the length of the boundary of the second object as the length of the boundary of the 3D box corresponding to the third 2D box with the shortest distance.
[0189] Exemplarily, the target detection device determines the length of the 3D box corresponding to the third 2D box with the shortest distance as the length of the second object; determines the width of the 3D box corresponding to the third 2D box with the shortest distance as the width of the second object; and determines the height of the 3D box in step 402 as the height of the second object.
[0190] It can be understood that the larger the M value, the smaller the error in determining the length of the boundary of the second object by the target detection device, and the more accurate the length of the boundary of the second object obtained.
[0191] Optionally, after step 903, the target detection device can also calculate the centroid of the second object according to the length of the boundary of the second object.
[0192] Based on Figure 9 According to the method shown, the target detection device can compare the third 2D box corresponding to the adjusted 3D box with the first 2D box to obtain the third 2D box closest to the first 2D box, thereby obtaining the length of the boundary of the second object. In this way, the target detection device can obtain more accurate dimensions of the second object.
[0193] It can be understood that in order to implement the above functions, the above target detection device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm operations of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0194] The embodiments of this application can divide the function modules of the target detection device according to the above method examples. For example, each function module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software function module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0195] For example, in the case of dividing each function module in an integrated manner, Figure 10 FIG. shows a schematic structural diagram of a target detection device. This target detection device can be used to execute the functions of the target detection device involved in the above embodiments.
[0196] As a possible implementation manner,Figure 10 The target detection device shown includes: an acquisition unit 1001 and a determination unit 1002.
[0197] The acquisition unit 1001 is configured to acquire two-dimensional (2D) information of a second object in an image acquired by a first object; the 2D information of the second object includes: coordinate information of endpoints of a first demarcation line of the second object and type information of the second object, where the first demarcation line is a demarcation line between a first surface and a second surface of the second object, and at least one of the first surface and the second surface is included in a first 2D frame of the second object, and the first 2D frame is a polygon that encloses the second object in the image. For example, in combination with Figure 4 , the acquisition unit 1001 is configured to perform step 401.
[0198] The acquisition unit 1001 is further configured to acquire three-dimensional (3D) information of the second object according to the 2D information of the second object; the 3D information of the second object includes coordinate information of endpoints of a second demarcation line, where the second demarcation line is a demarcation line between the first surface and the second surface in a 3D frame of the second object, and the 3D frame is a 3D model of the second object, and the length of the boundary of the 3D frame corresponds to the type information of the second object. For example, in combination with Figure 4 , the acquisition unit 1001 is further configured to perform step 402.
[0199] The determination unit 1002 is configured to obtain a first transformation relationship according to the coordinate information of the endpoints of the first demarcation line and the coordinate information of the endpoints of the second demarcation line; the first transformation relationship is a transformation relationship between the 3D coordinate system corresponding to the 3D frame and the 3D coordinate system of the first object when the distance between the second object and the first object is K, where K > 0. For example, in combination with Figure 4 , the determination unit 1002 is configured to perform step 403.
[0200] The determination unit 1002 is further configured to obtain the distance between the second object and the first object according to the coordinate information of the endpoints of the first demarcation line, the coordinate information of the endpoints of the second demarcation line, and the first transformation relationship. For example, in combination with Figure 4 , the determination unit 1002 is further configured to perform step 404.
[0201] In a possible implementation, the 2D information of the second object further includes: coordinate information of the first 2D frame; the acquisition unit 1001 is specifically configured to construct the 3D frame according to the type information of the second object to obtain coordinate information of the 3D frame; the acquisition unit 1001 is further specifically configured to determine the first surface and the second surface according to the coordinate information of the first 2D frame and the coordinate information of the endpoints of the first demarcation line; the acquisition unit 1001 is further specifically configured to obtain the coordinate information of the endpoints of the second demarcation line according to the first surface, the second surface, and the coordinate information of the 3D frame.
[0202] A possible implementation manner, the 2D information of the second object further includes: the surface information of the second object, and the surface information of the second object is used to indicate the first surface and the second surface; the obtaining unit 1001 is specifically configured to construct the 3D box according to the type information of the second object to obtain the coordinate information of the 3D box; the obtaining unit 1001 is further specifically configured to obtain the coordinate information of the endpoints of the second dividing line according to the surface information of the second object and the coordinate information of the 3D box.
[0203] A possible implementation manner, the determining unit 1002 is specifically configured to perform coordinate transformation on the first coordinate through the first transformation relationship, the internal parameter matrix, and the external parameter matrix to obtain a second coordinate, the first coordinate is on the second dividing line, the distance between the second coordinate and the first object is K, the second coordinate, the third coordinate, and the first object are on a straight line, the third coordinate is the coordinate corresponding to the first coordinate on the first dividing line, the internal parameter matrix is the internal parameter matrix of the device for capturing the image, and the external parameter matrix is the external parameter matrix of the device; the determining unit 1002 is further specifically configured to perform a composite operation on the second coordinate, the third coordinate, and the K to obtain the distance between the second object and the first object.
[0204] A possible implementation manner, the determining unit 1002 is further specifically configured to obtain a second transformation relationship according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the distance between the second object and the first object; the second transformation relationship is the transformation relationship between the three-dimensional coordinate system corresponding to the 3D box and the three-dimensional coordinate system of the first object.
[0205] A possible implementation manner, the obtaining unit 1001 is further configured to rotate the 3D box N times with the second dividing line as the center to obtain the corresponding second 2D box for each rotated 3D box, the second 2D box is obtained by performing coordinate transformation on the rotated 3D box, and N is a positive integer; the determining unit 1002 is further configured to obtain, according to the N second 2D boxes corresponding to the 3D boxes after N rotations and the first 2D box, the second 2D box with the shortest distance between the boundary and the boundary of the first 2D box among the N second 2D boxes; the determining unit 1002 is further configured to determine the orientation angle of the second object according to the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance.
[0206] A possible implementation manner, the second 2D box is obtained by performing coordinate transformation on the rotated 3D box, including: the second 2D box is obtained by performing coordinate transformation on the rotated 3D box through the second transformation relationship, the internal parameter matrix, and the external parameter matrix, the internal parameter matrix is the internal parameter matrix of the device for capturing the image, and the external parameter matrix is the external parameter matrix of the device.
[0207] A possible implementation manner, the obtaining unit 1001 is further configured to perform M adjustments on the length of the boundary of the 3D box corresponding to the second 2D box with the shortest distance, and obtain the third 2D box corresponding to the 3D box after each adjustment. The third 2D box is obtained by performing coordinate transformation on the adjusted 3D box, and M is a positive integer; the determining unit 1002 is further configured to obtain, according to the M third 2D boxes corresponding to the 3D boxes after M adjustments and the first 2D box, the third 2D box with the shortest distance between the boundary and the boundary of the first 2D box among the M third 2D boxes; the determining unit 1002 is further configured to determine the length of the boundary of the second object according to the length of the boundary of the 3D box corresponding to the third 2D box with the shortest distance.
[0208] A possible implementation manner, the third 2D box is obtained by performing coordinate transformation on the adjusted 3D box, including: the third 2D box is obtained by performing coordinate transformation on the adjusted 3D box through the second transformation relationship, the intrinsic matrix, and the extrinsic matrix. The intrinsic matrix is the intrinsic matrix of the device for capturing the image, and the extrinsic matrix is the extrinsic matrix of the device.
[0209] A possible implementation manner, the obtaining unit 1001 is specifically configured to input the image into a neural network model to obtain the 2D information of the second object.
[0210] Wherein, all relevant contents of each operation involved in the above method embodiment can be cited to the function description of the corresponding functional module, and will not be elaborated here.
[0211] In this embodiment, the target detection device is presented in the form of dividing each functional module in an integrated manner. Here, "module" may refer to a specific ASIC, circuit, processor and memory for executing one or more software or firmware programs, integrated logic circuit, and / or other devices that can provide the above functions. In a simple embodiment, those skilled in the art can think that the target detection device can adopt Figure 3 the form shown.
[0212] For example, Figure 3 the processor 301 in [reference] can call the computer execution instructions stored in the memory 303 to enable the target detection device to execute the target detection method in the above method embodiment.
[0213] Exemplarily, Figure 10 the functions / implementation processes of the obtaining unit 1001 and the determining unit 1002 in [reference] can be implemented by Figure 3 the processor 301 in [reference] calling the computer execution instructions stored in the memory 303.
[0214] Since the target detection device provided in this embodiment can execute the above-mentioned target detection method, the technical effects it can obtain can refer to the above method embodiments and will not be elaborated here.
[0215] Figure 11 This is a schematic structural diagram of a chip provided in an embodiment of the present application. The chip 110 includes one or more processors 1101 and an interface circuit 1102. Optionally, the chip 110 may further include a bus 1103. Among them:
[0216] The processor 1101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor 1101 or the instructions in the form of software. The above-mentioned processor 1101 may be a general-purpose processor, a digital communicator (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0217] The interface circuit 1102 is used for sending or receiving data, instructions, or information. The processor 1101 may utilize the data, instructions, or other information received by the interface circuit 1102 for processing, and may send the processed information through the interface circuit 1102.
[0218] Optionally, the chip 110 further includes a memory, and the memory may include a read-only memory and a random access memory, and provides operation instructions and data to the processor. A part of the memory may also include a non-volatile random access memory (NVRAM).
[0219] Optionally, the memory stores an executable software module or data structure, and the processor may execute corresponding operations by calling the operation instructions stored in the memory (the operation instructions may be stored in the operating system).
[0220] Optionally, the chip 110 may be used in the target detection device involved in the embodiments of the present application. Optionally, the interface circuit 1102 may be used to output the execution result of the processor 1101. Regarding the target detection method provided by one or more embodiments of the present application, reference may be made to the foregoing respective embodiments and will not be elaborated here.
[0221] It should be noted that the respective functions corresponding to the processor 1101 and the interface circuit 1102 may be implemented through hardware design, may also be implemented through software design, or may be implemented through a combination of software and hardware, and are not limited here.
[0222] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0223] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0224] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0225] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit exists physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0226] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs and other various media that can store program codes.
[0227] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claimed rights.
Claims
1. A target detection method, characterized in that, the method includes: obtaining 2D information of a second object in an image obtained by a first object; the 2D information of the second object includes: coordinate information of endpoints of a first dividing line of the second object and type information of the second object, the first dividing line being a dividing line between a first surface and a second surface of the second object, at least one of the first surface and the second surface being included in a first 2D frame of the second object, the first 2D frame being a polygon that encloses the second object in the image; obtaining 3D information of the second object according to the 2D information of the second object; the 3D information of the second object includes coordinate information of endpoints of a second dividing line, the second dividing line being a dividing line between the first surface and the second surface in a 3D frame of the second object, the 3D frame being a 3D model of the second object, and the length of the boundary of the 3D frame corresponding to the type information of the second object; obtaining a first transformation relationship according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and a third transformation relationship; the third transformation relationship is a mapping relationship between a three-dimensional coordinate system of a camera that obtains the image and a three-dimensional coordinate system corresponding to the 3D frame, the camera being located in the first object, and the first transformation relationship is a transformation relationship between the three-dimensional coordinate system corresponding to the 3D frame and the three-dimensional coordinate system of the first object when the distance between the second object and the first object is K, where K>0; obtaining the distance between the second object and the first object according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the first transformation relationship.
2. The method according to claim 1, characterized in that, the 2D information of the second object further includes: coordinate information of the first 2D frame; the obtaining 3D information of the second object according to the 2D information of the second object includes: constructing the 3D frame according to the type information of the second object to obtain coordinate information of the 3D frame; determining the first surface and the second surface according to the coordinate information of the first 2D frame and the coordinate information of the endpoints of the first dividing line; obtaining the coordinate information of the endpoints of the second dividing line according to the first surface, the second surface, and the coordinate information of the 3D frame.
3. The method according to claim 1, characterized in that, the 2D information of the second object further includes: surface information of the second object, the surface information of the second object being used to indicate the first surface and the second surface; the obtaining 3D information of the second object according to the 2D information of the second object includes: constructing the 3D frame according to the type information of the second object to obtain coordinate information of the 3D frame; obtaining the coordinate information of the endpoints of the second dividing line according to the surface information of the second object and the coordinate information of the 3D frame.
4. The method according to any one of claims 1-3, characterized in that, Obtaining the distance between the second object and the first object according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the first transformation relationship includes: Performing coordinate transformation on the first coordinate through the first transformation relationship, the internal parameter matrix, and the external parameter matrix to obtain a second coordinate. The first coordinate is the coordinate of the endpoint of the second dividing line, the second coordinate is the two-dimensional coordinate of a point at a distance of K from the camera, the second coordinate, the third coordinate, and the camera are on a straight line, the third coordinate is the coordinate corresponding to the first coordinate on the first dividing line, the internal parameter matrix is the internal parameter matrix of the camera, and the external parameter matrix is the external parameter matrix of the camera; Performing a composite operation on the second coordinate, the third coordinate, and the K to obtain the distance between the second object and the first object.
5. The method according to any one of claims 1-3, characterized in that, the method further includes: Obtaining a second transformation relationship according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the distance between the second object and the first object; the second transformation relationship is the transformation relationship between the three-dimensional coordinate system corresponding to the 3D box and the three-dimensional coordinate system of the first object.
6. The method according to claim 5, characterized in that, the method further includes: Rotating the 3D box N times with the second dividing line as the center to obtain a second 2D box corresponding to the 3D box after each rotation. The second 2D box is obtained by performing coordinate transformation on the rotated 3D box, and N is a positive integer; According to the N second 2D boxes corresponding to the 3D boxes after N rotations and the first 2D box, obtaining the second 2D box with the shortest distance between the boundaries and the boundaries of the first 2D box among the N second 2D boxes; Determining the orientation angle of the second object according to the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance.
7. The method according to claim 6, characterized in that, the second 2D box is obtained by performing coordinate transformation on the rotated 3D box, including: The second 2D box is obtained by performing coordinate transformation on the rotated 3D box through the second transformation relationship, the internal parameter matrix, and the external parameter matrix. The internal parameter matrix is the internal parameter matrix of the camera, and the external parameter matrix is the external parameter matrix of the camera.
8. The method according to claim 6 or 7, characterized in that, the method further includes: Adjusting the length of the boundary of the 3D box corresponding to the second 2D box with the shortest distance M times to obtain a third 2D box corresponding to the 3D box after each adjustment. The third 2D box is obtained by performing coordinate transformation on the adjusted 3D box, and M is a positive integer; According to the M third 2D boxes corresponding to the 3D boxes after M adjustments and the first 2D box, obtaining the third 2D box with the shortest distance between the boundaries and the boundaries of the first 2D box among the M third 2D boxes; Determine the length of the boundary of the second object according to the length of the boundary of the 3D box corresponding to the third 2D box with the shortest distance.
9. The method according to claim 8, wherein, the third 2D box is obtained by performing coordinate transformation on the adjusted 3D box, including: the third 2D box is obtained by performing coordinate transformation on the adjusted 3D box through the second transformation relationship, the intrinsic matrix, and the extrinsic matrix, the intrinsic matrix is the intrinsic matrix of the camera, and the extrinsic matrix is the extrinsic matrix of the camera.
10. The method according to any one of claims 1-3, 6, 7, or 9, wherein, obtaining the 2D information of the second object in the image obtained by the first object includes: inputting the image into a neural network model to obtain the 2D information of the second object.
11. An object detection device, wherein, the object detection device includes: an acquisition unit and a determination unit; the acquisition unit is configured to acquire the 2D information of the second object in the image acquired by the first object; the 2D information of the second object includes: the coordinate information of the endpoints of the first dividing line of the second object and the type information of the second object, the first dividing line is the dividing line between the first surface and the second surface of the second object, and at least one of the first surface and the second surface is included in the first 2D box of the second object, and the first 2D box is a polygon that encloses the second object in the image; the acquisition unit is further configured to acquire the three-dimensional 3D information of the second object according to the 2D information of the second object; the 3D information of the second object includes the coordinate information of the endpoints of the second dividing line, and the second dividing line is the dividing line between the first surface and the second surface in the 3D box of the second object, and the 3D box is the 3D model of the second object, and the length of the boundary of the 3D box corresponds to the type information of the second object; the determination unit is configured to obtain a first transformation relationship according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the third transformation relationship; the third transformation relationship is the mapping relationship between the three-dimensional coordinate system of the camera that acquires the image and the three-dimensional coordinate system corresponding to the 3D box, and the first transformation relationship is the transformation relationship between the three-dimensional coordinate system corresponding to the 3D box and the three-dimensional coordinate system of the first object when the distance between the second object and the first object is K, and K>0; the determination unit is further configured to obtain the distance between the second object and the first object according to the coordinate information of the endpoints of the first dividing line, the coordinate information of the endpoints of the second dividing line, and the first transformation relationship.
12. The object detection device according to claim 11, wherein, the 2D information of the second object further includes: the coordinate information of the first 2D box; the acquisition unit is specifically configured to construct the 3D box according to the type information of the second object to obtain the coordinate information of the 3D box. The obtaining unit is further specifically configured to determine the first surface and the second surface according to the coordinate information of the first 2D box and the coordinate information of the endpoints of the first demarcation line; The obtaining unit is further specifically configured to obtain the coordinate information of the endpoints of the second demarcation line according to the coordinate information of the first surface, the second surface, and the 3D box.
13. The object detection device according to claim 11, wherein, The 2D information of the second object further includes: the surface information of the second object, and the surface information of the second object is used to indicate the first surface and the second surface; The obtaining unit is specifically configured to construct the 3D box according to the type information of the second object to obtain the coordinate information of the 3D box; The obtaining unit is further specifically configured to obtain the coordinate information of the endpoints of the second demarcation line according to the surface information of the second object and the coordinate information of the 3D box.
14. The object detection device according to any one of claims 11-13, wherein, The determining unit is specifically configured to perform coordinate transformation on the first coordinate through the first transformation relationship, the internal parameter matrix, and the external parameter matrix to obtain a second coordinate, where the first coordinate is the coordinate of the endpoint of the second demarcation line, the second coordinate is the two-dimensional coordinate of a point at a distance of K from the camera, the second coordinate and the third coordinate are on a straight line with the camera, the third coordinate is the coordinate corresponding to the first coordinate on the first demarcation line, the internal parameter matrix is the internal parameter matrix of the camera, and the external parameter matrix is the external parameter matrix of the camera; The determining unit is further specifically configured to perform a composite operation on the second coordinate, the third coordinate, and the K to obtain the distance between the second object and the first object.
15. The object detection device according to any one of claims 11-13, wherein, The determining unit is further configured to obtain a second transformation relationship according to the coordinate information of the endpoints of the first demarcation line, the coordinate information of the endpoints of the second demarcation line, and the distance between the second object and the first object; the second transformation relationship is the transformation relationship between the three-dimensional coordinate system corresponding to the 3D box and the three-dimensional coordinate system of the first object.
16. The object detection device according to claim 15, wherein, The obtaining unit is further configured to rotate the 3D box N times with the second demarcation line as the center to obtain the corresponding second 2D box of the 3D box after each rotation, where the second 2D box is obtained by performing coordinate transformation on the rotated 3D box, and N is a positive integer; The determining unit is further configured to obtain, according to the N second 2D boxes corresponding to the 3D boxes after N rotations and the first 2D box, the second 2D box with the shortest distance between the boundary and the boundary of the first 2D box among the N second 2D boxes; The determining unit is further configured to determine the orientation angle of the second object according to the rotation angle of the 3D box corresponding to the second 2D box with the shortest distance.
17. The object detection device according to claim 16, wherein, The second 2D box is obtained by performing coordinate transformation on the rotated 3D box, including: The second 2D box is obtained by performing coordinate transformation on the rotated 3D box through the second transformation relationship, the intrinsic matrix, and the extrinsic matrix. The intrinsic matrix is the intrinsic matrix of the camera, and the extrinsic matrix is the extrinsic matrix of the camera.
18. The object detection device according to claim 16 or 17, wherein, The obtaining unit is further configured to perform M adjustments on the length of the boundary of the 3D box corresponding to the second 2D box with the shortest distance, and obtain a third 2D box corresponding to the 3D box after each adjustment. The third 2D box is obtained by performing coordinate transformation on the adjusted 3D box, and M is a positive integer; The determining unit is further configured to obtain, according to the M third 2D boxes corresponding to the 3D boxes after M adjustments and the first 2D box, a third 2D box with the shortest distance between the boundary and the boundary of the first 2D box among the M third 2D boxes; The determining unit is further configured to determine the length of the boundary of the second object according to the length of the boundary of the 3D box corresponding to the third 2D box with the shortest distance.
19. The object detection device according to claim 18, wherein, The third 2D box is obtained by performing coordinate transformation on the adjusted 3D box, including: The third 2D box is obtained by performing coordinate transformation on the adjusted 3D box through the second transformation relationship, the intrinsic matrix, and the extrinsic matrix. The intrinsic matrix is the intrinsic matrix of the camera, and the extrinsic matrix is the extrinsic matrix of the camera.
20. The object detection device according to any one of claims 11-13, 16, 17, or 19, wherein, The obtaining unit is specifically configured to input the image into a neural network model to obtain 2D information of the second object.
21. An intelligent driving vehicle, wherein, including: The object detection device according to any one of claims 11-20.
22. A computer-readable storage medium, wherein, Computer program code is stored on the computer-readable storage medium, and when the computer program code is executed by a processing circuit, it implements the object detection method according to any one of claims 1-10.
23. A chip, wherein, The chip includes a processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the chip executes the object detection method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Video monitoring method and device
CN104954747A
Method and apparatus for detecting side of object using ground boundary information of obstacle
CN107491065A