Vehicle three-dimensional target detection method and device based on fisheye image

By correcting the distortion of pixel points in the fisheye image, building a three-dimensional cone point cloud and obtaining semantic feature information, the problem of poor detection effect of three-dimensional targets in the fisheye image in the prior art is solved, and higher detection accuracy and recognition accuracy are achieved.

CN120164203APending Publication Date: 2025-06-17BEIJING JINGWEI HIRAIN TECH CO INC
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510192584.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When using circumferential fish eye images for three-dimensional object detection, the detection effect is poor, the recognition accuracy is low, and it is not yet mature.

Method used

By acquiring fisheye images from multiple perspectives in the preset area of ​​the vehicle, distortion correction is performed on each pixel point, a three-dimensional cone point cloud is constructed, and semantic feature information is obtained through external product calculation, and allocated to the bird's eye view to achieve target detection.

Benefits of technology

Eliminate image distortion caused by fisheye image distortion, and improve the accuracy and recognition accuracy of vehicle three-dimensional object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164203A_ABST
    Figure CN120164203A_ABST
Patent Text Reader

Abstract

The invention discloses a fisheye image-based vehicle three-dimensional target detection method and device, and the method comprises the steps: carrying out the distortion correction of each pixel point in a fisheye image of each visual angle in a preset region of a vehicle, and carrying out the correction of the distortion of each pixel point in a space coordinate system of the vehicle, determining a three-dimensional view cone point cloud of each three-dimensional space point in the fisheye camera coordinate system according to a first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and a main point of the fisheye camera coordinate system and the depth information of the three-dimensional space point; performing outer product calculation on the feature vector and depth distribution of each point cloud in the three-dimensional view cone point cloud to obtain semantic feature information of each point cloud in the three-dimensional view cone point cloud; and distributing the semantic feature information of each point cloud in the three-dimensional view cone point cloud to a bird's-eye view feature map of a preset area of the vehicle to obtain a target bird's-eye view. The effect of accurately detecting the three-dimensional target in the preset area of the vehicle is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of autonomous driving, and particularly to a method and device for three-dimensional object detection of vehicles based on fisheye images. Background Art

[0002] With the rapid development of fields such as robot applications and autonomous driving, three-dimensional object detection has become increasingly important. In the field of intelligent driving, it is necessary to obtain information such as the relative positions, sizes, and orientations of pedestrians and vehicles on the road through various perception algorithms, so as to control the host vehicle to avoid pedestrians and vehicles.

[0003] Since the panoramic fisheye camera has a wide viewing angle range, no blind spots, and can also reduce the occlusion between objects, using fisheye images for three-dimensional object detection is a commonly used technical means at present. However, the current research on using panoramic fisheye images for three-dimensional object detection is not yet mature, the object detection effect is very poor, and the object recognition accuracy is low. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a method and device for three-dimensional object detection of vehicles based on fisheye images, so as to achieve the effect of accurately detecting three-dimensional objects within a preset area of the vehicle.

[0005] The technical solution of the present application is as follows:

[0006] In the first aspect, a method for three-dimensional object detection of vehicles based on fisheye images is provided, and the method includes:

[0007] Obtain fisheye images of multiple viewpoints within a preset area of the vehicle;

[0008] For each fisheye image, perform distortion correction on each pixel point in the fisheye image to obtain each pixel point after distortion correction;

[0009] For each three-dimensional space point in the spatial coordinate system of the vehicle, determine the three-dimensional frustum point cloud of each three-dimensional space point in the fisheye camera coordinate system according to the first distance between the pixel point after distortion correction corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, and the depth information of the three-dimensional space point;

[0010] Perform an outer product calculation on the feature vector and depth distribution of each point cloud in the three-dimensional frustum point cloud to obtain the semantic feature information of each point cloud in the three-dimensional frustum point cloud;

[0011] Assign the semantic feature information of each point cloud in the three-dimensional frustum point cloud to the bird's-eye view of the preset area of the vehicle to obtain a target bird's-eye view.

[0012] In the second aspect, a device for three-dimensional object detection of vehicles based on fisheye images is provided, and the device includes:

[0013] A first acquisition module, configured to acquire fisheye images of multiple perspectives within a preset area of the vehicle;

[0014] A first determination module, configured to perform distortion correction on each pixel point in the fisheye image for each fisheye image, to obtain the pixel points after distortion correction;

[0015] A second determination module, configured to, for each three-dimensional space point in the spatial coordinate system of the vehicle, determine the three-dimensional cone point cloud of each three-dimensional space point in the fisheye camera coordinate system according to a first distance between the pixel point after distortion correction corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, and the depth information of the three-dimensional space point;

[0016] A third determination module, configured to perform an outer product calculation on the feature vector and depth distribution of each point cloud in the three-dimensional cone point cloud, to obtain the semantic feature information of each point cloud in the three-dimensional cone point cloud;

[0017] A fourth determination module, configured to allocate the semantic feature information of each point cloud in the three-dimensional cone point cloud to a bird's-eye view of the preset area of the vehicle, to obtain a target bird's-eye view.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of any one of the vehicle three-dimensional target detection methods based on fisheye images in the embodiments of the present application are implemented.

[0019] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of any one of the vehicle three-dimensional target detection methods based on fisheye images in the embodiments of the present application are implemented.

[0020] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device can execute the steps of any one of the vehicle three-dimensional target detection methods based on fisheye images in the embodiments of the present application.

[0021] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:

[0022] In the embodiments of the present application, by respectively performing distortion correction on each pixel point in the fisheye images of multiple perspectives within the preset area of the vehicle, the problem of image distortion caused by the distortion of the fisheye images can be eliminated, thereby improving the accuracy of detecting three-dimensional targets of the vehicle. After respectively performing distortion correction on each pixel point in the fisheye images, for each three-dimensional space point in the spatial coordinate system of the vehicle, according to the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, and the depth information of the three-dimensional space point, the position information of the three-dimensional space point in the fisheye camera coordinate system is determined. Furthermore, a three-dimensional visual cone point cloud of each three-dimensional space point in the spatial coordinates within the preset area of the vehicle in the fisheye camera coordinate system can be constructed to accurately feedback the position and depth of the three-dimensional space point in the spatial coordinate system, thereby enabling accurate detection of three-dimensional targets within the preset area of the vehicle and further improving the accuracy of detecting three-dimensional targets within the preset area of the vehicle.

[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application, and do not constitute an improper limitation to the present application.

[0025] Figure 1 is a schematic flowchart of a method for detecting three-dimensional targets of a vehicle based on fisheye images provided by an embodiment of the present application;

[0026] Figure 2 is a schematic diagram of a fisheye camera projection model provided by an embodiment of the present application;

[0027] Figure 3 is a schematic diagram of a fisheye camera projection model provided by an embodiment of the present application;

[0028] Figure 4 is a schematic diagram of a three-dimensional visual cone point cloud of each three-dimensional space point in the fisheye camera coordinate system provided by an embodiment of the present application;

[0029] Figure 5 is a schematic diagram of a three-dimensional visual cone point cloud of each three-dimensional space point in the vehicle coordinate system provided by an embodiment of the present application;

[0030] Figure 6 is a schematic structural diagram of a device for detecting three-dimensional targets of a vehicle based on fisheye images provided by an embodiment of the present application;

[0031] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0032] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than limiting the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only provided to provide a better understanding of the present application by showing examples of the present application.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples consistent with some aspects of the present application as detailed in the appended claims.

[0034] Before introducing the solutions of the embodiments of the present application, the background art of the embodiments of the present application will be introduced first.

[0035] Currently, when performing three-dimensional object detection using fisheye images, it is usually based on a spatial cross-attention model. Through the projection model of the fisheye camera, the spatial three-dimensional pixel point coordinates are projected into the two-dimensional space of the fisheye camera as reference points for corresponding position encoding queries. At the same time, multi-scale features are sampled around the spatial three-dimensional pixel point coordinates, and then the position encoding and multi-scale features are weighted and summed to obtain a bird's-eye view (BEV) of the vehicle's surrounding environment. However, in the above method, the fisheye images used for training the spatial cross-attention model are virtual synthetic fisheye images. However, the difference between the virtual synthetic fisheye images and real fisheye images is still relatively large. Therefore, when using real fisheye images for processing, the spatial cross-attention model has a poor effect on the feature processing of spatial pixel points. Moreover, the above method of using the spatial cross-attention model to process the features of spatial pixel points can only obtain the positions and multi-scale features of spatial pixel points, and cannot construct a bird's-eye view of the vehicle's surrounding environment based on the depth information of spatial pixel points. As a result, the object detection effect is poor and the object recognition accuracy is low.

[0036] To solve the above problems, the embodiments of the present application provide a method and device for three-dimensional object detection of a vehicle based on fisheye images. By separately performing distortion correction on each pixel point in the fisheye images of multiple viewpoints within a preset area of the vehicle, the problem of image distortion caused by the distortion of the fisheye images can be eliminated, thereby improving the accuracy of detecting three-dimensional objects of the vehicle. After separately performing distortion correction on each pixel point in the fisheye images, for each three-dimensional space point in the spatial coordinate system of the vehicle, according to the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, and the depth information of the three-dimensional space point, the position information of the three-dimensional space point in the fisheye camera coordinate system is determined. Furthermore, a three-dimensional visual cone point cloud of each three-dimensional space point in the spatial coordinate within the preset area of the vehicle in the fisheye camera coordinate system can be constructed to accurately feedback the position and depth of the three-dimensional space points in the spatial coordinate system, and thus the three-dimensional objects within the preset area of the vehicle can be accurately detected, further improving the accuracy of detecting three-dimensional objects within the preset area of the vehicle.

[0037] It should be noted that the solution of the embodiments of the present application can be applied to scenarios that require a large field of view, such as the automatic valet parking scenario. This is because, different from common highways and urban roads, the parking scenario mainly involves the perception of the near-field area of the vehicle. The detection range of traditional pinhole cameras in the near-field area is small, it is difficult to cover the surrounding environment of the vehicle, and insufficient information can be provided. In contrast, fisheye cameras have a larger field of view and can capture the entire near-field area, making up for the deficiency of pinhole cameras in perception under close-range conditions. Therefore, multiple fisheye cameras are usually used in valet parking to capture the near-field area around the vehicle for three-dimensional object recognition of the vehicle.

[0038] Before introducing the solution of the embodiments of the present application, first introduce the professional terms involved in the embodiments of the present application:

[0039] Spatial coordinate system: That is, the world coordinate system.

[0040] Normalized plane coordinate system: In the normalized plane coordinate system, points in the three-dimensional world coordinate system are first transformed into the camera coordinate system, and then projected onto the plane of Z = 1 through homogeneous coordinate transformation. In this way, the coordinates of each point on the normalized plane are represented as (X / Z, Y / Z, 1), where X and Y are the normalized X and Y coordinates, and Z is the depth value.

[0041] The relationship between the normalized plane coordinate system and the image coordinate system is a ratio scaling relationship, that is, when the camera focal length f = 1, the normalized plane coordinate system coincides with the image coordinate system. In this case, only by multiplying the coordinates in the camera coordinate system by f can they be transformed into the image coordinate system.

[0042] The following will combine the accompanying drawings and, through specific embodiments and their application scenarios, elaborate in detail on the vehicle three-dimensional target detection method based on fisheye images provided by the embodiments of the present application.

[0043] Figure 1 FIG. 4 is a schematic flowchart of a vehicle three-dimensional target detection method based on fisheye images provided by an embodiment of the present application. The execution subject of the vehicle three-dimensional target detection method based on fisheye images may be a server. It should be noted that the above execution subject does not constitute a limitation on the embodiments of the present application.

[0044] As Figure 1 shown, the vehicle three-dimensional target detection method based on fisheye images provided by the embodiments of the present application may include step 110-step 150.

[0045] Step 110, obtain fisheye images of multiple perspectives within a preset area of the vehicle.

[0046] Among them, the preset area may be a certain area of the vehicle set in advance. For example, it may be within a range with a radius of 5 meters centered on the vehicle. The specific range of the preset area can be set according to user needs.

[0047] The fisheye image may be an image within the preset area of the vehicle obtained based on a fisheye camera deployed on the vehicle.

[0048] In some embodiments of the present application, fisheye cameras may be deployed at different positions of the vehicle to obtain fisheye images of different perspectives around the vehicle. For example, fisheye cameras may be respectively deployed on the front hood, trunk, left center (such as the position between the driver's door and the rear seat door of the driver's seat) and right center (such as the position between the passenger's door and the rear seat door of the passenger's seat) of the vehicle to obtain fisheye images around the vehicle.

[0049] It should be noted that the above fisheye images of multiple perspectives may be images obtained by actually photographing three-dimensional space points in the three-dimensional space of the vehicle based on the fisheye camera, rather than virtual synthesized images. In this way, three-dimensional target detection of the vehicle surrounding environment can be performed based on real fisheye images, improving the accuracy of vehicle three-dimensional target detection.

[0050] Step 120, for each fisheye image, perform distortion correction on each pixel point in the fisheye image to obtain each pixel point after distortion correction.

[0051] In some embodiments of the present application, in order to improve the accuracy of vehicle three-dimensional target detection, step 120 may specifically include:

[0052] Convert the first coordinate information of the first pixel point in the fisheye image from the fisheye image coordinate system to the normalized plane coordinate system to obtain the second coordinate information of the first pixel point in the normalized plane coordinates;

[0053] Determine the second angle of the incident light ray of the fisheye camera based on the second distance between the second coordinate information and the principal point of the fisheye camera coordinate system, and the first angle of the outgoing light ray of the fisheye camera;

[0054] Perform distortion correction on the second coordinate information according to the second angle and the projection model of the pinhole camera to obtain the first pixel point after distortion correction.

[0055] Wherein, the first pixel point is any pixel point among the pixel points in the fisheye image.

[0056] The first coordinate information can be the coordinate information of the first pixel point in the fisheye image.

[0057] The second coordinate information can be the coordinate information of the first pixel point in the normalized plane coordinate system after converting the first pixel point from the fisheye image coordinate system to the normalized plane coordinate system.

[0058] The second distance can be the distance between the second coordinate information and the principal point of the fisheye camera coordinate system.

[0059] The principal point of the fisheye camera coordinate system here can be understood as the origin of the fisheye camera coordinate system.

[0060] The first angle can be the angle between the outgoing light ray of the fisheye camera and the optical axis of the fisheye camera.

[0061] The second angle can be the angle between the incident light ray and the optical axis when using the fisheye camera for shooting.

[0062] It should be noted that each pixel in the fisheye image can be considered to be obtained by the light ray emitted from the fisheye camera and intersecting with the object in the real world.

[0063] In some embodiments of the present application, the first coordinate information of the first pixel point in the fisheye image can be converted from the fisheye image coordinate system to the normalized plane coordinate system to obtain the second coordinate information of the first pixel point in the normalized plane coordinates, and then based on the second distance between the second coordinate information and the principal point of the fisheye camera coordinate system, and the first angle of the outgoing light ray of the fisheye camera, determine the second angle of the incident light ray of the fisheye camera, and further, according to the second angle and the projection model of the pinhole camera, perform distortion correction on the second coordinate information to obtain the first pixel point after distortion correction.

[0064] In an embodiment of the present application, by performing distortion correction on each pixel point in the fisheye image, the true position of the pixel point on the image plane is restored. Thus, compared with directly performing distortion correction on the entire fisheye image, the loss of feature information in the fisheye image is reduced, and furthermore, the accuracy of three-dimensional object detection of the vehicle can be improved.

[0065] In some embodiments of the present application, the first coordinate information of the first pixel point in the fisheye image can be converted from the fisheye image coordinate system to the normalized plane coordinate system through the following formula (1) to obtain the second coordinate information of the first pixel point in the normalized plane coordinates:

[0066]

[0067] Wherein, in the above formula (1), (u, v) is the first coordinate information of the first pixel point, and f x 、f y 、c x and c y are all internal parameters of the fisheye camera, and (x', y') is the second coordinate information of the first pixel point in the normalized plane coordinate system.

[0068] The origin of the above formula (1) is introduced in detail below:

[0069] Before introducing the origin of the above formula (1), the projection function of the fisheye camera is first introduced:

[0070] In some embodiments of the present application, the complex lens structure of the fisheye camera lens can be abstracted into a simplified unit spherical model. Referring to Figure 2 , Figure 2 is the projection model of the fisheye camera. Among them, the X C axis, Y C axis, and Z C axis constitute the fisheye camera coordinate system, the x-axis and the y-axis constitute the image coordinate system. Assuming that P is a spatial point in three-dimensional space and θ is the angle between its incident light ray and the optical axis, that is, the second angle.

[0071] In the pinhole camera model, the light ray travels in a straight line after passing through the lens. Therefore, the projection of the spatial point P on the imaging plane is the image point p'. However, in the fisheye camera model, the light ray refracts after passing through the lens, and the propagation path of the light ray changes, which makes the projection position of the spatial point P on the imaging plane become the image point p, and its polar coordinates are expressed as

[0072] The projection models of both the pinhole camera and the fisheye camera can be represented by r and It is described as follows. Here, r is the distance between the image point and the principal point (for example, for a fish-eye camera, r is the distance between the image point p and the principal point of the fish-eye camera; for a pinhole camera, r is the distance between the image point p' and the principal point of the pinhole camera), θ is the angle between the incident light ray and the principal axis, and f is the camera focal length. For a pinhole camera, its projection follows the classical perspective projection principle, and its projection function can be expressed as the following formula (2):

[0073] r = ftanθ (2)

[0074] The design of a fish-eye camera usually follows four different projection methods, which include: equiareal projection, equidistant projection, stereographic projection, and orthographic projection. However, in practical applications, the imaging characteristics of a fish-eye camera often deviate from the theoretical projection function. To simplify the calibration process of the fish-eye camera and enable it to adapt to various types of fish-eye lenses, the Taylor series expansion based on θ can be used to approximate the actual projection function of the fish-eye camera. The specific expression is as shown in the following formula (3):

[0075] r(θ) = k1θ + k2θ 3 + k3θ 5 + k4θ 7 + k5θ 9 +… (3)

[0076] When performing calculations, it is crucial to determine the number of terms in the Taylor expansion. In fact, the first five terms of the Taylor expansion, that is, up to the fifth power of the variable, already provide sufficient degrees of freedom to approximate various different projection curves. Therefore, in actual operations, usually only the first five terms of the expansion are used for calculation. It should be noted that in the Taylor expansion, the coefficient of the first-order term has a relatively small impact on the final result. To simplify the calibration process and reduce the number of parameters to be calibrated, the coefficient of the first term is usually set to 1. In this way, the expression of the expansion can be simplified to the following formula (4):

[0077] r(θ) = θ(1 + k1θ 2 + k2θ 4 + k3θ 6 + k4θ 8 ) (4)

[0078] where k1, k2, k3, and k4 are all distortion coefficients of the fish-eye camera.

[0079] Then, the projection model of the fish-eye camera is introduced. According to the projection model of the fish-eye camera, the above formula (1) can be obtained:

[0080] As Figure 3 shown Figure 3Schematic diagram of the projection model of a fish-eye camera. The normalized plane coordinate system is OXY, and the fish-eye camera coordinate system is O C X X Y C Z C , assuming that the coordinate vector of a spatial point P in three-dimensional space in the world coordinate system (i.e., the spatial coordinate system in the following embodiments) is The coordinate vector in the fish-eye camera coordinate system is Then the conversion relationship between the world coordinate system and the fish-eye camera coordinate system is shown in the following formula (5):

[0081]

[0082] In the above formula (5), R is the rotation matrix of the external parameters of the fish-eye camera, and T is the translation matrix of the external parameters of the fish-eye camera.

[0083] As Figure 3 shown, assuming that the coordinates of the spatial point P in the fish-eye camera coordinate system are (x, y, z), then the following formula (6) can be obtained:

[0084]

[0085] In the above formula (6), and are respectively The coordinate components on the X C axis, Y C axis and Z C axis.

[0086] As Figure 3 shown, assuming that the image point corresponding to the spatial point P according to the pinhole camera projection model is P0, and the coordinates are (a, b), then the following formula (7) can be obtained from the geometric characteristics of the pinhole camera:

[0087]

[0088] Assuming that the equivalent focal length f of the fish-eye camera is 1, according to the projection principle of the pinhole camera, the angle between the light ray and the optical axis remains unchanged. Therefore, the incident angle and the exit angle are equal and both are θ, then the distance r between the image point P0 and the principal point of the fish-eye camera can be expressed as shown in the following formula (8):

[0089] r 2 = a 2 + b 2 (8)

[0090] The above r relationship with θ can also be expressed as the following formula (9):

[0091] θ = arctan(r) (9)

[0092] However, due to the distortion of the fisheye camera, assuming that the image point corresponding to the spatial point P according to the fisheye camera projection model is p', with coordinates (x', y') (these coordinates are distorted), according to the above formula (4), the projection function of the fisheye camera can be expressed as the following formula (10):

[0093] θ d = θ(1 + k1θ 2 + k2θ 4 + k3θ 6 + k4θ 8 ) (10)

[0094] In the above formula (10), θ d is the equivalent angle (not the actual angle) between the outgoing light ray and the principal axis after distortion. Then, from the following formula (11):

[0095] r d = ftanθ d (11)

[0096] In the above formula (11), the equivalent focal length f of the fisheye camera is 1. And, in actual situations, the average size of the imaging sensor of the fisheye camera is usually only a few millimeters, while the focal length often reaches several hundred millimeters. This means that during the actual imaging process, θ d is extremely small, that is, r d ≈ tanθ d ≈ θ d , so, according to the principle of similar triangles, the following formula (12) can be obtained:

[0097]

[0098] Therefore, the coordinates (x', y') of the image point p' after distortion satisfy the following formula (13):

[0099]

[0100] Finally, assuming that the coordinates of the image point p' in the pixel coordinate system (i.e., the fisheye image coordinate system) are (u, v), then from the coordinate transformation, the following formula (14) can be obtained:

[0101]

[0102] By transforming the above formula (14), the above formula (1) can be obtained.

[0103] It should be noted that (x', y') is the position where the spatial point P is projected onto the normalized plane by the fisheye camera, and the coordinates (a, b) obtained after distortion correction are equivalent to the result of projecting the spatial point P onto the normalized plane by the pinhole camera.

[0104] The coordinate system corresponding to the above-mentioned normalized plane can be understood as a coordinate system involved in the conversion between the world coordinate system and the fisheye camera coordinate system.

[0105] In the embodiments of the present application, through the above formula (1), the first coordinate information of the first pixel point in the fisheye image can be accurately converted from the fisheye image coordinate system to the fisheye camera coordinate system, so as to facilitate the distortion correction of the pixel points in the fisheye image.

[0106] In some embodiments of the present application, according to the principle of the above formula (8), and the formula r d ≈tanθ d ≈θ d , the second distance between the second coordinate information and the principal point of the fisheye camera coordinate system can be expressed as shown in the following formula (15):

[0107]

[0108] According to the above formula (9), the first angle of the outgoing light of the fisheye camera can be expressed as shown in the following formula (16):

[0109]

[0110] According to the projection function of the fisheye camera shown in the above formula (15), formula (16) and formula (10), θ can be solved, which is the second angle of the incident light of the fisheye camera.

[0111] In some embodiments of the present application, in order to accurately obtain the first pixel point after distortion correction, the distortion correction of the second coordinate information is performed according to the second angle and the projection model of the pinhole camera to obtain the first pixel point after distortion correction, which may specifically include:

[0112] Determine the third distance between the projection point of the first pixel point in the pinhole camera coordinate system and the principal point of the pinhole camera coordinate system according to the second angle and the projection model of the pinhole camera;

[0113] Perform distortion correction on the second coordinate information according to the third distance to obtain the first pixel point after distortion correction.

[0114] Among them, the third distance is the distance between the projection point of the first pixel point in the pinhole camera coordinate system and the principal point of the pinhole camera coordinate system obtained according to the second angle and the projection model of the pinhole camera.

[0115] In some embodiments of the present application, according to the second angle and the projection model of the pinhole camera, the third distance between the projection point of the first pixel point in the pinhole camera coordinate system and the principal point of the pinhole camera coordinate system can be obtained; then, based on the third distance, the distortion of the second coordinate information is corrected to obtain the first pixel point after distortion correction.

[0116] In the embodiments of the present application, by obtaining the third distance between the projection point of the first pixel point in the pinhole camera coordinate system and the principal point of the pinhole camera coordinate system according to the second angle and the projection model of the pinhole camera; then, based on the third distance, the distortion of the second coordinate information is corrected, so that the first pixel point after distortion correction can be accurately obtained.

[0117] In some embodiments of the present application, according to the above formula (2), it can be known that according to the second angle and the projection model of the pinhole camera, the third distance between the projection point of the first pixel point in the pinhole camera coordinate system and the principal point of the pinhole camera coordinate system can be obtained. In formula (2), f is the camera focal length, and its value is 1, then the third distance can be expressed as the following formula (17):

[0118] r = tan(θ) (17)

[0119] In some embodiments of the present application, the above formula (13) is transformed to obtain the following formula (18), so as to realize the distortion correction of the second coordinate information (x', y') and obtain the first pixel point (a, b) after distortion correction:

[0120]

[0121] In the above formula (18), (a, b) are the coordinates of the first pixel point after distortion correction.

[0122] Step 130: For each three-dimensional space point in the vehicle's spatial coordinate system, according to the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, and the depth information of the three-dimensional space point, determine the three-dimensional perspective cone point cloud of each three-dimensional space point in the fisheye camera coordinate system.

[0123] Among them, the distortion-corrected pixel point corresponding to the three-dimensional space point can be the pixel point after distortion correction corresponding to the three-dimensional space point in the fisheye image.

[0124] The first distance can be the distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system.

[0125] In some embodiments of the present application, in order to accurately and quickly determine the position information of the three-dimensional space point in the fisheye camera coordinate system, step 130 may specifically include:

[0126] Determine the position information of the three-dimensional space point in the fisheye camera coordinate system according to the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, and the depth information of the three-dimensional space point;

[0127] Determine the three-dimensional frustum point cloud of each three-dimensional space point in the fisheye camera coordinate system according to the position information of the three-dimensional space point in the fisheye camera coordinate system.

[0128] In some embodiments of the present application, according to the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, and the depth information of the three-dimensional space point, according to the following formula (19), the position information of the three-dimensional space point in the fisheye camera coordinate system can be determined:

[0129]

[0130] In the above formula (19), r1 is the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, (a, b) is the coordinate of the distortion-corrected pixel point corresponding to the three-dimensional space point, r2 is the distance between the three-dimensional space point and the principal point of the fisheye camera coordinate system, and f is the equivalent focal length of the fisheye camera coordinate system.

[0131] In some embodiments of the present application, according to the principle of the above formula (8), on the normalized plane, the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system can be

[0132] In some embodiments of the present application, a set containing D discrete depths can be predefined. The set of the D discrete depths can be a set of depth information corresponding to the three-dimensional space points of the vehicle in the space coordinate system, denoted as S: {d0 + Δ,... d0 + DΔ}, and each three-dimensional space point in the space coordinate system is associated with each discrete depth in S.

[0133] f = 1 in the above formula (19). Thus, according to formula (19), the following formula (20) can be obtained:

[0134]

[0135] In the normalized plane coordinate system, the angle between the distorted-corrected pixel point corresponding to the three-dimensional space point and the X-axis of the normalized plane coordinate system is arctan(b / a). Assuming that the depth value of a certain three-dimensional space point in the space coordinate system is d, and d ∈ S, then according to the principle of similar triangles, the coordinates of a certain three-dimensional space point with a depth value of d in the space coordinate system in the fisheye camera coordinate system can be expressed by the following formula (21):

[0136]

[0137] According to the above calculation method of the coordinates of the three-dimensional space point in the fisheye camera coordinate system, the same calculation is performed on each three-dimensional space point in the three-dimensional space to obtain the coordinates of each three-dimensional space point in the fisheye camera coordinate system. In this way, the three-dimensional frustum point cloud of each three-dimensional space point in the fisheye camera coordinate system can be obtained.

[0138] Step 140: Perform an outer product calculation on the feature vector and depth distribution of each point cloud in the three-dimensional frustum point cloud to obtain the semantic feature information of each point cloud in the three-dimensional frustum point cloud.

[0139] Among them, the semantic feature information can be information used to characterize the semantic features of the point cloud. For example, for a tree outside the vehicle, the corresponding semantic feature information can be an obstacle of the vehicle.

[0140] In some embodiments of the present application, it can be a three-dimensional object detection algorithm for constructing BEV features from bottom to top to perform an outer product calculation on the feature vector and depth distribution of each point cloud in the three-dimensional frustum point cloud to obtain the semantic feature information of each point cloud in the three-dimensional frustum point cloud. The above-mentioned three-dimensional object detection algorithm for constructing BEV features from bottom to top can be, for example, the LSS (Lift, Splat, Shoot) model.

[0141] In an embodiment of the present application, the LSS model can be used to calculate the features of each point cloud in the three-dimensional frustum point cloud to obtain the feature vector of each point cloud, and to explicitly estimate the depth distribution of each point cloud in the three-dimensional frustum point cloud to obtain the depth distribution of each point cloud. Then, an outer product calculation is performed on the feature vector and depth distribution of each point cloud in the three-dimensional frustum point cloud to obtain the semantic feature information of each point cloud in the three-dimensional frustum point cloud, and thus a three-dimensional frustum point cloud containing the semantic feature information of each point cloud can be obtained.

[0142] In some embodiments of the present application, in order to accurately obtain the semantic feature information of each space point in each three-dimensional space point, before step 140, the above-mentioned method may further include:

[0143] Based on the internal and external parameters of the fisheye camera, the three-dimensional frustum point cloud of each three-dimensional space point in the fisheye camera coordinate system is transformed from the fisheye camera coordinate system to the vehicle coordinate system, and the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system is obtained;

[0144] Step 140 may specifically include:

[0145] Perform an outer product calculation on the feature vector and depth distribution of each point cloud in the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system to obtain the semantic feature information of each point cloud in the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system.

[0146] In some embodiments of the present application, refer to Figure 4 , Figure 4 is a schematic diagram of the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system. The three-dimensional frustum point cloud of each three-dimensional space point in the fisheye camera coordinate system and the internal and external parameters of the camera can be input into the Lift algorithm in the LSS model. In this way, based on the Lift algorithm, the three-dimensional frustum point cloud can be transformed from the fisheye camera coordinate system to the vehicle coordinate system, and the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system as shown in Figure 5 can be obtained. Then, based on the Splat algorithm in the LSS model, an outer product calculation is performed on the feature vector and depth distribution of each point cloud in the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system, and the semantic feature information of each point cloud in the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system can be obtained.

[0147] It should be noted that Figure 4 and Figure 5 are both frustum point clouds corresponding to the front camera.

[0148] In the embodiments of the present application, based on the internal and external parameters of the fisheye camera, the three-dimensional frustum point cloud of each three-dimensional space point in the fisheye camera coordinate system is transformed from the fisheye camera coordinate system to the vehicle coordinate system, and the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system is obtained. Then, an outer product calculation is performed on the feature vector and depth distribution of each point cloud in the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system, so that the semantic feature information of each point cloud in the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinate system can be accurately obtained.

[0149] Step 150: Assign the semantic feature information of each point cloud in the three-dimensional frustum point cloud to the bird's-eye view feature map of the preset area of the vehicle to obtain the target bird's-eye view.

[0150] Among them, the target bird's-eye view may be the finally obtained bird's-eye view of the vehicle surrounding environment.

[0151] In some embodiments of the present application, the semantic feature information of each point cloud in the three-dimensional frustum point cloud of each three-dimensional space point in the vehicle coordinates can be assigned to each grid of the bird's-eye view of the vehicle based on the Splat algorithm in the LSS model, and then the frustum point cloud in each grid of the bird's-eye view is summarized and calculated to form a BEV feature map, and then the Shoot algorithm in the LSS model is used to process the BEV feature map to obtain the target bird's-eye view.

[0152] It should be noted that when the LSS model executes the above steps 140 - 150, the LSS model only needs to construct the frustum point cloud once before training, that is, after constructing the frustum point cloud of each three-dimensional space point in the fisheye camera coordinate system in step 130, it is stored, and then when processing each fisheye image, only the three-dimensional frustum point cloud of the corresponding fisheye camera needs to be directly retrieved.

[0153] It should be noted that for the vehicle three-dimensional target detection method based on fisheye images provided in the embodiments of the present application, the execution subject can be a vehicle three-dimensional target detection device based on fisheye images, or a control module in the vehicle three-dimensional target detection device based on fisheye images for executing the vehicle three-dimensional target detection method based on fisheye images.

[0154] Based on the same inventive concept as the above-mentioned vehicle three-dimensional target detection method based on fisheye images, the present application also provides a vehicle three-dimensional target detection device based on fisheye images. The following combines Figure 6 to elaborate in detail on the vehicle three-dimensional target detection device based on fisheye images provided in the embodiments of the present application.

[0155] Figure 6 is a schematic structural diagram of a vehicle three-dimensional target detection device based on fisheye images shown according to an exemplary embodiment.

[0156] As Figure 6 shown, the vehicle three-dimensional target detection device 600 based on fisheye images may include:

[0157] A first acquisition module 610, configured to acquire fisheye images of multiple perspectives of the vehicle;

[0158] A first determination module 620, configured to perform distortion correction on each pixel point in the fisheye image for each fisheye image to obtain each pixel point after distortion correction;

[0159] A second determination module 630, configured to determine, for each three-dimensional spatial point in the spatial coordinate system of the vehicle, a three-dimensional frustum point cloud of each three-dimensional spatial point in the fish-eye camera coordinate system according to a first distance between a distortion-corrected pixel point corresponding to the three-dimensional spatial point and a principal point of the fish-eye camera coordinate system, and depth information of the three-dimensional spatial point;

[0160] A third determination module 640, configured to perform an outer product calculation on a feature vector and a depth distribution of each point cloud in the three-dimensional frustum point cloud to obtain semantic feature information of each point cloud in the three-dimensional frustum point cloud;

[0161] A fourth determination module 650, configured to allocate the semantic feature information of each point cloud in the three-dimensional frustum point cloud to a bird's-eye view feature map of a preset area of the vehicle to obtain a target bird's-eye view map.

[0162] In an embodiment of the present application, by separately performing distortion correction on each pixel point in fish-eye images of multiple viewpoints within a preset area of the vehicle, the problem of image distortion caused by the distortion of the fish-eye images can be eliminated, thereby improving the accuracy of detecting three-dimensional targets of the vehicle. After separately performing distortion correction on each pixel point in the fish-eye images, for each three-dimensional spatial point in the spatial coordinate system of the vehicle, according to a first distance between a distortion-corrected pixel point corresponding to the three-dimensional spatial point and a principal point of the fish-eye camera coordinate system, and depth information of the three-dimensional spatial point, the position information of the three-dimensional spatial point in the fish-eye camera coordinate system is determined, and then a three-dimensional frustum point cloud of each three-dimensional spatial point in the spatial coordinates within the preset area of the vehicle in the fish-eye camera coordinate system can be constructed to accurately feedback the position and depth of the three-dimensional spatial points in the spatial coordinate system, and further the three-dimensional targets within the preset area of the vehicle can be accurately detected, further improving the accuracy of detecting three-dimensional targets within the preset area of the vehicle.

[0163] In some embodiments of the present application, the first determination module 620 may include:

[0164] A first determination unit, configured to convert first coordinate information of a first pixel point in the fish-eye image from a fish-eye image coordinate system to a normalized plane coordinate system to obtain second coordinate information of the first pixel point in the normalized plane coordinate, where the first pixel point is any one of the pixel points in the fish-eye image;

[0165] A second determination unit, configured to determine a second angle of an incident light ray of the fish-eye camera according to a second distance between the second coordinate information and a principal point of the fish-eye camera coordinate system, and a first angle of an outgoing light ray of the fish-eye camera;

[0166] A third determination unit, configured to correct the distortion of the second coordinate information according to the second angle and the projection model of the pinhole camera, so as to obtain the first pixel point after distortion correction.

[0167] In some embodiments of the present application, the first determination unit is specifically configured to:

[0168] According to the following formula, convert the first coordinate information of the first pixel point in the fisheye image from the fisheye image coordinate system to the normalized plane coordinate system, so as to obtain the second coordinate information of the first pixel point in the normalized plane coordinates:

[0169]

[0170] where (u, v) is the first coordinate information of the first pixel point, and f x , f y , c x and c y are all internal parameters of the fisheye camera, and (x', y') is the second coordinate information of the first pixel point in the normalized plane coordinate system.

[0171] In some embodiments of the present application, the second determination unit is specifically configured to:

[0172] According to the second distance between the second coordinate information and the principal point of the fisheye camera coordinate system, and the first angle of the outgoing light ray of the fisheye camera, solve for the second angle of the incident light ray of the fisheye camera according to the following formula:

[0173]

[0174] θ d = θ(1 + k1θ 2 + k2θ 4 + k3θ 6 + k4θ 8 )

[0175] where r d is the second distance between the second coordinate information and the principal point of the fisheye camera coordinate system, θ d is the first angle of the outgoing light ray of the fisheye camera, θ is the second angle of the incident light ray of the fisheye camera, and k1, k2, k3, and k4 are all distortion coefficients of the fisheye camera.

[0176] In some embodiments of the present application, the third determination unit specifically includes:

[0177] A first determination subunit, configured to determine a third distance between a projection point of a first pixel point in a pinhole camera coordinate system and a principal point of the pinhole camera coordinate system according to the second angle and a projection model of the pinhole camera;

[0178] A second determination subunit, configured to correct the distortion of the second coordinate information according to the third distance to obtain a first pixel point after distortion correction.

[0179] In some embodiments of the present application, the first determination subunit is specifically configured to:

[0180] According to the second angle and the projection model of the pinhole camera, determine the third distance between the projection point of the first pixel point in the pinhole camera coordinate system and the principal point of the pinhole camera coordinate system according to the following formula:

[0181] r = tan(θ)

[0182] where r is the third distance between the projection point of the first pixel point in the pinhole camera coordinate system and the principal point of the pinhole camera coordinate system.

[0183] In some embodiments of the present application, the second determination subunit is specifically configured to:

[0184] According to the third distance, correct the distortion of the second coordinate information according to the following formula to obtain a first pixel point after distortion correction:

[0185]

[0186] where (a, b) are the coordinates of the first pixel point after distortion correction.

[0187] In some embodiments of the present application, the second determination module 630 may include:

[0188] A fourth determination unit, configured to determine position information of the three-dimensional space point in the fisheye camera coordinate system according to a first distance between a pixel point after distortion correction corresponding to the three-dimensional space point and a principal point of the fisheye camera coordinate system, and depth information of the three-dimensional space point;

[0189] A fifth determination unit, configured to determine a three-dimensional visual cone point cloud of each three-dimensional space point in the fisheye camera coordinate system according to the position information of the three-dimensional space point in the fisheye camera coordinate system.

[0190] In some embodiments of the present application, the fourth determination unit is specifically configured to:

[0191] Determine the position information of the three-dimensional space point in the fisheye camera coordinate system according to the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, and the depth information of the three-dimensional space point, according to the following formula:

[0192]

[0193] where r1 is the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, (a, b) is the coordinate of the distortion-corrected pixel point corresponding to the three-dimensional space point, r2 is the distance between the three-dimensional space point and the principal point of the fisheye camera coordinate system, d is the depth information of the three-dimensional space point, f is the equivalent focal length of the fisheye camera coordinate system, and f = 1.

[0194] The vehicle three-dimensional target detection device based on fisheye images provided in the embodiments of the present application can be used to execute the vehicle three-dimensional target detection method based on fisheye images provided in the above method embodiments. The implementation principles and technical effects are similar. For the sake of brevity, they will not be elaborated here.

[0195] Based on the same inventive concept, the embodiments of the present application also provide an electronic device.

[0196] Figure 7 is a schematic structural diagram of an electronic device provided in the embodiments of the present application. As Figure 7 shown, the electronic device may include a processor 701 and a memory 702 storing computer programs or instructions.

[0197] Specifically, the above-mentioned processor 701 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0198] The memory 702 may include a mass storage for data or instructions. By way of example and not limitation, the memory 702 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 702 may include removable or non-removable (or fixed) media. Where appropriate, the memory 702 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, the memory 702 is a non-volatile solid-state memory. The memory may include a read-only memory (ROM), a random-access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described by the vehicle three-dimensional object detection method based on fisheye images provided by the above embodiments.

[0199] The processor 701 reads and executes the computer program instructions stored in the memory 702 to implement any one of the vehicle three-dimensional object detection methods based on fisheye images in the above embodiments.

[0200] In one example, the electronic device may further include a communication interface 703 and a bus 710. Among them, as Figure 7 shown, the processor 701, the memory 702, and the communication interface 703 are connected through the bus 710 and complete communication with each other.

[0201] The communication interface 703 is mainly used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present invention.

[0202] The bus 710 includes hardware, software, or both, and couples components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable bus or a combination of two or more of these. Where appropriate, the bus 710 may include one or more buses. Although embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.

[0203] The electronic device can execute the method for three-dimensional vehicle target detection based on a fisheye image in the embodiments of the present invention, so as to implement Figure 1 the described method for three-dimensional vehicle target detection based on a fisheye image.

[0204] In addition, in combination with the method for three-dimensional vehicle target detection based on a fisheye image in the above embodiments, embodiments of the present invention can provide a readable storage medium to implement. Program instructions are stored on the readable storage medium, and when the program instructions are executed by a processor, any one of the methods for three-dimensional vehicle target detection based on a fisheye image in the above embodiments is implemented.

[0205] In addition, in combination with the method for three-dimensional vehicle target detection based on a fisheye image in the above embodiments, embodiments of the present invention can provide a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device can execute any one of the methods for three-dimensional vehicle target detection based on a fisheye image in the above embodiments.

[0206] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0207] The functional blocks shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted over a transmission medium or communication link via a data signal carried in a carrier wave. A "machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0208] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0209] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block in the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It is also understood that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware for performing the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0210] As described above, this is only the specific implementation manner of the present invention. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A method for detecting three-dimensional vehicle targets based on fisheye images, characterized in that: The method comprises: Acquire fisheye images of multiple viewing angles within a preset area of ​​the vehicle; For each fisheye image, performing distortion correction on each pixel point in the fisheye image to obtain each pixel point after distortion correction; For each three-dimensional space point in the space coordinate system of the vehicle, determine a three-dimensional view frustum point cloud of each three-dimensional space point in the fisheye camera coordinate system according to a first distance between a distortion-corrected pixel point corresponding to the three-dimensional space point and a principal point of the fisheye camera coordinate system, and depth information of the three-dimensional space point; Performing outer product calculation on the feature vector and depth distribution of each point cloud in the three-dimensional view cone point cloud to obtain semantic feature information of each point cloud in the three-dimensional view cone point cloud; The semantic feature information of each point cloud in the three-dimensional view cone point cloud is assigned to the bird's-eye view feature map of the preset area of ​​the vehicle to obtain a target bird's-eye view map.

2. The method according to claim 1, characterized in that The step of respectively performing distortion correction on each pixel point in the fisheye image to obtain each pixel point after distortion correction comprises: Converting first coordinate information of a first pixel point in the fisheye image from a fisheye image coordinate system to a normalized plane coordinate system to obtain second coordinate information of the first pixel point in the normalized plane coordinate system, wherein the first pixel point is any pixel point among the pixels in the fisheye image; Determine a second angle of the incident light of the fisheye camera according to a second distance between the second coordinate information and the principal point of the fisheye camera coordinate system, and a first angle of the outgoing light of the fisheye camera; According to the second angle and the projection model of the pinhole camera, the second coordinate information is distorted and corrected to obtain the first pixel point after the distortion correction.

3. The method according to claim 2, characterized in that The converting the first coordinate information of the first pixel point in the fisheye image from the fisheye image coordinate system to the normalized plane coordinate system to obtain the second coordinate information of the first pixel point in the normalized plane coordinate system includes: According to the following formula, the first coordinate information of the first pixel in the fisheye image is converted from the fisheye image coordinate system to the normalized plane coordinate system to obtain the second coordinate information of the first pixel in the normalized plane coordinate system: Wherein, (u, v) is the first coordinate information of the first pixel point, f x 、f y 、c x and c y are all internal parameters of the fisheye camera, and (x', y') is the second coordinate information of the first pixel point in the normalized plane coordinate system.

4. The method according to claim 3, characterized in that The determining, according to the second distance between the second coordinate information and the principal point of the fisheye camera coordinate system and the first angle of the outgoing light of the fisheye camera, the second angle of the incident light of the fisheye camera comprises: According to the second distance between the second coordinate information and the principal point of the fisheye camera coordinate system, and the first angle of the outgoing light of the fisheye camera, the second angle of the incident light of the fisheye camera is obtained by solving the following formula: i d =θ(1+k1θ 2 +k2θ 4 +k3θ 6 +k4θ 8 ) Among them, r d is the second distance between the second coordinate information and the principal point of the fisheye camera coordinate system, θ d is the first angle of the outgoing light of the fisheye camera, θ is the second angle of the incident light of the fisheye camera, and k1, k2, k3 and k4 are all distortion coefficients of the fisheye camera.

5. The method according to claim 4, characterized in that The step of correcting the distortion of the second coordinate information according to the second angle and the projection model of the pinhole camera to obtain a first pixel point after the distortion is corrected includes: Determine, according to the second angle and the projection model of the pinhole camera, a third distance between a projection point of the first pixel point in the pinhole camera coordinate system and a principal point of the pinhole camera coordinate system; According to the third distance, the second coordinate information is distorted and corrected to obtain a first pixel point after distortion correction.

6. The method according to claim 5, characterized in that The determining, according to the second angle and the projection model of the pinhole camera, a third distance between a projection point of the first pixel point in the pinhole camera coordinate system and a principal point of the pinhole camera coordinate system comprises: According to the second angle and the projection model of the pinhole camera, a third distance between the projection point of the first pixel point in the pinhole camera coordinate system and the principal point of the pinhole camera coordinate system is determined according to the following formula: r = tan(θ) in, r is a third distance between the projection point of the first pixel in the pinhole camera coordinate system and the principal point of the pinhole camera coordinate system.

7. The method according to claim 6, characterized in that The step of correcting the distortion of the second coordinate information according to the third distance to obtain a first pixel point after the distortion is corrected includes: According to the third distance, the second coordinate information is distorted and corrected according to the following formula to obtain a first pixel point after distortion correction: Among them, (a, b) is the coordinate of the first pixel after distortion correction.

8. The method according to claim 1, characterized in that The method of determining a 3D view frustum point cloud of each 3D space point in the fisheye camera coordinate system according to a first distance between a distortion-corrected pixel point corresponding to the 3D space point and a principal point of the fisheye camera coordinate system and depth information of the 3D space point comprises: Determine the position information of the three-dimensional space point in the fisheye camera coordinate system according to a first distance between a distortion-corrected pixel point corresponding to the three-dimensional space point and a principal point of the fisheye camera coordinate system, and depth information of the three-dimensional space point; According to the position information of the three-dimensional space point in the fisheye camera coordinate system, a three-dimensional view cone point cloud of each three-dimensional space point in the fisheye camera coordinate system is determined.

9. The method according to claim 8, characterized in that The determining, according to a first distance between a distortion-corrected pixel point corresponding to the three-dimensional space point and a principal point of a fisheye camera coordinate system, and depth information of the three-dimensional space point, position information of the three-dimensional space point in the fisheye camera coordinate system comprises: According to a first distance between a distortion-corrected pixel point corresponding to the three-dimensional space point and a principal point of the fisheye camera coordinate system, and depth information of the three-dimensional space point, position information of the three-dimensional space point in the fisheye camera coordinate system is determined according to the following formula: Wherein, r1 is the first distance between the distortion-corrected pixel point corresponding to the three-dimensional space point and the principal point of the fisheye camera coordinate system, (a, b) are the coordinates of the distortion-corrected pixel point corresponding to the three-dimensional space point, r2 is the distance between the three-dimensional space point and the principal point of the fisheye camera coordinate system, d is the depth information of the three-dimensional space point, f is the equivalent focal length of the fisheye camera coordinate system, and f=1.

10. A vehicle three-dimensional target detection device based on fisheye images, characterized in that: The device comprises: A first acquisition module is used to acquire fisheye images of the vehicle from multiple perspectives; A first determination module is used to perform distortion correction on each pixel in each fisheye image to obtain each pixel after distortion correction; a second determination module, configured to determine, for each three-dimensional space point in the space coordinate system of the vehicle, a three-dimensional view frustum point cloud of each three-dimensional space point in the fisheye camera coordinate system according to a first distance between a distortion-corrected pixel point corresponding to the three-dimensional space point and a principal point of the fisheye camera coordinate system, and depth information of the three-dimensional space point; A third determination module is used to perform outer product calculation on the feature vector and depth distribution of each point cloud in the three-dimensional view cone point cloud to obtain semantic feature information of each point cloud in the three-dimensional view cone point cloud; The fourth determination module is used to assign the semantic feature information of each point cloud in the three-dimensional view cone point cloud to the bird's-eye view of the vehicle to obtain a target bird's-eye view.

Citation Information

Cited By

  • Around-looking image generation method and device, electronic equipment and computer readable storage medium

    CN121329759A

  • Target detection method and device based on fisheye image, and vehicle

    CN121330650A

  • Target detection method and device based on fisheye image and vehicle

    CN121330650B