Fish-eye camera target detection method and electronic device

CN122820428APending Publication Date: 2026-09-25UISEE SHANGHAI AUTOMOTIVE TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611133123.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0002]当前工业车辆、园区物流自动驾驶场景,普遍采用四路环视鱼眼相机实现360°无死角感知,但传统针孔基线稀疏检测方案仅适配无畸变的平视图像输入,依赖针孔透视投影公式完成三维稀疏关键点到二维像素的映射,将传统针孔基线稀疏检测方案直接应用于鱼眼场景,存在无法规避的技术缺陷,核心缺陷如下:

Benefits of technology

[0013]本公开实施例提供的一种鱼眼相机目标检测方法,获取各空间关键点以及当前鱼眼相机采集的目标图像,针对每一个空间关键点,将该空间关键点从车体坐标系转换至当前鱼眼相机的极坐标系,得到该空间关键点的极径与极角,并通过鱼眼畸变系数对极角进行校正,根据极径与校正后的极角确定该空间关键点在目标图像的像素平面中的像素坐标,进而通过预先训练的目标检测模型确定目标图像的当前特征图,并在当前特征图中对每个空间关键点的像素坐标对应的位置进行采样,得到采样特征,最终通过每个空间关键点对应的采样特征确定目标检测结果,实现基于鱼眼相机采集图像的3D目标检测,该方法通过对空间关键点在鱼眼相机极坐标系下的极角进行校正并确定其在像素平面中的像素坐标,能够达到对空间关键点的去畸变效果,在完成对空间关键点的去畸变之后进行投影采样,可以适配鱼眼相机非线性畸变成像的特性,避免投影错位与特征采样偏移,提高三维目标检测的准确性,并且,无需对鱼眼相机图像进行去畸变预处理,能够避免丢失图像边缘有效信息,进一步提高三维目标检测的精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820428A_ABST
    Figure CN122820428A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a fisheye camera target detection method and an electronic device. The method obtains each spatial key point and a target image collected by a current fisheye camera, converts each spatial key point from a vehicle body coordinate system to a polar coordinate system of the current fisheye camera to obtain a polar radius and a polar angle, corrects the polar angle through a fisheye distortion coefficient, determines a pixel coordinate in a pixel plane of the target image according to the polar radius and the corrected polar angle, further determines a current feature map through a pre-trained target detection model, samples a position corresponding to the pixel coordinate to obtain a sampling feature, and determines a target detection result. The method can adapt to the characteristics of fisheye camera nonlinear distortion imaging after completing the de-distortion of the spatial key points, does not need to perform de-distortion preprocessing on the image, avoids projection misplacement and feature sampling deviation, avoids losing effective information of an image edge, and ensures the accuracy of three-dimensional target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of autonomous driving technology, and in particular to a fisheye camera target detection method and electronic device. Background Technology

[0002] In current industrial vehicles and park logistics autonomous driving scenarios, four-way surround-view fisheye cameras are commonly used to achieve 360° perception without blind spots. However, traditional pinhole baseline sparse detection solutions are only compatible with distortion-free planar image input. They rely on pinhole perspective projection formulas to map 3D sparse key points to 2D pixels. Directly applying traditional pinhole baseline sparse detection solutions to fisheye scenarios has unavoidable technical defects, the core defects of which are as follows:

[0003] 1. Extremely large geometric projection error: Fisheye lenses have strong radial distortion, and the target at the image edge is severely deformed. The standard pinhole perspective projection model cannot adapt to the distortion imaging rules, resulting in the offset of the 3D anchor point projection pixel position, misalignment of sparse sampling points, and a significant decrease in detection accuracy.

[0004] 2. Distortion removal causes irreversible information loss: Preprocessing the original fisheye image to remove distortion will crop and stretch the effective pixels at the edges, lose edge perception information of the large field of view, compress the effective perception field of view, and introduce secondary geometric distortion. Summary of the Invention

[0005] To address or at least partially address the aforementioned technical problems, this disclosure provides a fisheye camera target detection method and electronic device that can adapt to the characteristics of nonlinear distortion imaging in fisheye cameras, avoid projection misalignment and feature sampling offset, and eliminate the need for distortion correction processing of fisheye camera images, thereby avoiding the loss of effective image edge information and improving the accuracy of 3D target detection.

[0006] In a first aspect, embodiments of this disclosure provide a target detection method using a fisheye camera, the method comprising:

[0007] Acquire key points in space and target images currently captured by the fisheye camera;

[0008] For each of the aforementioned spatial key points, the following steps are performed: transforming the spatial key point from the vehicle coordinate system to the polar coordinate system of the current fisheye camera to obtain the polar radius and polar angle of the spatial key point; correcting the polar angle according to the fisheye distortion coefficient of the current fisheye camera; and determining the pixel coordinates of the spatial key point in the pixel plane of the target image based on the polar radius and the corrected polar angle.

[0009] The current feature map of the target image is determined by a pre-trained target detection model, and the position corresponding to the pixel coordinates of each spatial key point in the current feature map is sampled to obtain the sampled features.

[0010] The target detection result of the target image is determined based on the sampling features corresponding to each of the spatial key points.

[0011] Secondly, embodiments of this disclosure also provide an electronic device, the electronic device comprising: one or more processors; a storage device for storing one or more programs; and when the one or more programs are executed by the one or more processors, the one or more processors implement the fisheye camera target detection method as described above.

[0012] Thirdly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the fisheye camera target detection method as described above.

[0013] This disclosure provides a fisheye camera target detection method that acquires spatial key points and the target image captured by the fisheye camera. For each spatial key point, the key point is transformed from the vehicle coordinate system to the polar coordinate system of the fisheye camera to obtain its polar radius and polar angle. The polar angle is corrected using fisheye distortion coefficients. Based on the polar radius and the corrected polar angle, the pixel coordinates of the spatial key point in the pixel plane of the target image are determined. Then, a pre-trained target detection model is used to determine the current feature map of the target image. The positions corresponding to the pixel coordinates of each spatial key point in the current feature map are sampled to obtain sampled features. Finally, the target image is detected using a fisheye camera target detection method. The sampling features corresponding to each spatial key point determine the target detection result, realizing 3D target detection based on images acquired by a fisheye camera. This method corrects the polar angle of the spatial key points in the polar coordinate system of the fisheye camera and determines their pixel coordinates in the pixel plane, which can achieve the distortion removal effect of the spatial key points. After completing the distortion removal of the spatial key points, projection sampling is performed, which can adapt to the nonlinear distortion imaging characteristics of the fisheye camera, avoid projection misalignment and feature sampling offset, improve the accuracy of 3D target detection, and avoid the loss of effective image edge information by eliminating the need for distortion removal preprocessing of the fisheye camera image, further improving the accuracy of 3D target detection. Attached Figure Description

[0014] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0015] Figure 1 This is a flowchart of a target detection method using a fisheye camera according to an embodiment of the present disclosure;

[0016] Figure 2 This is a schematic diagram of the structure of a fisheye camera target detection device according to an embodiment of the present disclosure;

[0017] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0021] To address the shortcomings of existing technologies, this disclosure provides a fisheye camera target detection method. This method addresses the strong radial distortion characteristic of fisheye camera images by correcting the polar angles of spatial key points in the fisheye camera's polar coordinate system and determining their pixel coordinates in the pixel plane. This achieves distortion correction of spatial key points, avoiding projection misalignment and feature sampling offset in the fisheye camera image. Furthermore, it avoids the loss of effective image edge information caused by distortion correction preprocessing of the fisheye camera image, further improving the accuracy of 3D target detection.

[0022] Figure 1 This is a flowchart illustrating a fisheye camera target detection method according to an embodiment of this disclosure. This method can be applied to purely visual perception scenarios such as industrial vehicles, park logistics vehicles, and low-speed autonomous driving, and is adapted to the 3D target detection characteristics of fisheye cameras, including large distortion, wide field of view, and multi-camera surround view. This method can be executed by a fisheye camera target detection device, which can be implemented in software and / or hardware and can be configured in an electronic device. Figure 1 As shown, the method may specifically include the following steps:

[0023] S110: Acquire key points in space and target images captured by the current fisheye camera.

[0024] In this embodiment, there can be multiple fisheye cameras, each capable of capturing a corresponding target image. For example, in a vehicle panoramic surround view system for parking assistance, there can be four fisheye cameras, which together constitute a 360-degree field of view.

[0025] Among them, spatial keypoints can be spatial locations of predicted potential targets, which can be understood as 3D anchor points; the number of spatial keypoints can be multiple, and the specific spatial keypoints can be determined by the distribution of real targets.

[0026] For example, a pre-labeled dataset of pre-labeled targets can be obtained, and the center coordinates of each labeled target can be extracted from the pre-labeled dataset. These center coordinates can be the coordinates of the center point of the labeled target in the vehicle coordinate system. Further, based on the center coordinates of each labeled target, labeled targets whose center coordinates are outside the effective perception range of the vehicle can be filtered out, thus removing invalid targets far from the vehicle and improving perception efficiency. Next, all remaining labeled targets can be clustered, for example, using the K-Means clustering algorithm to cluster the remaining labeled targets into multiple clusters, obtaining 3D bounding boxes for each cluster. Finally, the corner points and planar center points of each cluster's 3D bounding box can be identified as spatial keypoints.

[0027] In this embodiment of the disclosure, by learning spatial key points that are closer to the actual location of the target from the real target distribution, the distribution of spatial key points can be made uniform, ensuring the reliability of subsequent feature sampling.

[0028] S120. For each spatial key point, perform the following: transform the spatial key point from the vehicle coordinate system to the polar coordinate system of the current fisheye camera to obtain the polar radius and polar angle of the spatial key point; correct the polar angle according to the fisheye distortion coefficient of the current fisheye camera, and determine the pixel coordinates of the spatial key point in the pixel plane of the target image according to the polar radius and the corrected polar angle.

[0029] After obtaining each spatial key point, multiple spatial key points can be input to perform distortion correction on multiple spatial key points simultaneously, thereby achieving the goal of running a batch of spatial key points in parallel and improving distortion correction efficiency.

[0030] Specifically, for each spatial keypoint, it can first be transformed from the vehicle coordinate system to the polar coordinate system of the current fisheye camera, thereby obtaining the polar radius and polar angle of the spatial keypoint in the current fisheye camera's polar coordinate system. For example, the spatial keypoint can first be transformed from the vehicle coordinate system to the current fisheye camera's camera coordinate system, and then the spatial keypoint can be transformed from the camera coordinate system to the polar coordinate system.

[0031] In one specific implementation, the spatial key points are transformed from the vehicle coordinate system to the polar coordinate system of the current fisheye camera to obtain the polar radius and polar angle of the spatial key points, including the following steps:

[0032] Step A1: Based on the current fisheye camera's extrinsic parameter matrix, transform the spatial key points from the vehicle coordinate system to the current fisheye camera's camera coordinate system to obtain the camera spatial coordinates of the spatial key points.

[0033] Step A2: Determine the polar radius and polar angle of the key points in space based on the camera's spatial coordinates.

[0034] The camera extrinsic parameter matrix describes the spatial relationship between the vehicle coordinate system and the current fisheye camera's camera coordinate system. It maintains the three-dimensional geometric reference and represents the current fisheye camera's three-dimensional spatial position relative to the vehicle chassis. This matrix can be determined through calibration. The camera extrinsic parameter matrix can consist of a rotation matrix and a translation matrix. The pose transformation matrix.

[0035] Specifically, in step A1, the original coordinates of the spatial key points in the vehicle coordinate system can be obtained first, and the original coordinates can be homogeneously expanded, that is, the three-dimensional original coordinates can be expanded into four-dimensional coordinates to obtain four-dimensional original coordinates, so as to adapt to the dimension of the camera extrinsic matrix.

[0036] Furthermore, the original four-dimensional coordinates can be multiplied by the camera extrinsic parameter matrix of the current fisheye camera to obtain the transformed four-dimensional coordinates of the spatial keypoints. The effective spatial components can then be extracted from these transformed four-dimensional coordinates to obtain the camera spatial coordinates of the spatial keypoints in the current fisheye camera's camera coordinate system. This achieves the transformation from the vehicle coordinate system to the camera coordinate system, as shown in the following equation:

[0037] ;

[0038] In the formula, The transformed four-dimensional coordinates of the spatial key points. For the camera extrinsic matrix, These are the original four-dimensional coordinates of the spatial keypoints. The effective spatial components that can be obtained by disassembling the middle These can be used as camera spatial coordinates for spatial key points.

[0039] After transforming from the vehicle coordinate system to the camera coordinate system, further, in step A2, the polar radius and polar angle of the spatial keypoints in the current fisheye camera's polar coordinate system can be calculated based on the camera's spatial coordinates in the camera coordinate system. The polar radius describes the projected distance of the spatial keypoint on the camera's perpendicular optical axis plane, and the polar angle is the angle between the spatial keypoint and the camera's optical axis. As shown in the following formula:

[0040] , ;

[0041] In the formula, The camera spatial coordinates of the spatial key points. Let the polar radius of the spatial key point be defined in the polar coordinate system of the current fisheye camera. The polar angle of the spatial key point in the current fisheye camera's polar coordinate system.

[0042] It should be noted that if there are multiple fisheye cameras, the polar radius and polar angle of the spatial key points in the polar coordinate system of different current fisheye cameras can be determined according to the camera extrinsic parameter matrix of each current fisheye camera.

[0043] Steps A1-A2 above first use the camera extrinsic parameter matrix of the current fisheye camera to transform the spatial key points from the vehicle coordinate system to the camera coordinate system. Then, using the camera spatial coordinates of the transformed spatial key points, the polar radius and polar angle are solved. This ensures that the solved polar radius and polar angle are based on the lens optical center as the origin and the optical axis as the reference, thereby ensuring that the polar radius and polar angle can match the lens imaging physical model and guarantee geometric consistency.

[0044] After obtaining the polar radius and polar angle of the spatial key points, the polar angle can be further corrected using the fisheye distortion coefficients of the current fisheye camera to correct the radial distortion of the fisheye. The fisheye distortion coefficients describe the physical characteristics of the radial distortion of the fisheye and are composed of core distortion parameters k1 to k4.

[0045] For example, a polar angle multi-power matrix can be constructed based on the fisheye distortion coefficient, and then the nonlinear distortion correction of the polar angle can be completed through the polar angle multi-power matrix to accurately fit the true imaging law of the fisheye.

[0046] In one specific implementation, the polar angle is corrected based on the fisheye distortion coefficient of the current fisheye camera, including the following steps:

[0047] Step B1: Construct a polar angle multi-power matrix based on the polar angle, and determine the distortion correction coefficient based on the polar angle multi-power matrix and the fisheye distortion coefficient.

[0048] Step B2: Correct the polar angle using the distortion correction factor.

[0049] Specifically, in step B1, considering that fisheye cameras use spherical imaging and exhibit radial distortion, the larger the incident angle of light (the closer the image is to the edge), the more non-linearly the pixel position will be stretched by the lens. Therefore, a multi-power matrix of the polar angle can be constructed using even powers of the polar angle to accommodate the rotational symmetry of the radial distortion, as shown in the following equation:

[0050] ;

[0051] In the formula, It is a polar angle multi-order power matrix. The polar angle of the key point in space.

[0052] Furthermore, the polar angle multi-power matrix and the fisheye distortion coefficient can be substituted into the fourth-order even-degree polynomial model to obtain the distortion correction coefficient, as shown in the following equation:

[0053] ;

[0054] In the formula, This is the distortion correction factor. It is a polar angle multi-order power matrix. Let represent the fourth-order radial distortion of the fisheye camera. This distortion correction factor can be understood as a distortion correction amplification factor.

[0055] After determining the distortion correction coefficient, in step B2, the distortion correction coefficient can be multiplied by the polar angle to obtain the corrected polar angle, as shown in the following formula:

[0056] ;

[0057] In the formula, The corrected polar angle, This is the polar angle before correction. It's understandable that if there are multiple fisheye cameras, the polar angle of the spatial keypoint under each current fisheye camera can be corrected based on its fisheye distortion coefficient.

[0058] Steps B1-B2 above, by constructing a polar angle multi-power matrix and determining the distortion correction coefficients based on this polar angle multi-power matrix, can nonlinearly amplify or shrink the polar angle in the polar coordinate system of the current fisheye camera, simulating the radial stretching effect of the spherical surface of the fisheye lens, and providing reliable geometric parameters for subsequent projection onto the pixel plane.

[0059] After completing the distortion correction of the polar angle, the corrected polar angle and polar radius can be restored to two-dimensional planar coordinates. Then, combined with the camera intrinsic parameter matrix, the projection onto the pixel plane is completed to obtain the pixel coordinates of the spatial key points on the pixel plane.

[0060] In one example, the pixel coordinates of spatial keypoints in the pixel plane of the target image are determined based on the polar radius and the corrected polar angle, including the following steps:

[0061] Step C1: Based on the polar radius and the corrected polar angle, transform the camera spatial coordinates to the camera plane of the current fisheye camera to obtain the camera plane coordinates of the spatial key points;

[0062] Step C2: Based on the camera intrinsic parameter matrix of the current fisheye camera, transform the camera plane coordinates to the pixel plane of the target image to obtain the pixel coordinates of the spatial key points.

[0063] Specifically, in step C1, the camera plane coordinates of the spatial key points below the current fisheye camera's camera plane can be calculated using the camera spatial coordinates, polar radius, and corrected polar angle, as shown in the following formula:

[0064] ;

[0065] In the formula, The camera plane coordinates of the spatial key points, These are the values ​​of the camera spatial coordinates of the spatial keypoints in the first two dimensions. The corrected polar angle, This represents the radius. To avoid the division-to-zero anomaly in the above formula, a lower limit can be set for the radius to truncate it, such as... To avoid the case where the radius is zero.

[0066] After obtaining the camera plane coordinates of the spatial key points, further, in step C2, the camera plane coordinates can be extended into three-dimensional coordinates. The camera intrinsic parameter matrix of the current fisheye camera is multiplied by these three-dimensional camera plane coordinates to obtain the pixel coordinates of the spatial key points, thus achieving projection of the spatial key points. The camera intrinsic parameter matrix describes the mapping relationship between the physical space of the fisheye camera and the image pixel plane, and can be obtained through calibration. As shown in the following equation:

[0067] ;

[0068] in, These are the pixel coordinates of the spatial key points. For the camera intrinsic parameter matrix, These are the extended 3D camera plane coordinates.

[0069] Steps C1-C2 above, by using the polar radius and the corrected polar angle, first restore the camera plane coordinates of the spatial key points under the current fisheye camera's camera plane, which can accurately match the physical laws of fisheye spherical imaging. Then, combined with the camera intrinsic parameter matrix of the current fisheye camera, it is projected onto the pixel plane of the target image, which can predict the pixel coordinates of the spatial key points under real hardware imaging, ensuring projection accuracy.

[0070] It should be noted that the present invention implements distortion correction in the geometric projection link, which can avoid distortion correction being implemented in the image preprocessing stage, fully preserve all pixel information of the edge of the fisheye large field of view, and avoid edge field compression, irreversible loss of information, etc., thereby improving the ability to detect targets at the edge of the surrounding view.

[0071] In this embodiment of the disclosure, step S120 can be implemented by an algorithm independent of the target detection model, or step S120 can be embedded in the multi-scale feature sampling and aggregation module of the target detection model. The multi-scale feature sampling and aggregation module performs distortion correction and projection on each spatial key point, and then samples it in the current feature map of the target image.

[0072] S130. Using a pre-trained target detection model, determine the current feature map of the target image, and sample the position corresponding to the pixel coordinates of each spatial key point in the current feature map to obtain the sampled features.

[0073] The target detection model can adopt a neural network architecture and be trained by a batch of sample key points and sample fisheye images.

[0074] In one specific implementation, the method provided by this disclosure further includes the following steps:

[0075] Step D1: Obtain key points for each sample, fisheye images for each sample, and target detection bounding boxes for the fisheye images of the samples.

[0076] Step D2: For each sample key point, perform the following: transform the sample key point from the vehicle coordinate system to the polar coordinate system of the sample fisheye camera to obtain the polar radius and polar angle of the sample key point; correct the polar angle according to the fisheye distortion coefficient of the sample fisheye camera, and determine the pixel coordinates of the sample key point in the pixel plane of the sample fisheye image according to the polar radius and the corrected polar angle.

[0077] Step D3: Construct a sample dataset based on the pixel coordinates of key points of each sample, the fisheye image of each sample, and the target detection box of the fisheye image of the sample.

[0078] Step D4: Train the pre-built target detection model using the sample dataset to obtain the trained target detection model.

[0079] In step D1, the sample key points can be determined by the pre-labeled dataset corresponding to the fisheye camera that acquired the sample fisheye image. The specific determination process can be referred to the above description of spatial key points, and will not be repeated here.

[0080] Specifically, in step D2, distortion correction can be performed on multiple sample key points simultaneously to achieve parallel processing of batch sample key points. For each sample key point, it is first transformed from the vehicle coordinate system to the camera coordinate system of the fisheye camera, and then from the camera coordinate system to the polar coordinate system to obtain the polar radius and polar angle. Then, a polar angle multi-power matrix can be constructed based on the fisheye distortion coefficients, and the nonlinear distortion correction of the polar angle can be achieved through this matrix. Next, the corrected polar angle and polar radius can be restored to two-dimensional planar coordinates, and then combined with the camera intrinsic parameter matrix to complete the projection onto the pixel plane, obtaining the pixel coordinates of the sample key point on the pixel plane.

[0081] Furthermore, in step D3, the pixel coordinates of each sample key point, the sample fisheye image, and the target detection box of the sample fisheye image can be used as a sample data to construct a sample dataset.

[0082] Specifically, in step D4, sample data can be input into a pre-built target detection model so that the target detection model outputs a predicted detection box based on the pixel coordinates of each sample key point and the sample fisheye image. The loss value is calculated based on the deviation between the predicted detection box and the corresponding target detection box, and the model parameters are adjusted in reverse based on the loss value until the convergence condition is met or the number of iterations exceeds the set number, thus obtaining a trained target detection model.

[0083] Steps D1-D4 above correct the polar angle of the sample key points in the polar coordinate system of the fisheye camera and determine their pixel coordinates in the pixel plane, which can achieve the distortion removal effect of the sample key points. After completing the distortion removal of the sample key points, projection sampling can be performed, which can adapt to the nonlinear distortion imaging characteristics of the fisheye camera, avoid projection misalignment and feature sampling offset. Furthermore, there is no need to perform distortion removal preprocessing on the fisheye camera image, which can avoid losing effective image edge information and ensure the training reliability of the target detection model.

[0084] To increase sample diversity, the target detection box can be rotated. However, in related technologies, conventional data augmentation operations such as rotating 3D targets within the camera are usually performed using a single projection matrix. This approach disrupts the matching relationship of distortion parameters within the original camera, making it impossible to distinguish between physical imaging transformation and image augmentation transformation. This leads to the offset of the 3D anchor point projection coordinates, generating a large amount of geometric noise during training and making it difficult for the model to converge. Therefore, to solve this problem, in this embodiment, the 3D target rotation and distortion correction are decoupled. Inverse rotation transformation correction is performed before distortion correction of the sample key points to avoid the data augmentation operation disrupting the inherent imaging parameter matching relationship of the camera and causing misalignment between the 3D ground truth and the 2D feature sampling position.

[0085] In one example, before transforming the sample keypoints from the vehicle coordinate system to the polar coordinate system of the sample fisheye camera, the following steps are also included:

[0086] Step D21: Rotate the target detection box and record the corresponding inverse rotation transformation matrix;

[0087] Step D22: Perform homogeneous dimension expansion on the three-dimensional coordinates of the sample key points in the vehicle coordinate system to obtain four-dimensional homogeneous coordinates;

[0088] Step D23: Correct the four-dimensional homogeneous coordinates using the inverse rotation transformation matrix;

[0089] Step D24: Based on the corrected four-dimensional homogeneous coordinates, determine the corrected coordinates of the sample key points in the vehicle coordinate system.

[0090] Specifically, in step D21, to improve sample diversity or ensure model detection accuracy, the target detection box can be rotated, and a corresponding inverse rotation transformation matrix can be generated based on the rotation angle and direction. This inverse rotation transformation matrix can be used to compensate for the geometric offset caused by the 3D rotation of the target detection box.

[0091] Furthermore, in step D22, the three-dimensional coordinates of the sample keypoints in the vehicle coordinate system can be homogeneously expanded, i.e., an additional dimension is added to obtain four-dimensional homogeneous coordinates. In step D23, the inverse rotation transformation matrix is ​​multiplied by the four-dimensional homogeneous coordinates to compensate for the spatial coordinate offset caused by the rotation of the target detection box, thus correcting the position of the sample keypoints in the vehicle coordinate system. Next, in step D24, the effective spatial components are extracted from the corrected four-dimensional homogeneous coordinates to obtain the corrected coordinates of the sample keypoints in the vehicle coordinate system.

[0092] Steps D21-D24 above record the corresponding inverse rotation transformation matrix while rotating the target detection box. Before distortion removal and projection of the sample key points, the inverse rotation transformation matrix is ​​used to correct the sample key points. This ensures the accuracy of the subsequent projection geometric reference. Furthermore, it avoids coupling the enhancement operations such as rotation transformation with the camera intrinsic parameters, thus preventing the enhancement operations from disrupting the inherent imaging parameter matching relationship of the camera and causing misalignment between the 3D ground truth and the 2D feature sampling position.

[0093] To increase sample diversity, the fisheye images of the samples can be scaled, cropped, or flipped to perform 2D image enhancement transformations. Considering that in related technologies, a single projection matrix is ​​usually used to bind the in-camera image enhancement transformation operations, this method will destroy the matching relationship of the original in-camera distortion parameters, making it impossible to distinguish between physical imaging transformation and image enhancement transformation, which will lead to the offset of the 3D anchor point projection coordinates, generate a lot of geometric noise during the training process, and make it difficult for the model to converge. Therefore, in order to solve this problem, in this embodiment of the disclosure, the image enhancement transformation is decoupled from distortion correction. After distortion correction is performed on the key points of the samples, the enhancement transformation correction is achieved to avoid the data enhancement operation from destroying the inherent imaging parameter matching relationship of the camera and causing the 3D ground truth and the 2D feature sampling position to be misaligned.

[0094] In one example, after obtaining the key points of each sample, the fisheye image of each sample, and the target detection box of the fisheye image of the sample, the method further includes: performing image enhancement processing on the fisheye image of the sample and recording the corresponding enhancement transformation matrix.

[0095] After determining the pixel coordinates of the sample key points in the pixel plane of the sample fisheye image based on the polar radius and the corrected polar angle, the process also includes correcting the pixel coordinates based on the enhancement transformation matrix.

[0096] Specifically, while performing image enhancement processing on the sample fisheye image, the corresponding enhancement transformation matrix can be recorded. This enhancement transformation matrix can be used to describe the two-dimensional enhancement transformation relationship of the sample fisheye image, and is used to realize the secondary mapping of pixel coordinates after image enhancement. Among them, image enhancement transformation processing includes, but is not limited to, scaling, cropping, and flipping. The enhancement transformation matrix is ​​initially an identity matrix, and can accumulate enhancement transformation operations such as scaling, cropping, and flipping.

[0097] Furthermore, after correcting the polar angle of the sample keypoints and determining the pixel coordinates of the sample keypoints in the pixel plane of the sample fisheye image based on the polar radius and the corrected polar angle, these pixel coordinates can be multiplied by the enhancement transformation matrix to accommodate the two-dimensional coordinate offset caused by image enhancement processing such as scaling and cropping. Finally, based on the true size of the input sample fisheye image, the pixel coordinates of all sample keypoints can be normalized to the [0,1] interval, outputting accurate pixel coordinates that can be used for feature sampling.

[0098] In the above example, while performing image enhancement processing on the sample fisheye image, the corresponding enhancement transformation matrix is ​​recorded. After distortion correction and projection of the sample key points, the pixel coordinates of the sample key points are corrected by enhancement transformation using the enhancement transformation matrix. This can adapt to the two-dimensional coordinate offset caused by image enhancement processing such as scaling and cropping, ensuring the accuracy of the projection geometric reference. Furthermore, there is no need to couple the image enhancement processing with the camera intrinsic parameters, avoiding the enhancement operation from disrupting the inherent imaging parameter matching relationship of the camera and causing global three-dimensional to two-dimensional geometric mapping disorder. This ensures the stability and detection accuracy of fisheye pure vision three-dimensional detection.

[0099] In this embodiment of the disclosure, for each target image currently acquired by the fisheye camera, it can be input into a pre-trained target detection model to extract the current feature map of each target image through the target detection model.

[0100] To simultaneously capture fine-grained texture and high-level semantics in a target image, multi-level features can be mined from the target image to fully preserve multi-level semantics, providing rich information support for subsequent feature sampling and reducing false positives and false negatives caused by the failure of single-scale features. Different scales correspond to different semantic levels. For example, low-level features can correspond to low-level visual details such as edges, corners, textures, and colors; mid-level features can correspond to local semantics such as component structures (e.g., wheels, windows, body); and high-level features can correspond to global semantics such as target category, overall pose, and scene context. For instance, in situations involving occlusion, changes in lighting, viewpoint shifts, and long-distance blur, shallow features have weak anti-blurring capabilities but are rich in detail, making them suitable for unoccluded areas. Deeper features are more robust, less sensitive to local occlusion and changes in lighting, making them suitable for incomplete targets.

[0101] In one specific implementation, the object detection model includes a backbone network and a feature pyramid network. The pre-trained object detection model determines the current feature map of the target image, including the following steps:

[0102] Step E1: The target image is sequentially convolved through each convolutional layer in the backbone network to obtain the feature maps of the corresponding levels output by each convolutional layer.

[0103] Step E2: Through the feature pyramid network, the feature maps of adjacent levels are fused to obtain the current feature map at multiple scales.

[0104] The backbone network can be a bottom-up feature extraction network. It can consist of multiple convolutional layers, each with a different number of channels. These layers are connected sequentially, and the feature map resolution is gradually reduced and the number of channels is increased through internal residual blocks.

[0105] In this embodiment, considering that some networks (such as the ResNet backbone network) have weak generalization ability for feature extraction of fisheye images with strong distortion, which easily leads to the loss of semantic details, the backbone network of DINOv2 (DIstillation of No On-label version 2, second-generation unsupervised knowledge distillation visual model) can be selected as the main backbone network. DINOv2 can complete self-supervised pre-training through massive amounts of unlabeled image data, learn highly generalizable visual representations that are highly robust to geometric deformation and image distortion, and adapt to the imaging characteristics of nonlinear radial distortion and edge stretching distortion of fisheye cameras.

[0106] Specifically, in step E1, the target image can be input into the backbone network of the pre-trained target detection model. First, it undergoes initial convolution and initial pooling in the backbone network for preliminary downsampling. Then, it passes through each convolutional layer in sequence, and the residual blocks in each convolutional layer perform convolution processing to gradually reduce the feature map resolution, thereby obtaining the feature maps of the corresponding levels output by each convolutional layer.

[0107] For example, the backbone network can include four convolutional layers, namely Layer 1 to Layer 4, with corresponding input channel numbers of 256, 512, 1024, and 2048, respectively. Assuming the input target image size is H×W, the feature map output by Layer 1 has a size of H / 4×W / 4, representing low-level texture features (edges, details), and 256 channels; the feature map output by Layer 2 has a size of H / 8×W / 8, representing mid-level semantic features (local structure), and 512 channels; the feature map output by Layer 3 has a size of H / 16×W / 16, representing mid-to-high-level semantic features (part information), and 1024 channels; and the feature map output by Layer 4 has a size of H / 32×W / 32, representing high-level abstract features (global semantics, category information), and 2048 channels.

[0108] Furthermore, in step E2, the feature maps of each level output by the backbone network are used as input to the feature pyramid network. The feature pyramid network performs lateral connections and top-down upsampling to fuse feature maps of different levels, solving the problems of weak semantics in low-level features and lack of details in high-level features, and finally outputting multi-scale current feature maps.

[0109] For example, the feature maps of each level are input into the feature pyramid network. The feature pyramid network can use 1×1 convolution to uniformly adjust the number of channels of the feature maps of each level to a preset number of channels (such as 256). Furthermore, the feature pyramid network can start from the feature map of the highest level, gradually upsample the feature map and add it to the feature map of the previous level to fuse it, so as to allow high-level semantic information to be passed to low-level features.

[0110] Continuing with the previous example, the feature maps of the four layers output by the backbone network are C2 to C5, where C2 corresponds to 256 channels and C5 corresponds to 2048 channels. First, the feature pyramid network can perform a 1×1 convolution on C5 to reduce its dimensionality to 256 channels, resulting in P5. Then, P5 is upsampled by a factor of 2, and the upsampled result is added to C4 (1024 channels, reduced to 256 channels by 1×1) to obtain P4. Then, P4 is upsampled by a factor of 2, and the upsampled result is added to C3 (512 channels, reduced to 256 channels by 1×1) to obtain P3. Finally, P3 is upsampled by a factor of 2, and the upsampled result is added to C2 (256 channels, reduced to 256 channels by 1×1) to obtain P2. Finally, the feature pyramid network can use 3×3 convolutions to process P2~P5 respectively to eliminate the aliasing effect caused by upsampling, and output four multi-scale feature maps with 256 channels each.

[0111] Steps E1-E2 above extract feature maps at different levels through the backbone network and perform cross-scale fusion of feature maps at different levels through the feature pyramid network, finally outputting a multi-scale current feature map. This can simultaneously take into account global high-level semantic features and local spatial detail features, effectively suppressing feature degradation, semantic confusion and detail loss caused by fisheye distortion, providing a highly robust and high-resolution image feature base for subsequent feature sampling, and significantly improving the accuracy of 3D target detection in distorted scenes.

[0112] After obtaining the current feature map, the multi-scale feature sampling and aggregation module in the object detection model can further sample the current feature map based on the pixel coordinates of spatial key points to obtain sampled features.

[0113] S140. Determine the target detection result of the target image based on the sampling features corresponding to each spatial key point.

[0114] Specifically, after obtaining the sampling features corresponding to each spatial key point, the target detection result of the target image can be determined through the sampling features of all spatial key points.

[0115] In one specific implementation, the target detection model includes a multi-scale feature sampling and aggregation module and a target detection head; it samples the position corresponding to the pixel coordinates of each spatial key point in the current feature map to obtain sampled features, including:

[0116] The multi-scale feature sampling and aggregation module samples the pixel coordinates at the corresponding positions in the feature layer of each scale of the current feature map to obtain the sampled features of the pixel coordinates at each scale.

[0117] Based on the sampling features corresponding to each spatial keypoint, the target detection results of the target image are determined, including:

[0118] The sampling features of each pixel coordinate at all scales are fused to obtain fused features; the fused features of each pixel coordinate are input into the target detection head to obtain the target detection result of the target image.

[0119] Specifically, the multi-scale feature sampling and aggregation module can locate the pixel coordinates of spatial key points in each feature layer of the current feature map at each scale, and then sample at the located locations to obtain the sampled features at each scale. During the sampling process, the sampling weights and sampling range can also be adaptively adjusted to complete refined visual feature sampling at the corresponding locations and extract robust image semantic features.

[0120] Furthermore, the sampling features of pixel coordinates at all scales can be fused to obtain fused features. The fused features of each pixel coordinate can be used as input to the target detection head to determine the target detection result.

[0121] Specifically, the target detection head determines the target detection result based on the fused features. This can be achieved by fusing the fused features with the initial instance feature vector of the spatial keypoint to obtain the final features, which can describe the target corresponding to the spatial keypoint. Then, the detection box correction amount is determined based on the final features, and the detection box of the target corresponding to the spatial keypoint is corrected based on the detection box correction amount to obtain the target detection result. This process can be executed iteratively and repeatedly. Through multiple rounds of iterative correction, the target detection result can be made closer to the true location of the target.

[0122] The above implementation method samples different feature layers in the current feature map at multiple scales and fuses the sampled features at all scales to achieve target detection. It can continuously correct three-dimensional spatial deviations and refine target features, ultimately achieving high-precision three-dimensional target detection in fisheye distortion scenes. From the dual dimensions of geometric mapping and feature representation, it solves the technical problems of low accuracy and weak robustness in fisheye surround view three-dimensional detection.

[0123] The fisheye camera target detection method provided in this disclosure acquires spatial key points and the target image captured by the fisheye camera. For each spatial key point, it transforms the key point from the vehicle coordinate system to the polar coordinate system of the current fisheye camera to obtain the polar radius and polar angle of the key point. The polar angle is corrected using fisheye distortion coefficients. Based on the polar radius and the corrected polar angle, the pixel coordinates of the key point in the pixel plane of the target image are determined. Then, a pre-trained target detection model is used to determine the current feature map of the target image. The positions corresponding to the pixel coordinates of each spatial key point in the current feature map are sampled to obtain sampled features. Finally, the target image is detected by... The method determines the target detection result by sampling features corresponding to each spatial key point, realizing 3D target detection based on images acquired by a fisheye camera. This method corrects the polar angle of the spatial key points in the polar coordinate system of the fisheye camera and determines their pixel coordinates in the pixel plane, which can achieve the distortion removal effect of the spatial key points. After completing the distortion removal of the spatial key points, projection sampling is performed, which can adapt to the nonlinear distortion imaging characteristics of the fisheye camera, avoid projection misalignment and feature sampling offset, improve the accuracy of 3D target detection, and avoid the loss of effective image edge information by eliminating the need for distortion removal preprocessing of the fisheye camera image, further improving the accuracy of 3D target detection.

[0124] It should be noted that the fisheye camera target detection method provided in this disclosure has the following innovative features:

[0125] 1. End-to-end projection mechanism of fisheye distortion

[0126] A complete fisheye coordinate transformation chain is built through a batch-parallel fourth-order fisheye radial distortion correction process. For the model training phase, it follows a fixed logic of 3D coordinate correction, camera coordinate system transformation, fisheye polar coordinate distortion correction, original intrinsic parameter mapping, enhancement matrix quadratic transformation, and coordinate normalization. For the model inference phase, it follows a fixed logic of camera coordinate system transformation, fisheye polar coordinate distortion correction, and original intrinsic parameter mapping. No distortion preprocessing of fisheye camera images is required. Combined with multi-scale feature maps, it accurately optimizes 3D sparse anchor points and completes 3D object detection, achieving native adaptation of the sparse detection framework to fisheye panoramic distortion scenes from the algorithm's underlying layer.

[0127] 2. Dual-channel decoupled storage mechanism for physical and enhancement parameters

[0128] The design incorporates a five-category independent parameter decoupled storage architecture, which completely isolates the fisheye camera intrinsic parameters, fourth-order fisheye distortion coefficients, fisheye camera extrinsic parameters, rotation inverse transformation matrix, and enhancement transformation matrix. The physical imaging parameters remain fixed throughout the process, and the two-dimensional image transformation is accumulated only through the enhancement transformation matrix. This completely solves the pain point of image enhancement destroying the fisheye distortion imaging model and ensures the geometric mapping consistency throughout the model training and inference process.

[0129] 3. A multi-scale feature extraction mechanism for fisheye-specific applications based on DINOv2+FPN (Feature Pyramid Network).

[0130] Abandoning the ResNet convolutional backbone with weak native generalization ability of Sparse4D, this paper adopts the DINOv2 self-supervised pre-trained model combined with the FPN multi-scale fusion architecture. Relying on DINOv2's natural robust representation ability of geometric distortion and deformation, it integrates multi-layer and multi-scale features, taking into account both global semantics and local details. It perfectly adapts to the nonlinear distortion imaging characteristics of fisheye, solves the problems of traditional backbone feature degradation and semantic confusion, and provides a highly robust feature base for fisheye projection sampling and 3D decoding.

[0131] This disclosed embodiment relies on core improvements such as parameter decoupling geometry mechanism, fourth-order fisheye tensor projection, and DINOv2+FPN feature extraction to construct a complete engineering system from fisheye camera parameter configuration, original image feature extraction, distortion projection, feature aggregation to model training and inference. The entire process adapts to the characteristics of fisheye panoramic distortion, exhibiting strong engineering feasibility, versatility, and stability. Compared to existing technologies, it possesses the following technical advantages:

[0132] 1. Significantly improved accuracy in distorted scene perception: Fourth-order fisheye radial distortion correction is achieved, accurately fitting the nonlinear imaging law of fisheye, and eliminating geometric mapping deviation with parameter decoupling mechanism. In sparse projection sampling, the target effective pixels are accurately aligned, completely solving the problems of target offset, missed detection, and false detection in distorted scenes.

[0133] 2. Complete preservation of full field of view information: No distortion correction preprocessing is required for fisheye camera images, there is no loss of edge pixels, and the advantages of fisheye's 360° ultra-large field of view perception are maximized;

[0134] 3. Significantly enhanced training stability and generalization: Physical parameters and enhancement parameters are completely decoupled. Combined with a unified geometric projection interface and a dedicated training strategy, geometric noise caused by image enhancement is completely eliminated, the model converges faster, and has a stronger generalization ability for fisheye scenes with different degrees of distortion.

[0135] 4. Dual-layer anti-distortion mechanism, maximizing feature robustness: Constructing a dual-layer protection system of DINOv2+FPN anti-distortion feature extraction + fisheye accurate distortion projection sampling, adapting to fisheye distortion characteristics from two core dimensions of feature representation and geometric mapping, effectively avoiding feature degradation, semantic confusion and sampling misalignment, and greatly improving detection accuracy and model stability in complex surround distortion scenarios.

[0136] Based on the fisheye surround view dataset and following the official nuScenes 3D detection evaluation standard, the method provided in this embodiment is quantitatively evaluated. The effective evaluation space is the entire surround view area in the horizontal direction [-40m, 40m]. The evaluation uses four industry-standard distance thresholds: 0.5m, 1.0m, 2.0m, and 4.0m. First, the corresponding AP value for each category is calculated. Then, the average distance accuracy for each category is obtained by taking the average AP value of the four thresholds for each category. Finally, the global mAP index is obtained by weighted averaging of the average values ​​of all categories. The measured accuracy results for each category provided in this embodiment are as follows: car category 0.3771, truck category 0.3411, full_dolly category 0.1850, unfull_dolly category 0.3775, and trailer category 0.2697. Quantitative results show that the method provided in this embodiment has stable and excellent multi-category 3D target detection capability in fisheye strong distortion surround view scenarios, effectively verifying the effectiveness of the dual-layer anti-distortion mechanism and geometric decoupling architecture, and possessing extremely high engineering practical accuracy and reliability.

[0137] Figure 2 This is a schematic diagram of the structure of a fisheye camera target detection device according to an embodiment of this disclosure. Figure 2 As shown: The device includes a data acquisition module 210, a key point projection module 220, a sampling module 230, and a result determination module 240, wherein:

[0138] Data acquisition module 210 is used to acquire key points in space and target images currently captured by the fisheye camera;

[0139] The key point projection module 220 is used to perform the following for each spatial key point: transform the spatial key point from the vehicle coordinate system to the polar coordinate system of the current fisheye camera to obtain the polar radius and polar angle of the spatial key point; correct the polar angle according to the fisheye distortion coefficient of the current fisheye camera; and determine the pixel coordinates of the spatial key point in the pixel plane of the target image according to the polar radius and the corrected polar angle.

[0140] The sampling module 230 is used to determine the current feature map of the target image through a pre-trained target detection model, and to sample the position corresponding to the pixel coordinates of each spatial key point in the current feature map to obtain the sampled features;

[0141] The result determination module 240 is used to determine the target detection result of the target image based on the sampling features corresponding to each spatial key point.

[0142] Based on the above embodiments, optionally, the key point projection module 220 includes a polar angle correction unit, which is used for:

[0143] Construct a polar angle multi-power matrix based on the polar angle, and determine the distortion correction coefficient based on the polar angle multi-power matrix and the fisheye distortion coefficient.

[0144] Polar angles are corrected using distortion correction factors.

[0145] Based on the above embodiments, optionally, the key point projection module 220 includes a polar coordinate transformation unit, which is used for:

[0146] Based on the current fisheye camera's extrinsic parameter matrix, the spatial key points are transformed from the vehicle coordinate system to the current fisheye camera's camera coordinate system to obtain the camera spatial coordinates of the spatial key points.

[0147] Determine the polar radius and polar angle of key points in space based on the camera's spatial coordinates.

[0148] Based on the above embodiments, optionally, the key point projection module 220 includes a pixel projection unit, which is used for:

[0149] Based on the polar radius and the corrected polar angle, the camera spatial coordinates are transformed to the camera plane of the current fisheye camera to obtain the camera plane coordinates of the spatial key points;

[0150] Based on the camera intrinsic parameter matrix of the current fisheye camera, the camera plane coordinates are transformed to the pixel plane of the target image to obtain the pixel coordinates of the spatial key points.

[0151] Based on the above embodiments, optionally, the fisheye camera target detection device further includes a model training module, which is used for:

[0152] Obtain key points for each sample, fisheye images for each sample, and target detection bounding boxes for the fisheye images of the samples.

[0153] For each sample key point, the following steps are performed: transform the sample key point from the vehicle coordinate system to the polar coordinate system of the sample fisheye camera to obtain the polar radius and polar angle of the sample key point; correct the polar angle according to the fisheye distortion coefficient of the sample fisheye camera, and determine the pixel coordinates of the sample key point in the pixel plane of the sample fisheye image based on the polar radius and the corrected polar angle.

[0154] A sample dataset is constructed based on the pixel coordinates of key points of each sample, the fisheye image of each sample, and the target detection box of the fisheye image of the sample.

[0155] The pre-built object detection model is trained using a sample dataset to obtain the trained object detection model.

[0156] Optionally, based on the above implementation method, the model training module is further configured to:

[0157] Before transforming the sample key points from the vehicle coordinate system to the polar coordinate system of the sample fisheye camera, the target detection box is rotated and the corresponding inverse rotation transformation matrix is ​​recorded.

[0158] The three-dimensional coordinates of the sample key points in the vehicle coordinate system are homogeneously expanded to obtain four-dimensional homogeneous coordinates.

[0159] The four-dimensional homogeneous coordinates are corrected by the inverse rotation transformation matrix;

[0160] Based on the corrected four-dimensional homogeneous coordinates, determine the corrected coordinates of the sample key points in the vehicle coordinate system.

[0161] Optionally, based on the above implementation method, the model training module is further configured to:

[0162] After obtaining the key points of each sample, the fisheye image of each sample, and the target detection box of the fisheye image of the sample, the fisheye image of the sample is subjected to image enhancement processing, and the corresponding enhancement transformation matrix is ​​recorded.

[0163] Furthermore, after determining the pixel coordinates of the sample key points in the pixel plane of the sample fisheye image based on the polar radius and the corrected polar angle, the pixel coordinates are corrected according to the enhancement transformation matrix.

[0164] Based on the above implementation method, optionally, the target detection model includes a backbone network and a feature pyramid network, and the sampling module 230 is specifically used for:

[0165] The target image is processed by convolutional layers in the backbone network in sequence to obtain the feature maps of the corresponding levels output by each convolutional layer.

[0166] By using a feature pyramid network, feature maps from adjacent levels are fused to obtain current feature maps at multiple scales.

[0167] Based on the above implementation, optionally, the target detection model includes a multi-scale feature sampling and aggregation module and a target detection head, wherein the sampling module 230 is specifically used for:

[0168] The multi-scale feature sampling and aggregation module samples the pixel coordinates at the corresponding positions in the feature layer of each scale of the current feature map to obtain the sampled features of the pixel coordinates at each scale.

[0169] Furthermore, the sampling features of each pixel coordinate at all scales are fused to obtain fused features; the fused features of each pixel coordinate are input into the target detection head to obtain the target detection result of the target image.

[0170] The fisheye camera target detection device provided in this disclosure can execute the steps in the fisheye camera target detection method provided in this disclosure method embodiment, and has the execution steps and beneficial effects, which will not be repeated here.

[0171] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. See below for details. Figure 3 It shows a schematic diagram of a structure suitable for implementing the electronic device 500 in the embodiments of this disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0172] like Figure 3 As shown, the electronic device 500 may include a processing device 501, a ROM 502, a RAM 503, a bus 504, an input / output (I / O) interface 505, an input device 506, an output device 507, a storage device 508, and a communication device 509. The processing device (e.g., a central processing unit, a graphics processor, etc.) 501 can perform various appropriate actions and processes to implement the methods of the embodiments described herein, based on a program stored in the read-only memory (ROM) 502 or a program loaded from the storage device 508 into the random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing device 501, the ROM 502, and the RAM 503 are interconnected via the bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0173] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts, thereby implementing the fisheye camera target detection method as described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0174] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0175] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods provided in any embodiment of this disclosure.

[0176] Optionally, when one or more of the above-described procedures are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.

[0177] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0178] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

Claims

1. A target detection method using a fisheye camera, characterized in that, The method includes: Acquire key points in space and target images currently captured by the fisheye camera; For each of the aforementioned spatial key points, the following steps are performed: transforming the spatial key point from the vehicle coordinate system to the polar coordinate system of the current fisheye camera to obtain the polar radius and polar angle of the spatial key point; correcting the polar angle according to the fisheye distortion coefficient of the current fisheye camera; and determining the pixel coordinates of the spatial key point in the pixel plane of the target image based on the polar radius and the corrected polar angle. The current feature map of the target image is determined by a pre-trained target detection model, and the position corresponding to the pixel coordinates of each spatial key point in the current feature map is sampled to obtain the sampled features. The target detection result of the target image is determined based on the sampling features corresponding to each of the spatial key points.

2. The method according to claim 1, characterized in that, The step of correcting the polar angle based on the fisheye distortion coefficient of the current fisheye camera includes: Construct a polar angle multi-power matrix based on the polar angle, and determine the distortion correction coefficient based on the polar angle multi-power matrix and the fisheye distortion coefficient. The polar angle is corrected by the distortion correction factor.

3. The method according to claim 1, characterized in that, The step of transforming the spatial key points from the vehicle coordinate system to the polar coordinate system of the current fisheye camera to obtain the polar radius and polar angle of the spatial key points includes: Based on the camera extrinsic matrix of the current fisheye camera, the spatial key points are transformed from the vehicle coordinate system to the camera coordinate system of the current fisheye camera to obtain the camera spatial coordinates of the spatial key points. The polar radius and polar angle of the spatial key point are determined based on the camera's spatial coordinates.

4. The method according to claim 3, characterized in that, Determining the pixel coordinates of the spatial key points in the pixel plane of the target image based on the polar radius and the corrected polar angle includes: Based on the polar radius and the corrected polar angle, the camera spatial coordinates are transformed to the camera plane of the current fisheye camera to obtain the camera plane coordinates of the spatial key points; Based on the camera intrinsic parameter matrix of the current fisheye camera, the camera plane coordinates are transformed to the pixel plane of the target image to obtain the pixel coordinates of the spatial key points.

5. The method according to claim 1, characterized in that, The method further includes: Obtain key points for each sample, fisheye images for each sample, and target detection bounding boxes for the fisheye images of the samples. For each of the sample key points, the following steps are performed: transform the sample key point from the vehicle coordinate system to the polar coordinate system of the sample fisheye camera to obtain the polar radius and polar angle of the sample key point; correct the polar angle according to the fisheye distortion coefficient of the sample fisheye camera; and determine the pixel coordinates of the sample key point in the pixel plane of the sample fisheye image according to the polar radius and the corrected polar angle. A sample dataset is constructed based on the pixel coordinates of each sample key point, the fisheye image of each sample, and the target detection box of the fisheye image of the sample. The pre-built target detection model is trained using the sample dataset to obtain the trained target detection model.

6. The method according to claim 5, characterized in that, Before transforming the sample key points from the vehicle coordinate system to the polar coordinate system of the sample fisheye camera, the method further includes: The target detection box is rotated, and the corresponding inverse rotation transformation matrix is ​​recorded. The three-dimensional coordinates of the sample key points in the vehicle coordinate system are homogeneously expanded to obtain four-dimensional homogeneous coordinates. The four-dimensional homogeneous coordinates are corrected using the inverse rotation transformation matrix; Based on the corrected four-dimensional homogeneous coordinates, the corrected coordinates of the sample key points in the vehicle coordinate system are determined.

7. The method according to claim 5, characterized in that, After acquiring the key points of each sample, the fisheye image of each sample, and the target detection bounding box of the fisheye image of the sample, the method further includes: The sample fisheye image is subjected to image enhancement processing, and the corresponding enhancement transformation matrix is ​​recorded; After determining the pixel coordinates of the sample key points in the pixel plane of the sample fisheye image based on the polar radius and the corrected polar angle, the method further includes: The pixel coordinates are corrected based on the enhancement transformation matrix.

8. The method according to claim 1, characterized in that, The target detection model includes a backbone network and a feature pyramid network. The process of determining the current feature map of the target image using the pre-trained target detection model includes: The target image is processed by convolutional layers in the backbone network in sequence to obtain feature maps of the corresponding level output by each convolutional layer. The feature pyramid network is used to fuse feature maps from adjacent levels to obtain a multi-scale current feature map.

9. The method according to claim 8, characterized in that, The target detection model includes a multi-scale feature sampling and aggregation module and a target detection head; The sampling features are obtained by sampling the pixel coordinates of each spatial key point in the current feature map, including: The multi-scale feature sampling and aggregation module samples the pixel coordinates at the corresponding positions in the feature layers of each scale of the current feature map to obtain the sampled features of the pixel coordinates at each scale. Based on the sampling features corresponding to each of the spatial key points, the target detection result of the target image is determined, including: The sampled features of each pixel coordinate at all scales are fused to obtain the fused features; The fused features of each pixel coordinate are input into the target detection head to obtain the target detection result of the target image.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-9.