A target global domain recognition method based on ring domain convolution

CN122714743APending Publication Date: 2026-09-08ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610683425.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

本发明要解决的技术问题是如何提供一种基于环域卷积的目标全域识别方法,以解决传统深度学习网络平移等变性与鱼眼图像径向对称性不适配的问题

Benefits of technology

本发明提出一种基于环域卷积的目标全域识别方法,本发明直接以鱼眼图像作为输入,能够更完整地保留鱼眼图像的宽视场信息和原始畸变特征;同时,通过构建环域卷积核,使卷积核方向能够随目标位置相对于鱼眼中心的方位变化而自适应调整,提高了卷积特征提取与鱼眼径向畸变分布之间的匹配性,增强了对强畸变区域目标的特征表征能力;进一步结合圆周扫描方式对鱼眼图像全域范围进行特征提取,实现了对鱼眼有效成像区域的全覆盖检测,并通过多层级特征融合综合利用浅层细节信息与深层语义信息,从而提升了对不同尺度目标、不同径向位置目标及复杂畸变目标的识别精度与鲁棒性,最终能够实现鱼眼图像中目标的精准识别与定位;在保证检测精度的同时,减少了计算复杂度,提升了目标识别的实时性,适配无人机拍摄、自动驾驶等实时性要求较高的应用场景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122714743A_ABST
    Figure CN122714743A_ABST
Patent Text Reader

Abstract

The application relates to a target global domain recognition method based on ring domain convolution and belongs to the target recognition field in computer vision. A fisheye camera is used to collect a global target distortion image, and through data cleaning and labeling, a distortion image dataset is formed; a target three-dimensional field distribution is converted into an image two-dimensional pixel array, and a nonlinear mapping relationship between three-dimensional space points and distortion two-dimensional pixels is established; a distortion target recognition model based on a deep learning network is constructed, a deep learning network framework is optimized, a contradiction between translation invariance of a traditional deep learning network and radial symmetry adaptation of a fisheye image is relieved, and a precise mapping relationship between a two-dimensional pixel domain and target semantic features is established; the distortion image dataset is input into the target recognition model, model training and parameter optimization are carried out, and target global domain recognition is realized. While ensuring detection accuracy, the application reduces the calculation complexity and improves the real-time performance of target recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of target recognition in computer vision, specifically relating to a global target recognition method based on annular domain convolution. Background Technology

[0002] Target recognition is a key technology that uses image processing, pattern recognition, artificial intelligence and other methods to locate and identify targets from visual data such as images, videos and point clouds. It is widely used in fields such as intelligent monitoring, autonomous driving, non-destructive testing and photoelectric reconnaissance.

[0003] In the field of wide field-of-view perception, target image acquisition methods are mainly divided into multi-camera array stitching imaging and single-lens fisheye imaging. Multi-camera arrays are difficult to deploy flexibly in lightweight scenarios, while fisheye lenses have a simple structure and high reliability, making them more suitable for highly mobile application environments such as low-altitude detection. However, while acquiring images with a wide field of view, fisheye lenses introduce strong nonlinear distortion, leading to geometric distortion of the target and increasing the difficulty of target recognition. To ensure the accuracy of target detection, recognition, localization, and measurement in distorted images, existing methods mostly employ deep learning methods to correct distorted images, correcting the geometric distortion of the image caused by the optical physical characteristics of the imaging system, and restoring the true spatial shape, size ratio, and positional relationship of the target. However, distortion correction can lead to overcorrection or loss of correction feasibility in ultra-wide field-of-view and extreme nonlinear distortion scenarios. Therefore, combining deep learning methods with distortion-free correction techniques for feature extraction is key to achieving full-field target recognition. However, the convolutional operation, the core carrier of deep learning feature extraction, suffers from a mismatch between its translational isovariability and the radial symmetry of fisheye images. This leads to difficulties in extracting distorted features and can even result in feature perception and cognitive dilemmas. Therefore, in the field of object recognition, there is still a lack of a method that can effectively solve the mismatch between the translational isovariability of traditional convolutional networks and the radial symmetry of fisheye images. Summary of the Invention

[0004] (a) Technical problems to be solved The technical problem to be solved by this invention is how to provide a target global recognition method based on annular domain convolution to solve the problem of the incompatibility between the translational equivariance of traditional deep learning networks and the radial symmetry of fisheye images.

[0005] (II) Technical Solution To address the aforementioned technical problems, this invention proposes a target global recognition method based on ring domain convolution, which includes the following steps: Step S1: Building a Distorted Image Dataset: Use a fisheye camera to collect distorted images of the target across the entire field. Through data cleaning and annotation, a distorted image dataset is formed. Step S2, Nonlinear Imaging Model Construction: The three-dimensional field-of-view distribution of the target is converted into a two-dimensional pixel array of the image, and a nonlinear mapping relationship is established between the object points in the three-dimensional space and the distorted two-dimensional pixels; Step S3: Construction of the Distorted Target Recognition Model: Construct a distorted target recognition model based on a deep learning network, optimize the deep learning network framework, alleviate the contradiction between the translational equivariance of traditional deep learning networks and the radial symmetry of fisheye images, and establish a precise mapping relationship between the two-dimensional pixel domain and the semantic features of the target. Step S4, Model Training and Recognition: Input the distorted image dataset into the target recognition model. After model training and parameter optimization, target recognition is achieved across the entire domain.

[0006] (III) Beneficial Effects This invention proposes a target global recognition method based on annular domain convolution. Using fisheye images directly as input, this method more completely preserves the wide field of view and original distortion features of the fisheye image. Simultaneously, by constructing annular domain convolution kernels, the kernel direction can adaptively adjust according to the target's position relative to the fisheye center, improving the matching between convolution feature extraction and the radial distortion distribution of the fisheye, and enhancing the feature representation capability for targets in strongly distorted regions. Furthermore, combining a circular scanning method to extract features across the entire fisheye image range achieves full coverage detection of the effective imaging area of ​​the fisheye. Multi-level feature fusion comprehensively utilizes shallow detail information and deep semantic information, thereby improving the recognition accuracy and robustness for targets of different scales, targets at different radial positions, and targets with complex distortions. Ultimately, it enables accurate target recognition and localization in fisheye images. While ensuring detection accuracy, it reduces computational complexity and improves the real-time performance of target recognition, making it suitable for applications with high real-time requirements such as drone photography and autonomous driving. Attached Figure Description

[0007] Figure 1 This is an overall flowchart of the target global recognition method based on annular domain convolution of the present invention; Figure 2 This is a schematic diagram of a nonlinear imaging model in an embodiment of the present invention; Figure 3 This is a schematic diagram of the ring domain convolution kernel in an embodiment of the present invention. Detailed Implementation

[0008] To make the objectives, contents, and advantages of the present invention clearer, the specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples.

[0009] The technical problem to be solved by this invention is how to provide a target global recognition method based on annular domain convolution to solve the problem of the incompatibility between the translational equivariance of traditional deep learning networks and the radial symmetry of fisheye images.

[0010] This invention provides the following technical solution: a target global recognition method based on annular domain convolution, the method comprising the following steps: Step S1: Building a Distorted Image Dataset: Use a fisheye camera to collect distorted images of the target across the entire field. Through data cleaning and annotation, a distorted image dataset is formed. Step S2, Nonlinear Imaging Model Construction: The three-dimensional field-of-view distribution of the target is converted into a two-dimensional pixel array of the image, and a nonlinear mapping relationship is established between the object points in the three-dimensional space and the distorted two-dimensional pixels; Step S3: Construction of the Distorted Target Recognition Model: Construct a distorted target recognition model based on a deep learning network, optimize the deep learning network framework, alleviate the contradiction between the translational equivariance of traditional deep learning networks and the radial symmetry of fisheye images, and establish a precise mapping relationship between the two-dimensional pixel domain and the semantic features of the target. Step S4, Model Training and Recognition: Input the distorted image dataset into the target recognition model. After model training and parameter optimization, target recognition is achieved across the entire domain.

[0011] Furthermore, in step S1, the fisheye camera is a wide-angle camera with a field of view of not less than 180°, which can acquire scene information within a wide field of view at once; the global target refers to targets of different scales, azimuth angles, and radial distances from the center to the edge of the image within the effective imaging range of the fisheye camera; the distorted image refers to the nonlinear projection design adopted by the fisheye camera to achieve ultra-wide-angle imaging of more than 180°, where the incident light deviates from the linear projection law after being refracted by the lens, causing the target to produce distortions such as radial stretching, edge compression, and directional shift during the imaging process.

[0012] Furthermore, in step S2, the nonlinear imaging model is represented by an equidistant projection model, and its imaging relationship is as follows:

[0013] In the formula, such as Figure 2 As shown, The radial distance from the image point to the center of the image. For the focal length of the fisheye lens, The angle between the incident ray and the optical axis. For the true angle of incidence Distortion angle of incidence after mapping by the fisheye distortion model.

[0014] Based on the centrally symmetric imaging characteristics of fisheye nonlinear imaging about the optical axis, the distorted incident angle... With the angle of incidence Since it satisfies the odd function relation, it can be expanded using an odd power polynomial, and its Taylor series form can be expressed as:

[0015] In the formula, is the term number of the polynomial expansion term.

[0016] Since low-order polynomials are insufficient to fully describe the strong nonlinear distortion in the edge region of a fisheye camera, while high-order polynomials significantly increase the computational complexity of the model, are overly sensitive to noise and errors, and prone to overfitting, in order to balance the accuracy of the distortion model and computational efficiency, the Taylor series is retained up to the ninth degree, resulting in the general model representation of a fisheye camera as follows:

[0017] Substituting equation (3) into (1), we can obtain the projection function of the distorted points in the fisheye image:

[0018] Based on the above projection function of the distortion points in the fisheye image, a nonlinear mapping relationship between three-dimensional spatial object points and two-dimensional pixels in the image is established, and the transformation relationship is as follows: World coordinate system ( (to camera coordinate system) The conversion relationship is as follows:

[0019] In the formula, ( () represents the coordinates of a point in space in the world coordinate system. () represents the coordinates of a point in space in the camera coordinate system. R For the camera rotation matrix, t It is a translation vector; The transformation relationship from the camera coordinate system to the image coordinate system (x, y) is as follows:

[0020] Pixel coordinate system Transformation between the coordinate system and the image coordinate system:

[0021] In the formula, dx , dy The actual size of the pixel. , , and Indicates the lens focal length The pixelated representation of the image in the x and y directions, i.e., the equivalent focal length in pixels; These are the coordinates of the camera's principal point. This represents the camera's intrinsic parameters. Based on the above formula, a nonlinear mapping relationship is established between 3D spatial object points and distorted 2D pixels.

[0022] Furthermore, in step S3, a distortion target recognition model based on a deep learning network is constructed based on the distortion features generated by the target during nonlinear imaging. The deep learning network specifically includes an input layer, a distortion feature extraction layer, a feature fusion layer, and a target recognition output layer. Furthermore, the input layer is used to receive distorted image information, normalize the image to a preset size, and then input it to the distortion feature extraction layer.

[0023] Furthermore, the distortion feature extraction layer specifically involves constructing a ring domain convolution kernel based on the nonlinear imaging model built in step S2, targeting the distortion features generated by the target during the nonlinear imaging process, and using a circular scanning method to extract features across the entire image domain.

[0024] Furthermore, the aforementioned annular convolution kernel refers to taking each coordinate point in the coordinate matrix of the standard convolution kernel as the coordinate point to be transformed, and rotating it around the coordinate center point of the convolution kernel. After rotation, a new kernel coordinate space is obtained; then, inverse mapping and bilinear interpolation are performed on the rotated kernel coordinates to obtain the rotated kernel weights (for the coordinate matrix parameters of the standard kernel; first, the coordinates are rotated around the center of the kernel coordinate system). The rotation is then calculated to obtain a new kernel coordinate space. Subsequently, the rotated coordinates are mapped back to the standard convolution kernel's coordinate space via inverse mapping. Then, the weight values ​​at adjacent integer grid coordinates in the standard convolution kernel are sampled, and bilinear interpolation is used to obtain the weight values ​​corresponding to the target coordinates. Finally, the rotation is calculated. (degree of ring-domain convolution kernel).

[0025] Furthermore, let the standard convolution kernel be... The size is The corresponding weight matrix is:

[0026] In the formula, The initial weights are denoted as k, which are updated during model training via backpropagation based on the loss function. In the formula, k... ij Let represent the learnable weight parameter (i, j = 1, 2, 3) at the i-th row and j-th column position in the standard convolutional kernel k. The initial value of the weight parameter can be obtained by random initialization, parameter initialization algorithm, or pre-trained model parameters, and is iteratively updated during model training through backpropagation.

[0027] Using the center point of the standard convolution kernel as the origin of the coordinate system, the coordinate matrix corresponding to the standard convolution kernel is established as follows:

[0028] In the formula, The coordinate matrix of the standard convolution kernel. You can choose any one of the coordinates.

[0029] Furthermore, the rotation The specific calculation process is as follows: extract any feature location in the entire image domain. =( Its azimuth angle relative to the center of the fisheye is Then the rotation angle of the rotated convolution kernel Defined as:

[0030] Furthermore, the fisheye center refers to the geometric center of the effective imaging area of ​​the fisheye; the coordinates of the fisheye center are obtained by calculating the center of the effective imaging area of ​​the fisheye. .

[0031] Furthermore, the standard convolution kernel is rotated using a rotation matrix, which is:

[0032] Then the standard convolution kernel any coordinate in After rotation angle New coordinates It can be represented as:

[0033] The coordinates obtained by rotation are inversely rotated and mapped to their corresponding positions in the original convolution kernel. The formula is as follows:

[0034] Due to the coordinates after inverse rotation mapping The coordinates are generally non-integer coordinates, requiring bilinear interpolation to calculate the weights of the rotated convolution kernel. Therefore, the rotated convolution kernel will be positioned... The weights are:

[0035] In the formula, Inverse mapping coordinates The four nearest adjacent integer grid coordinates, and The index variable value is used for variable traversal. Using the surrounding integer coordinates, bilinear interpolation is performed on the four nearest neighbors. and Indicates the offset distance; Representation and position The weight values ​​at the four nearest integer grid coordinates; Let represent the weight matrix after rotation.

[0036] Furthermore, the circular scan refers to taking the center of the fisheye image as the starting position, fixing the current radial distance, and scanning in angular steps. Complete a full-angle circular scan, then proceed radially... The step size is increased outward to the next radial distance, and the circular scan is repeated until the scan ends when it slides out of the fisheye effective imaging area; Furthermore, using each sampling center of the circular scan as the center of the convolution operation, the annular convolution kernel obtained above is used to process the input feature map. Perform convolution calculations to obtain the convolution output corresponding to the sampling center. The calculation formula is as follows:

[0037] In the formula, Indicates sampling center Convolution output at the point; Indicates the input feature map of the th Each channel is located in Eigenvalues ​​at; This indicates the number of channels in the input feature map; Indicates position The rotational convolution kernel corresponding to its azimuth angle is at the 1st Weights on each channel; This represents the local coordinate offset within the rotated convolution kernel.

[0038] Furthermore, the feature fusion layer is used to fuse features from different levels output by the distortion feature extraction layer, thereby enhancing the network's ability to represent targets of different scales and radial positions in fisheye images. Let the multi-level feature tensor output by the distortion feature extraction layer be:

[0039] In the formula, They represent the first The height, width, and number of channels of each level feature tensor.

[0040] For the first Each level of feature tensor After performing size alignment and channel mapping, the following calculations were performed:

[0041] In the formula, Represents the 1st generation after size alignment and channel mapping. Each level of feature tensor This represents the corresponding scaling and channel mapping functions. Further, the aligned feature tensors from each level are fused to obtain the fused features:

[0042] By using the aforementioned feature fusion method, the network can comprehensively utilize the local detail information in shallow features and the semantic information in deep features, thereby improving its ability to recognize targets of different scales and targets in strongly distorted regions in fisheye images.

[0043] Furthermore, after feature fusion is completed, the fused feature map is output. The target recognition output layer includes a classification branch, a regression branch, and a confidence branch. The classification branch outputs the target category probability, the regression branch outputs the target bounding box parameters (target center coordinates, length and width dimensions, and rotation angle) adapted to fisheye distortion, and the confidence branch outputs the target existence probability, thus achieving accurate identification and localization of distorted targets.

[0044] Further, in step S4, the distorted image dataset constructed in step S1 is used to train the target recognition model constructed in step S3. Specifically, the training samples are input into the target recognition model, and the target category, location, and confidence prediction results are obtained through forward propagation. A loss function is then constructed based on the difference between the prediction results and the true annotations.

[0045] Furthermore, the loss function includes classification loss, location regression loss, and confidence loss, and its total loss function can be expressed as:

[0046] In the formula, Represents classification loss. This represents the position regression loss. Indicates confidence loss. This represents the weighting coefficient corresponding to each loss term.

[0047] Furthermore, the backpropagation algorithm is used to calculate the gradient of the network parameters, and an optimization algorithm is used to iteratively update the network parameters. The update process can be expressed as follows:

[0048] In the formula, Indicates the first Network parameters at the next iteration Indicates the learning rate. Indicates the first The network parameters are updated in the next iteration.

[0049] Furthermore, after completing model training and parameter optimization, the test samples are input into the optimized target recognition model to detect and recognize targets in the entire range of the fisheye image, and output the target category, target location and corresponding confidence information, thereby realizing the full-range target recognition of the test samples.

[0050] Example 1: This invention provides a global target recognition method based on annular domain convolution to address the mismatch between the translational equivariance of traditional deep learning networks and the radial symmetry of fisheye images; such as Figure 1 As shown, the method includes the following steps: Step S1: Building a Distorted Image Dataset: Use a fisheye camera to collect distorted images of the target across the entire field, and through data cleaning and annotation, form a distorted image dataset; Step S2: Nonlinear imaging model construction: The three-dimensional field of view distribution of the target is converted into a two-dimensional pixel array of the image, and a nonlinear mapping relationship is established between the object points in three-dimensional space and the distorted two-dimensional pixels; Step S3: Construction of the distorted target recognition model: Optimize the deep learning network framework to alleviate the contradiction between the translational equivariance of traditional deep learning networks and the radial symmetry of fisheye images, and establish a precise mapping relationship between the two-dimensional pixel domain and the semantic features of the target. Step S4: Model Training and Recognition: Input the distorted image dataset into the target recognition model. After model training and parameter optimization, target recognition is achieved across the entire domain.

[0051] In step S1, the fisheye camera is a wide-angle camera with a field of view of not less than 180°, which can acquire scene information within a wide field of view at once; the global target refers to targets of different scales, azimuth angles, and radial distances from the center to the edge of the image within the effective imaging range of the fisheye camera; the distorted image refers to the nonlinear projection design adopted by the fisheye camera to achieve ultra-wide-angle imaging of more than 180°, where the incident light deviates from the linear projection law after being refracted by the lens, causing the target to produce distortions such as radial stretching, edge compression, and directional shift during the imaging process.

[0052] In step S2, the nonlinear imaging model is represented by an equidistant projection model, and its imaging relationship is as follows:

[0053] In the formula, such as Figure 2 As shown, The radial distance from the image point to the center of the image. For the focal length of the fisheye lens, The angle between the incident ray and the optical axis. For the true angle of incidence Distortion angle of incidence after mapping by the fisheye distortion model.

[0054] Based on the centrally symmetric imaging characteristics of fisheye nonlinear imaging about the optical axis, the distorted incident angle... With the angle of incidence Since it satisfies the odd function relation, it can be expanded using an odd power polynomial, and its Taylor series form can be expressed as:

[0055] In the formula, is the term number of the polynomial expansion term.

[0056] Since low-order polynomials are insufficient to fully describe the strong nonlinear distortion in the edge region of a fisheye camera, while high-order polynomials significantly increase the computational complexity of the model, are overly sensitive to noise and errors, and prone to overfitting, in order to balance the accuracy of the distortion model and computational efficiency, the Taylor series is retained up to the ninth degree, resulting in the general model representation of a fisheye camera as follows:

[0057] Substituting equation (3) into (1), we can obtain the projection function of the distorted points in the fisheye image:

[0058] Based on the above projection function of the distortion points in the fisheye image, a nonlinear mapping relationship between three-dimensional spatial object points and two-dimensional pixels in the image is established, and the transformation relationship is as follows: World coordinate system ( (to camera coordinate system) The conversion relationship is as follows:

[0059] In the formula, ( () represents the coordinates of a point in space in the world coordinate system. () represents the coordinates of a point in space in the camera coordinate system. R For the camera rotation matrix, t It is a translation vector; The transformation relationship from the camera coordinate system to the image coordinate system (x, y) is as follows:

[0060] Pixel coordinate system Transformation between the coordinate system and the image coordinate system:

[0061] In the formula, dx , dy The actual size of the pixel. and Indicates the lens focal length The pixelated representation of the image in the x and y directions, i.e., the equivalent focal length in pixels; These are the coordinates of the camera's principal point. This refers to the camera's internal parameters.

[0062] In step S3, a deep learning-based target recognition model is constructed based on the distortion features generated by the target during nonlinear imaging. The deep learning network specifically includes an input layer, a distortion feature extraction layer, a feature fusion layer, and a target recognition output layer. The input layer is used to receive distorted image information, normalize the image to a preset size, and then input it to the distortion feature extraction layer.

[0063] Specifically, the distortion feature extraction layer is constructed based on the nonlinear imaging model built in step S2. For the distortion features generated by the target during the nonlinear imaging process, a ring domain convolution kernel is constructed, and a circular scanning method is used to extract features from the entire image domain.

[0064] The aforementioned annular convolution kernel refers to a kernel that takes the coordinate points in the coordinate matrix of a standard convolution kernel as the coordinate points to be transformed and rotates around the coordinate center point of the convolution kernel. After rotation, a new kernel coordinate space is obtained; then, the rotated kernel coordinates are inversely mapped and bilinearly interpolated to obtain the rotated kernel weights. As shown in the first figure of Figure 3, the standard convolutional kernel size is assumed to be 3x3; Figure 3 The second figure shows a coordinate rotation transformation, which rotates the aforementioned coordinate points around the center of the convolution kernel. The new kernel coordinate space is obtained by rotating the kernel. At this point, the coordinates of the rotated kernel are no longer aligned with the integer grid of the original kernel and cannot directly correspond to the original weight values. Figure 3 The third figure illustrates inverse mapping and bilinear interpolation. To obtain the weights of the rotated convolutional kernel, the rotated coordinates need to be inversely mapped back to the coordinate space of the standard convolutional kernel. Since the coordinates after inverse mapping are usually not integers, the original weight values ​​cannot be directly obtained. Therefore, bilinear interpolation is needed to calculate the corresponding convolutional kernel weight values ​​after the rotation. Figure 3 As shown in the fourth figure, the interpolated weight values ​​are used to form a rotated convolution kernel.

[0065] Furthermore, let the standard convolution kernel be... Shape The corresponding weight matrix is:

[0066] In the formula, These are the initial weight values, which are continuously updated during model training using the backpropagation algorithm based on the loss function.

[0067] The center point of the standard convolution kernel is used as the origin of the coordinate system, and the coordinate matrix corresponding to the standard convolution kernel is established as follows:

[0068] In the formula, The coordinate matrix of the standard convolution kernel. You can choose any one of the coordinates.

[0069] Furthermore, the rotation angle The specific calculation process involves extracting features at any location across the entire image domain. =( Its azimuth angle relative to the center of the fisheye is Then the rotation angle of the rotated convolution kernel Defined as:

[0070] The fisheye center refers to the geometric center of the circular boundary of the effective imaging region of the fisheye. In this embodiment, the Hough circle detection algorithm is used to locate the fisheye center in the preprocessed fisheye image. By detecting and fitting the circular boundary of the effective imaging region of the fisheye, the coordinates of the center of the detected and fitted circle are used as the coordinates of the fisheye center. The specific calculation process is as follows: The pixel coordinates of a fisheye image are The pixel coordinates of the center of the fisheye are ( , ), all edge pixels of the effective fisheye region ( ) to the center ( , The Euclidean distance of ) is equal to the radius r 0, that is:

[0071] Each edge pixel can be obtained through deformation. The parametric space curve equation of ) is:

[0072] In the formula, each edge point in the image space will be represented by the formula. In a three-dimensional parameter space, a trajectory corresponding to the "possible circle parameters" is formed. Since multiple edge points belong to the same effective fisheye imaging circle, the parameter corresponding to the peak value of the accumulated values ​​of pixels that meet the conditions is the center of the effective fisheye region. With radius r 0. To avoid multi-circle interference and improve detection robustness, the radius is... r The value of 0 is constrained to be between the half-width and half-height of the fisheye image.

[0073] Furthermore, the fisheye center refers to the geometric center of the effective imaging area of ​​the fisheye; the coordinates of the fisheye center are obtained by calculating the center of the effective imaging area of ​​the fisheye. .

[0074] Furthermore, the standard convolution kernel is rotated using a rotation matrix, which is:

[0075] Then the standard convolution kernel any coordinate in After rotation angle New coordinates It can be represented as:

[0076] The coordinates obtained by rotation are inversely rotated and mapped to their corresponding positions in the original convolution kernel. The formula is as follows:

[0077] Due to the coordinates after inverse rotation mapping The coordinates are generally non-integer coordinates, requiring bilinear interpolation to calculate the weights of the rotated convolution kernel. Therefore, the rotated convolution kernel will be positioned... The weights are:

[0078] In the formula, Inverse mapping coordinates The four nearest adjacent integer grid coordinates, and The index variable value is used for variable traversal. Using the surrounding integer coordinates, bilinear interpolation is performed on the four nearest neighbors. and Indicates the offset distance; Representation and position The weight values ​​at the four nearest integer grid coordinates; Let represent the weight matrix after rotation.

[0079] Furthermore, the circular scan refers to taking the center of the fisheye image as the starting position, fixing the current radial distance, and scanning in angular steps. Complete a full-angle circular scan, then proceed radially... The step size is increased outward to the next radial distance, and the circular scan is repeated until the scan ends when it slides out of the fisheye effective imaging area; Furthermore, using each sampling center of the circular scan as the center of the convolution operation, the annular convolution kernel obtained above is used to process the input feature map. Perform convolution calculations to obtain the convolution output corresponding to the sampling center. The calculation formula is as follows:

[0080] In the formula, Indicates sampling center Convolution output at the point; Indicates the input feature map of the th Each channel is located in Eigenvalues ​​at; This indicates the number of channels in the input feature map; Indicates position The rotational convolution kernel corresponding to its azimuth angle is at the 1st Weights on each channel; This represents the local coordinate offset within the rotated convolution kernel.

[0081] Furthermore, the feature fusion layer is used to fuse features from different levels output by the distortion feature extraction layer, thereby enhancing the network's ability to represent targets of different scales and radial positions in fisheye images. Let the multi-level feature tensor output by the distortion feature extraction layer be:

[0082] In the formula, They represent the first The height, width, and number of channels of each level feature tensor.

[0083] For the first Each level of feature tensor After performing size alignment and channel mapping, the following calculations were performed:

[0084] In the formula, Represents the 1st generation after size alignment and channel mapping. Each level of feature tensor This represents the corresponding scaling and channel mapping functions. Further, the aligned feature tensors from each level are fused to obtain the fused features:

[0085] By using the aforementioned feature fusion method, the network can comprehensively utilize the local detail information in shallow features and the semantic information in deep features, thereby improving its ability to recognize targets of different scales and targets in strongly distorted regions in fisheye images.

[0086] Furthermore, after feature fusion is completed, the fused feature map is output. The target recognition output layer includes a classification branch, a regression branch, and a confidence branch. The classification branch outputs the target category probability, the regression branch outputs the target bounding box parameters (target center coordinates, length and width dimensions, and rotation angle) adapted to fisheye distortion, and the confidence branch outputs the target existence probability, thus achieving accurate identification and localization of distorted targets.

[0087] Further, in step S4, the distorted image dataset constructed in step S1 is used to train the target recognition model constructed in step S3. Specifically, the training samples are input into the target recognition model, and the target category, location, and confidence prediction results are obtained through forward propagation. A loss function is then constructed based on the difference between the prediction results and the true annotations.

[0088] Furthermore, the loss function includes classification loss, location regression loss, and confidence loss, and its total loss function can be expressed as:

[0089] In the formula, Represents classification loss. This represents the position regression loss. Indicates confidence loss. This represents the weighting coefficient corresponding to each loss term.

[0090] Furthermore, the backpropagation algorithm is used to calculate the gradient of the network parameters, and an optimization algorithm is used to iteratively update the network parameters. The update process can be expressed as follows:

[0091] In the formula, Indicates the first Network parameters at the next iteration Indicates the learning rate. Indicates the first The network parameters are updated in the next iteration.

[0092] Furthermore, after completing model training and parameter optimization, the test samples are input into the optimized target recognition model to detect and recognize targets in the entire range of the fisheye image, and output the target category, target location and corresponding confidence information, thereby realizing the full-range target recognition of the test samples.

[0093] Beneficial effects: This invention directly uses fisheye images as input, which can more completely preserve the wide field of view information and original distortion features of fisheye images. At the same time, by constructing a ring-domain convolution kernel, the direction of the convolution kernel can be adaptively adjusted according to the orientation of the target position relative to the center of the fisheye, which improves the matching between convolution feature extraction and the radial distortion distribution of the fisheye, and enhances the feature representation ability of targets in strongly distorted regions. Furthermore, by combining a circular scanning method to extract features across the entire range of the fisheye image, full coverage detection of the effective imaging area of ​​the fisheye is achieved. By integrating shallow detail information and deep semantic information through multi-level feature fusion, the recognition accuracy and robustness of targets of different scales, targets at different radial positions, and targets with complex distortions are improved, ultimately enabling accurate recognition and localization of targets in fisheye images. While ensuring detection accuracy, the invention reduces computational complexity and improves the real-time performance of target recognition, making it suitable for application scenarios with high real-time requirements such as drone photography and autonomous driving.

[0094] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A target global recognition method based on annular domain convolution, characterized in that, The method includes the following steps: Step S1: Building a Distorted Image Dataset: Use a fisheye camera to collect distorted images of the target across the entire field. Through data cleaning and annotation, a distorted image dataset is formed. Step S2, Nonlinear Imaging Model Construction: The three-dimensional field-of-view distribution of the target is converted into a two-dimensional pixel array of the image, and a nonlinear mapping relationship is established between the object points in the three-dimensional space and the distorted two-dimensional pixels; Step S3: Construction of the Distorted Target Recognition Model: Construct a distorted target recognition model based on a deep learning network, optimize the deep learning network framework, alleviate the contradiction between the translational equivariance of traditional deep learning networks and the radial symmetry of fisheye images, and establish a precise mapping relationship between the two-dimensional pixel domain and the semantic features of the target. Step S4, Model Training and Recognition: Input the distorted image dataset into the target recognition model. After model training and parameter optimization, target recognition is achieved across the entire domain.

2. The target global recognition method based on ring domain convolution as described in claim 1, characterized in that, In step S1, the fisheye camera is a wide-angle camera with a field of view of not less than 180°, which can acquire scene information within a wide field of view at once; the global target refers to targets of different scales, azimuth angles, and radial distances from the center to the edge of the image within the effective imaging range of the fisheye camera; the distorted image refers to the nonlinear projection design adopted by the fisheye camera to achieve ultra-wide-angle imaging of more than 180°, where the incident light deviates from the linear projection law after being refracted by the lens, causing radial stretching, edge compression, and directional shift distortion of the target during the imaging process.

3. The target global recognition method based on annular domain convolution as described in claim 1, characterized in that, In step S2, the nonlinear imaging model is represented by an equidistant projection model, and its imaging relationship is as follows: In the formula, The radial distance from the image point to the center of the image. For the focal length of the fisheye lens, The angle between the incident ray and the optical axis. For the true angle of incidence Distortion angle of incidence mapped by the fisheye distortion model; Based on the centrally symmetric imaging characteristics of fisheye nonlinear imaging about the optical axis, the distorted incident angle... With the angle of incidence Since it satisfies the odd function relation, it can be expanded using an odd power polynomial, and its Taylor series form is as follows: In the formula, The term number is the term index of the polynomial expansion. Retaining the above Taylor series up to the ninth degree, we obtain the general model representation of a fisheye camera as follows: Substituting equation (3) into (1), we obtain the projection function of the distorted points in the fisheye image: Based on the above projection function of the distortion points in the fisheye image, a nonlinear mapping relationship between three-dimensional spatial object points and two-dimensional pixels in the image is established.

4. The target global recognition method based on annular domain convolution as described in claim 1, characterized in that, Establishing a nonlinear mapping relationship between three-dimensional spatial points and two-dimensional image pixels includes: World coordinate system ( (to camera coordinate system) The conversion relationship is as follows: In the formula, ( () represents the coordinates of a point in space in the world coordinate system. () represents the coordinates of a point in space in the camera coordinate system. R For the camera rotation matrix, t It is a translation vector; The transformation relationship from the camera coordinate system to the image coordinate system (x, y) is as follows: Pixel coordinate system Transformation between the coordinate system and the image coordinate system: In the formula, dx , dy The actual size of the pixel. , , and Indicates the lens focal length The pixelated representation of the image in the x and y directions, i.e., the equivalent focal length in pixels; The coordinates of the camera's principal point; The camera intrinsic parameters are used; based on the above formula, a nonlinear mapping relationship is established between three-dimensional spatial object points and distorted two-dimensional pixels.

5. The target global recognition method based on annular domain convolution as described in any one of claims 1-4, characterized in that, In step S3, a distortion target recognition model based on a deep learning network is constructed based on the distortion features generated by the target during nonlinear imaging. The deep learning network specifically includes an input layer, a distortion feature extraction layer, a feature fusion layer, and a target recognition output layer. The input layer is used to receive distorted image information, normalize the image to a preset size, and then input it to the distortion feature extraction layer. Specifically, the distortion feature extraction layer is constructed based on the nonlinear imaging model built in step S2. For the distortion features generated by the target during the nonlinear imaging process, a ring domain convolution kernel is constructed, and a circular scanning method is used to extract features from the entire image domain. The feature fusion layer is used to fuse features from different levels output by the distortion feature extraction layer to enhance the network's ability to represent targets of different scales and radial positions in fisheye images. After feature fusion is completed, the fused feature map is output to the target recognition output layer; The target recognition output layer includes a classification branch, a regression branch, and a confidence branch. The classification branch outputs the target category probability, the regression branch outputs the target bounding box parameters adapted to fisheye distortion, and the confidence branch outputs the target existence probability, thus achieving accurate identification and localization of distorted targets.

6. The target global recognition method based on annular domain convolution as described in claim 5, characterized in that, In S3, the annular convolution kernel refers to taking each coordinate point in the coordinate matrix of the standard convolution kernel as the coordinate point to be transformed, and rotating it around the coordinate center point of the convolution kernel. After rotation, a new kernel coordinate space is obtained; then, the rotated kernel coordinates are inversely mapped and bilinearly interpolated to obtain the rotated kernel weight values.

7. The target global recognition method based on annular domain convolution as described in claim 6, characterized in that, let... The standard convolutional kernel The size is The corresponding weight matrix is: In the formula, These are the initial weight values, which are updated via backpropagation based on the loss function during model training. Using the center point of the standard convolution kernel as the origin of the coordinate system, the coordinate matrix corresponding to the standard convolution kernel is established as follows: In the formula, The coordinate matrix of the standard convolution kernel. Take any one of the coordinates; Furthermore, the rotation The specific calculation process is as follows: extract any feature location in the entire image domain. =( Its azimuth angle relative to the center of the fisheye is Then the rotation angle of the rotated convolution kernel Defined as: The fisheye center refers to the geometric center of the effective imaging area of ​​the fisheye; the coordinates of the fisheye center are obtained by calculating the center of the effective imaging area of ​​the fisheye. ; The standard convolution kernel is rotated using a rotation matrix, which is: Then the standard convolution kernel any coordinate in After rotation angle New coordinates Represented as: The coordinates obtained by rotation are inversely rotated and mapped to their corresponding positions in the original convolution kernel. The formula is as follows: Due to the coordinates after inverse rotation mapping For non-integer coordinates, the weights of the rotated convolution kernel are calculated using bilinear interpolation. Therefore, the rotated convolution kernel will be positioned... The weights are: In the formula, Inverse mapping coordinates The four nearest adjacent integer grid coordinates, and The index variable value is used for variable traversal. Using the surrounding integer coordinates, bilinear interpolation is performed on the four nearest neighbors. and Indicates the offset distance; Representation and position The weight values ​​at the four nearest integer grid coordinates; Let represent the weight matrix after rotation.

8. The target global recognition method based on annular domain convolution as described in claim 1, characterized in that, In S3, the circular scan refers to taking the center of the fisheye image as the starting position, fixing the current radial distance, and scanning in angular steps. Complete a full-angle circular scan, then proceed radially... The step size is increased outward to the next radial distance, and the circular scan is repeated until the scan ends when it slides out of the fisheye effective imaging area; Using each sampling center of the circular scan as the center of the convolution operation, the annular convolution kernel obtained above is used to process the input feature map. Perform convolution calculations to obtain the convolution output corresponding to the sampling center. The calculation formula is as follows: In the formula, Indicates sampling center The convolution output at the point; Indicates the input feature map of the th Each channel is located in Eigenvalues ​​at; This indicates the number of channels in the input feature map; Indicates position The rotational convolution kernel corresponding to its azimuth angle is at the 1st Weights on each channel; This represents the local coordinate offset within the rotated convolution kernel.

9. The target global recognition method based on annular domain convolution as described in claim 5, characterized in that, In S3, let the multi-level feature tensor output by the distortion feature extraction layer be: In the formula, They represent the first The height, width, and number of channels of each level of feature tensor; For the Each level of feature tensor After performing size alignment and channel mapping, the following calculations were performed: In the formula, Represents the 1st generation after size alignment and channel mapping. Each level of feature tensor This represents the corresponding scaling and channel mapping functions; the aligned feature tensors from each level are fused to obtain the fused features: Through the above feature fusion method, the network can comprehensively utilize the local detail information in shallow features and the semantic information in deep features.

10. The target global recognition method based on ring domain convolution as described in claim 5, characterized in that, In step S4, the distorted image dataset constructed in step S1 is used to train the target recognition model constructed in step S3; the training samples are input into the target recognition model, and the target category, location and confidence prediction results are obtained through forward propagation, and a loss function is constructed based on the difference between the prediction results and the real annotations. The loss function includes classification loss, location regression loss, and confidence loss, and its total loss function is expressed as follows: In the formula, Represents classification loss, This represents the position regression loss. Indicates confidence loss. This represents the weighting coefficient corresponding to each loss term; The backpropagation algorithm is used to calculate the gradient of the network parameters, and an optimization algorithm is used to iteratively update the network parameters. The update process is expressed as follows: In the formula, Indicates the first Network parameters at the next iteration Indicates the learning rate. Indicates the first Network parameters updated in the next iteration; After completing model training and parameter optimization, the test samples are input into the optimized target recognition model to detect and recognize targets in the entire range of the fisheye image, and output the target category, target location and corresponding confidence information, thereby realizing the full-range target recognition of the test samples.