Hyperbolic Space Deformable Convolution Method for Top-View Fish-Eye Vision Tasks
By applying the deformable convolution method in hyperbolic space, the problem of insufficient distortion representation of top-view fish eye images is solved, and the accuracy and feature extraction ability of top-view fish eye vision tasks of convolutional neural networks are improved.
Patent Information
- Application Number
- CN202211433024.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-11-16
AI Technical Summary
The existing convolutional neural networks lack the ability to characterize distortion in top-view fisheye images, resulting in low target perception accuracy in top-view fisheye visual tasks, and the existing dedistortion method has a large amount of calculation or reduces the field of view.
The hyperbolic space deformable convolution method is adopted to embed the top-view fish eye feature map into the Poincaré sphere, update the feature information in the Poincaré sphere space through the graph convolution neural network, and construct the Poincaré hyperplane to calculate the distance from the feature to the hyperplane, extract the distance information as the deformation parameter of the deformable convolution, and improve the adaptability of the convolution kernel.
The accuracy of convolutional neural networks in top-view fisheye visual tasks is improved, the ability to capture distorted features is enhanced, and feature extraction ability is improved.
Smart Images

Figure CN115731126B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and specifically relates to a hyperbolic space deformable convolution method for top-view fisheye vision tasks. Background Art
[0002] Optimizing the distortion in top-view fisheye images is crucial for improving the accuracy of target perception. Compared with other distorted images, the distortion problem of top-view fisheye has certain particularity. It does not conform to any current distortion model. Currently, the research usually uses a fourth-order polynomial mapping as an approximation of the fisheye camera model. However, in computer vision tasks, such an assumption lacks calibration information and also limits the field of view of the fisheye camera. And the deformable convolution proposed for distorted image vision tasks still has insufficient representation ability for fisheye image distortion. Using means such as calibration plates and unfolding methods to correct the distortion of top-view fisheye images is a common idea, but this method will introduce more computational complexity and reduce the field of view or integrity of the top-view fisheye images.
[0003] Using deformable convolution to optimize the convolutional neural network is an effective method to make the deep learning model adapt to image distortion. Currently, there are also various variants of applying deformable convolution to top-view fisheye images. For example, learning different weights for the grid sampling positions of each standard convolution, or restricting the bias learning of the grid sampling center positions of the standard convolution. The learning of these grid sampling weights and biases is achieved by adding an ordinary convolutional neural network, and all are optimized in the Euclidean space. However, fisheye images are similar to spherical signals. If they are mapped to the Riemannian space for optimization, the adaptability of the deformable convolution to image distortion can be effectively increased. Summary of the Invention
[0004] To solve the above problems, the present invention discloses a hyperbolic space deformable convolution method for top-view fisheye vision tasks. Based on the existing convolution operations, the hyperbolic space deformable convolution method disclosed by the present invention can achieve a more accurate approximation of the fisheye model by making an analogy between the Poincaré ball model in the hyperbolic space and the fisheye projection. And there is a bijective mapping between the hyperbolic space and the Euclidean space, that is, a one-to-one correspondence can be achieved. Therefore, geodesic metrics and graph convolutional neural networks can be applied in the hyperbolic space to learn the deformation parameters of the deformable convolution kernel in the Euclidean space, so as to better capture the smooth non-linear distortion of the top-view fisheye images. Furthermore, improve its ability to extract features of top-view fisheye images and enhance the accuracy of the convolutional neural network in top-view fisheye vision tasks.
[0005] To achieve the above object, the technical solution of the present invention is as follows:
[0006] A hyperbolic space deformable convolution method for top-view fisheye vision tasks, comprising:
[0007] Step 1: Embed the top-view fisheye feature map into the Poincaré sphere;
[0008] Step 2: Map the feature information to the tangent space of the Poincaré sphere;
[0009] Step 3: Update the feature information using a graph convolutional neural network;
[0010] Step 4: Map the feature information back to the Poincaré sphere space;
[0011] Step 5: Construct the Poincaré hyperplane and calculate the distance from the feature to the hyperplane;
[0012] Step 6: Extract the distance information as the deformation parameter for the deformable convolution.
[0013] Furthermore, the top-view fisheye feature described in Step 1 belongs to image structure data and only performs convolution operations in the Euclidean space. To achieve the aggregation and update of features in the Poincaré sphere space, it is necessary to embed the top-view fisheye feature map into an undirected graph structure, convert the information of each pixel on the feature map into node information in the undirected graph, and realize the aggregation and update of node features in the Poincaré sphere space through the subsequent graph convolutional neural network. During the graph embedding process, the adjacency matrix A of the graph needs to be calculated according to the 8-connectivity of the image.
[0014] Furthermore, in Step 2, since the parameters in the hyperbolic space cannot be optimized using the graph convolutional neural network in the Euclidean space, the feature vector on the Poincaré sphere is first mapped to a vector in the tangent space of the Poincaré sphere through a logarithmic mapping. For a Poincaré sphere with a radius of c, the calculation formula for the logarithmic mapping of vector y at vector x is:
[0015]
[0016] where is the Möbius addition on the Poincaré sphere with a radius of c, and ||·|| represents the norm of the vector in the Euclidean space.
[0017] Furthermore, in Step 3, the degree matrix D of the graph can be calculated through the adjacency matrix A obtained in Step 1, and the calculation formula is:
[0018]
[0019] where D is a diagonal matrix, so only the elements on the diagonal need to be calculated. To implement the graph convolution operation, the Laplacian matrix of the graph needs to be calculated through the adjacency matrix A and the degree matrix D The calculation formula is:
[0020]
[0021] For each node in an undirected graph, a graph convolutional neural network updates the feature information of each node by aggregating the feature information from its neighboring nodes. Before feature aggregation, it is necessary to first perform a linear transformation on the feature information from its neighboring nodes. After feature aggregation, the information of this node is activated using the ReLU function. The calculation formula for the graph convolutional neural network to update the node feature information is as follows:
[0022]
[0023] where W is a learnable linear transformation matrix acting on the feature information of neighboring nodes, and N(x) is the neighborhood of node x.
[0024] Furthermore, in step four, after updating the node feature information through the graph convolutional neural network in the tangent space of the Poincaré sphere, it is necessary to map the vector in the tangent space back to the Poincaré sphere space through the exponential map. For a Poincaré sphere with radius c, the calculation formula for the exponential map of vector y at vector x is:
[0025]
[0026] Furthermore, in step five, in order to transform the node information in the Poincaré sphere space into the Euclidean space, 27 learnable hyperplanes are constructed in the Poincaré sphere space. For each node, the geodesic distance to the hyperplane is calculated to obtain 27 parameters. The calculation formula for constructing each hyperplane is:
[0027]
[0028] where p is the learnable hyperplane bias parameter, and a is the learnable hyperplane normal vector parameter. The calculation formula for the geodesic distance from the node feature vector to each hyperplane is:
[0029]
[0030] Furthermore, in step six, the 27 geodesic distance parameters obtained in step five are applied to modulate the deformable convolution. The calculation formula for the modulated deformable convolution is:
[0031]
[0032] where p is the position coordinate in the feature map, p k ∈{(-1,-1),(-1,-0),...,(1,0),(1,1)} is the position coordinate in the feature map. For example, (-1,-1) and (0,0) are the indices of the upper left corner and the center respectively, Δp k 、Δm kThey are the position offset and modulation coefficient in deformable convolution respectively. For each node, there are 18 position offset parameters and 9 modulation coefficients. Therefore, among the 27 geodesic distances, the first 18 geodesic distances are activated by the Sigmoid function and assigned to the offset parameters, and the last 9 geodesic distances are directly assigned to the modulation coefficients. The calculation formula of the Sigmoid activation function is:
[0033]
[0034] The beneficial effect of the present invention is that the hyperbolic space deformable convolution method enables the convolution kernel to learn the deformation parameters for modulating the deformable convolution in the Poincaré ball model in the hyperbolic space, allowing the convolutional neural network to learn more information from the distorted features of the top-view fisheye image and improving the accuracy of the convolutional neural network in the top-view fisheye vision task. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a processing flowchart of a hyperbolic space deformable convolution method for a top-view fisheye vision task provided by an embodiment of the present specification;
[0036] Figure 2 is a schematic diagram of the principle of the hyperbolic space deformable convolution method provided by an embodiment of the present specification;
[0037] Figure 3 is a schematic diagram of constructing an 8-connected region according to an image provided by an embodiment of the present specification;
[0038] Figure 4 is a schematic diagram of converting picture data into graph structure data provided by an embodiment of the present specification;
[0039] Figure 5 is a schematic diagram of constructing an adjacency matrix of a graph provided by an embodiment of the present specification. DETAILED DESCRIPTION OF THE INVENTION
[0040] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.
[0041] Embodiment 1
[0042] Figure 1 shows the processing flow of the hyperbolic space deformable convolution method for a top-view fisheye vision task provided by the present specification, Figure 2 shows the calculation principle of the hyperbolic space deformable convolution method:
[0043] S101. Embed the top-view fisheye feature map into the Poincaré ball
[0044] Since the current optimization methods for Riemannian spaces are not applicable to convolutional neural networks, it is considered to embed the feature map into a graph convolutional neural network and then perform Riemannian optimization based on the Poincaré disk on the feature information. The structure includes nodes and edges. As Figure 3 shown, the information contained in each pixel in the feature map constitutes a node, and each pixel is connected to the surrounding 8 pixels to form an edge, forming the graph structure feature information as shown in Figure 4 . To make the graph convolution method similar to the 3×3 convolution, the edges of the graph and the [N, N] adjacency matrix are constructed according to the 8-connectivity of the image, where N represents the total number of nodes. The values in the adjacency matrix are only 0 or 1, where 0 means that the corresponding row and column nodes are not adjacent, and 1 is the opposite. The constructed adjacency matrix is as shown in Figure 5 .
[0045] Specifically, the input feature map has dimensions [B, C_in, W_in, H_in], where B is the Batch dimension of the feature map, C_in is the input channel dimension of the feature map, and W_in and H_in are the width and height of the input feature map respectively. There are a total of W_in×H_in nodes in the constructed undirected graph, each node contains feature information of dimension C_in, and the adjacency matrix A of the undirected graph has dimensions [(W_in×H_in), (W_in×H_in)].
[0046] S102. Map the feature information to the tangent space of the Poincaré ball
[0047] Since graph convolutional neural networks can only be applied to Euclidean spaces, and the tangent space of the Poincaré ball is a Euclidean space, in order to apply graph convolutional neural networks to the vectors in the tangent space, the feature vectors on the Poincaré ball need to be mapped to the vectors in the tangent space of the Poincaré ball through logarithmic mapping.
[0048] Specifically, a Poincaré ball with a radius of 1 and dimension C_in is selected. For the feature y = f(x) at each node x in the undirected graph, the tangent vector of this feature to the origin of the Poincaré ball is calculated through logarithmic mapping. The calculation formula for logarithmic mapping at the origin is as follows:
[0049]
[0050] S103. Update the feature information by the graph convolutional neural network
[0051] Specifically, the degree matrix D of the graph can be calculated from the adjacency matrix A obtained in S101. The calculation formula is:
[0052]
[0053] Among them, D is a diagonal matrix, so only the elements on the diagonal need to be calculated. To implement the operation of graph convolution, the Laplacian matrix of the graph needs to be calculated through the adjacency matrix A and the degree matrix D The calculation formula is:
[0054]
[0055] For each node in the undirected graph, the graph convolutional neural network updates the feature information of each node by aggregating the feature information from its neighboring nodes. The output feature channel dimension of the graph convolutional neural network in the Poincaré ball is C_out. Since the graph convolutional neural network cannot change the structure of the graph, the width and height of the output features are the same as those of the input features, only changing the channel dimension C_out. Before feature aggregation, the feature information from its neighboring nodes needs to be linearly transformed. After feature aggregation, the information of this node is activated with the ReLU function. The calculation formula for the graph convolutional neural network to update the node feature information is:
[0056]
[0057] Among them, W is a learnable linear transformation matrix acting on the feature information of neighboring nodes, with a dimension of [C_out, C_in], and N(x) is the neighborhood of node x
[0058] S104. Map the feature information back to the Poincaré ball space
[0059] After updating the node feature information through the graph convolutional neural network in the tangent space of the Poincaré ball, the vector in the tangent space needs to be mapped back to the Poincaré ball space through the exponential map
[0060] Specifically, select a Poincaré ball with a radius of 1 and a dimension of C_out. For the feature y = f(x) at each node x in the undirected graph, calculate its exponential map at the origin of the Poincaré ball. The calculation formula is:
[0061]
[0062] S105. Construct the Poincaré hyperplane and calculate the distance from the feature to the hyperplane
[0063] Specifically, to transform the node information in the Poincaré ball space into the Euclidean space, 27 learnable hyperplanes are constructed in the Poincaré ball space. For each node, calculate the geodesic distance to the hyperplane to obtain 27 distance values. The calculation formula for constructing each hyperplane is:
[0064]
[0065] Among them, p is a learnable hyperplane bias parameter with a dimension of 1, and a is a learnable hyperplane normal vector parameter with a dimension of C_out. The calculation formula for the geodesic distance value from the node feature vector to each hyperplane is as follows:
[0066]
[0067] S106. Extract the distance information as the deformation parameter of the deformable convolution
[0068] Apply the 27 geodesic distance parameters obtained in S105 to the modulated deformable convolution. The calculation formula for the modulated deformable convolution is as follows:
[0069]
[0070] Among them, p is the position coordinate in the feature map, p k ∈{(-1, -1), (-1, -0),..., (1, 0), (1, 1)} is the position coordinate in the feature map. For example, (-1, -1) and (0, 0) are the indexes of the upper left corner and the center respectively. Δp k and Δm k are the position offset and modulation coefficient in the deformable convolution respectively. For each node, there are 18 position offset parameters and 9 modulation coefficients. Therefore, among the 27 geodesic distances, the first 18 geodesic distances are activated by the Sigmoid function and assigned to the offset parameters, and the last 9 geodesic distances are directly assigned to the modulation coefficients. The calculation formula for the Sigmoid activation function is as follows:
[0071]
[0072] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
[0073] Embodiment 2
[0074] Taking the top-view fisheye feature map with an input dimension of [64, 64, 32] as an example, and the output feature channel dimension of the graph convolutional neural network in the Poincaré ball is 16, the hyperbolic space deformable convolution method proposed by the present invention is described.
[0075] As Figure 2 shown, the top-view fisheye feature map is embedded into an undirected graph through the graph embedding method described in step S101. The undirected Figure 1It contains a total of 4096 nodes, each of which contains feature information with a dimension of 32. The edges of the graph and the adjacency matrix with a dimension of [4096, 4096] are constructed according to the 8-connectivity of the image. The values in the adjacency matrix are only 0 or 1, where 0 represents that the corresponding row and column nodes are not adjacent, and 1 is the opposite. The graph node features are mapped to the tangent space at the origin of the Poincaré ball with a radius of 1 and a dimension of 32 through the logarithmic mapping described in step S102. After calculating the degree matrix and Laplacian matrix of the graph by the method described in step S103, the result of the graph convolutional neural network is calculated according to the Laplacian matrix of the graph, and the output node feature dimension is 16. After mapping the graph node features to the Poincaré ball space with a radius of 1 and a dimension of 16 through the exponential mapping described in step S104, 27 learnable hyperplanes with a dimension of 16 are constructed in this Poincaré ball space. For each node, the geodesic distance from the hyperplane is obtained through the calculation method described in step S105, resulting in 27 distance values. Finally, through the method described in step S105, the first 18 geodesic distances are activated by the Sigmoid function and assigned to the offset parameters of the modulated deformable convolution, and the last 9 geodesic distances are directly assigned to the modulation coefficients of the modulated deformable convolution, and the modulated deformable convolution is performed on the input features to implement the hyperbolic space deformable convolution method.
Claims
1. A hyperbolic space deformable convolution method for top-view fisheye vision tasks, characterized in that, The following steps are included in sequence: Step 1: embed the top-view fisheye feature map into the Poincare sphere; Step 2: Map feature information to the Poincare sphere tangent space; Step 3: The graph convolutional neural network updates feature information; Step 4: Map the feature information back to the Poincare sphere space; Step 5: Construct the Poincare hyperplane and calculate the distance from the feature to the hyperplane; Step 6: Extract distance information as deformation parameters of deformable convolution; In the step 1, the top-view fisheye feature belongs to the image structure data, and the convolution operation is only implemented in the Euclidean space. In order to realize the aggregate update of the features in the Poincare sphere space, the top-view fisheye feature map is embedded into the undirected graph structure, so that the information of each pixel on the feature map is converted into the node information in the undirected graph, and the aggregate update of the node features is realized in the Poincare sphere space through the subsequent graph convolutional neural network; in the process of graph embedding, the adjacency matrix A of the graph is calculated according to the 8-connectivity of the image.
2. The hyperbolic space deformable convolution method for top-view fisheye vision tasks according to claim 1, wherein: In the step 2, since the parameters in the hyperbolic space cannot be optimized using the graph convolutional neural network in the Euclidean space, the eigenvector on the Poincare sphere is first mapped to a vector in the tangent space of the Poincare sphere through a logarithmic mapping; for a Poincare sphere with a radius of c, the calculation formula for the logarithmic mapping of the vector y at the vector x is: wherein, is the Möbius addition on the Poincaré sphere with radius c, and ||·|| represents the norm of a vector in Euclidean space.
3. The hyperbolic space deformable convolution method for top-view fisheye vision tasks according to claim 2, characterized in that: In step 4, the degree matrix D of the graph can be calculated using the adjacency matrix A calculated in step 1. The calculation formula is: Where D is a diagonal matrix, so only the elements on the diagonal need to be calculated; in order to implement the graph convolution operation, the Laplacian matrix of the graph is calculated through the adjacency matrix A and the degree matrix D The calculation formula is: For each node in the undirected graph, the graph convolutional neural network updates the feature information of each node by aggregating the feature information from its neighboring nodes; Before feature aggregation, the feature information from its neighboring nodes is linearly transformed; after feature aggregation, the information of the node is activated with the ReLU function. The calculation formula for updating the node feature information of the graph convolutional neural network is: Where W is a learnable linear transformation matrix that acts on the feature information of neighboring nodes, and N(x) is the domain of node x.
4. The hyperbolic space deformable convolution method for top-view fisheye vision tasks according to claim 3, characterized in that: In step 4, after updating the node feature information in the tangent space of the Poincare sphere through the graph convolutional neural network, the vector in the tangent space is mapped back to the Poincare sphere space through exponential mapping; for a Poincare sphere with a radius of c, the calculation formula for the exponential mapping of vector y at vector x is:
5. The hyperbolic space deformable convolution method for top-view fisheye vision tasks according to claim 4, wherein: In step 5, in order to transform the node information in the Poincare sphere space into the Euclidean space, 27 learnable hyperplanes are constructed in the Poincare sphere space, and the geodesic distance between each node and the hyperplane is calculated to obtain 27 parameters; wherein the calculation formula for constructing each hyperplane is: Among them, p is the learnable hyperplane bias parameter, a is the learnable hyperplane normal vector parameter; the calculation formula of the geodesic distance from the node feature vector to each hyperplane is:
6. The hyperbolic space deformable convolution method for top-view fisheye vision tasks according to claim 5, characterized in that: In step 6, the 27 geodesic distance parameters obtained in step 5 are applied to the modulated deformable convolution, wherein the calculation formula of the modulated deformable convolution is: where p is the position coordinate in the feature map, p k ∈ {(-1, -1), (-1, -0),..., (1, 0), (1, 1)} is the position coordinate in the feature map, Δp k , Δm k are the position offset and modulation coefficient in the deformable convolution respectively; for each node, there are 18 position offset parameters and 9 modulation coefficients. Therefore, among the 27 geodesic distances, the first 18 geodesic distances are assigned to the offset parameters after being activated by the Sigmoid function, and the last 9 geodesic distances are directly assigned to the modulation coefficients; the calculation formula of the Sigmoid activation function is:
Citation Information
Patent Citations
Self-calibration method for radial distortion of fish-eye lens camera
CN104036496A
Driver behavior identification method based on multi-scale attention convolutional neural network
CN110059582A