Image processing method, system and device based on spherical convolutional neural network

By using inverse North Pole projection and generalized Fourier transform in spherical convolution neural networks, the problems of edge stretching and compression when spherical images are expanded are solved, efficient spherical image processing and three-dimensional feature retention are achieved, and the accuracy and computing efficiency of image recognition are improved.

CN119991421AInactive Publication Date: 2025-05-13INFORMATION SCI RES INST OF CETC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510467259.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the prior art expands the spherical image into a plane, it causes stretching and compression to occur at the edge of the image, affecting imaging accuracy, and increasing computational complexity and processing time, making it difficult to achieve good translation and rotation invariance.

Method used

The image processing method based on spherical convolution neural network is adopted to project two-dimensional image data onto the spherical substrate through inverse north pole projection, spherical grid mapping is performed, and convolution calculation is performed in the frequency domain space using generalized Fourier transform, spherical harmonic function and Wigner D matrix, which avoids the edge problem of image expansion and reduces the computational complexity.

Benefits of technology

It realizes direct convolution operation on the spherical surface, avoids image edge stretching and compression, maintains the geometric characteristics of the three-dimensional image, and improves the computing efficiency and image recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991421A_ABST
    Figure CN119991421A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method based on a spherical convolutional neural network, and the method comprises the following steps: projecting input two-dimensional image data to a spherical substrate through an inverse north pole projection method, and forming a spherical image; mapping the spherical image into a spherical grid; and carrying out spherical convolution calculation according to the image in the spherical grid, specifically, converting the image in the spherical grid from a spatial domain to a frequency domain space through generalized Fourier transform, fitting the image and a convolution kernel in the spherical grid by adopting a spherical harmonic function and a Wigner D matrix as orthogonal bases, and carrying out convolution calculation on the image in the spherical grid. Performing inner product summation operation in the frequency domain space to complete spherical convolution calculation; and reducing a frequency domain convolution result to a geometric space through inverse Fourier transform to obtain an output feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to an image processing method, system and device based on a spherical convolutional neural network. Background Art

[0002] With the rapid development of artificial intelligence and computer vision technology, bionic electronic eyes are increasingly used in automation, intelligent monitoring, robotics and other fields. The design of bionic electronic eyes is inspired by biological visual systems, especially the visual abilities of humans and other animals.

[0003] The biological visual system achieves efficient information processing through complex neural networks and biological structures, enabling organisms to quickly identify objects and dynamic changes in the surrounding environment. The goal of the bionic electronic eye is to mimic this complex visual ability to achieve more efficient target recognition and environmental perception, thus playing an important role in a variety of application scenarios.

[0004] In the field of automation, bionic electronic eyes can improve the intelligence level of industrial robots and drones through high-precision target recognition and tracking functions. Intelligent monitoring systems use bionic electronic eyes for real-time video analysis, which can effectively identify suspicious behaviors and abnormal events and enhance public safety. In the field of robotics, bionic electronic eyes enable robots to better understand and adapt to their surroundings, improving their autonomous navigation and interaction capabilities.

[0005] For the micro-hemispherical bionic electronic eye, the conventional imaging method is to transform the sphere into a plane by projection or expansion, and then perform convolution processing. However, when the spherical image is expanded into a plane, the change in geometric shape causes stretching and compression at the edge of the image, affecting the imaging accuracy. In addition, in order to effectively convert between the sphere and the plane, a complex interpolation algorithm is required, which increases the complexity of the calculation and the processing time. In addition, the projected images generated at different positions are different, and the translation and rotation invariance of the convolutional neural network is lost, making it difficult to achieve good results.

[0006] This application aims to propose an image processing method based on a spherical convolutional neural network based on a micro-hemispherical bionic electronic eye device in order to make full use of the characteristics of the bionic eye directly imaging on the spherical surface and achieve accurate recognition over a large range. Compared with the planar convolutional neural network algorithm, the spherical convolutional neural network algorithm has fast convergence and high accuracy, and the characteristics of the three-dimensional image are not lost. Summary of the invention

[0007] In a first aspect of the present disclosure, there is provided an image processing method based on a spherical convolutional neural network, the method comprising the following steps: The input two-dimensional image data is projected onto a spherical base through the inverse North Pole projection method to form a spherical image; Mapping the spherical image into a spherical grid; Performing spherical convolution calculation according to the image in the spherical grid, specifically comprising: converting the image in the spherical grid from the spatial domain to the frequency domain space by generalized Fourier transform, using spherical harmonics and Wigner D matrix as orthogonal basis, fitting the image and convolution kernel in the spherical grid, performing inner product summation operation in the frequency domain space, and completing the spherical convolution calculation; The frequency domain convolution result is restored to the geometric space through inverse Fourier transform to obtain the output feature map.

[0008] In combination with the first aspect, the two-dimensional image data includes handwritten digital images, and is trained and tested based on the MNIST data set.

[0009] In combination with the first aspect, the spherical grid is a uniformly distributed spherical base grid, and the spherical image is fitted in blocks through the grid.

[0010] In combination with the first aspect, the spherical convolution operation utilizes spherical harmonic functions for fitting, and combines the Wigner D matrix to decompose and reconstruct image features in the frequency domain space.

[0011] In combination with the first aspect, the method is used for three-dimensional image recognition tasks, including target detection, target classification and environment perception.

[0012] A second aspect of the present disclosure provides an image processing system based on a spherical convolutional neural network, comprising: A data projection module, used for converting an input two-dimensional image into a spherical image by inverse North Pole projection; A spherical grid mapping module, used for mapping a spherical image to a spherical base grid; Convolution calculation module, including: Generalized Fourier transform unit, used to transform spherical images into frequency domain space, Convolution operation unit, which performs convolution operation in frequency domain based on spherical harmonics and Wigner D matrix. Inverse Fourier transform unit, used to restore the convolution result to the spatial domain; The output module is used to output the image feature map.

[0013] In combination with the second aspect, the system also includes a comparison module for performing performance comparison analysis on the results of the spherical convolutional neural network and the planar convolutional neural network.

[0014] A third aspect of the present disclosure provides an image processing device based on a spherical convolutional neural network, comprising: An input unit, used for receiving input two-dimensional image data; A processing unit, including a data projection module, a spherical grid mapping module and a convolution calculation module, for projecting the two-dimensional image data onto a spherical surface and performing a spherical convolution operation; The output unit is used to output the feature results after convolution.

[0015] According to a fourth aspect of the present disclosure, an electronic device is provided, including: one or more processors; A storage unit is used to store one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the image processing method based on the spherical convolutional neural network.

[0016] A fifth aspect of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, it can implement the image processing method based on the spherical convolutional neural network.

[0017] Beneficial effects: The present invention performs convolution operations directly on the sphere, avoiding the edge stretching and compression phenomenon that occurs when the spherical image is unfolded into a plane, and can more accurately maintain the geometric features of the three-dimensional image. In addition, the use of spherical harmonic functions and Wigner D matrices as orthogonal bases for frequency domain convolution calculations effectively reduces the computational complexity and improves the computational efficiency. Experimental results show that the method based on spherical convolutional neural networks not only converges faster, but also has higher accuracy in image recognition and classification tasks, and is suitable for automation systems, intelligent monitoring, drone vision, robot vision and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic diagram of a flow chart of an image processing method based on a spherical convolutional neural network according to an embodiment of the present disclosure; Figure 2 A schematic diagram of spherical projection and three-dimensional imaging of an embodiment of the present disclosure; Figure 3 A schematic diagram of the calculation process of the spherical convolutional neural network according to an embodiment of the present disclosure; Figure 4 A comparison diagram of the loss functions of the spherical convolutional neural network algorithm and the planar convolutional neural network algorithm of the embodiment of the present disclosure; Figure 5 A comparison diagram of the accuracy functions of the spherical convolutional neural network algorithm and the planar convolutional neural network algorithm of the embodiment of the present disclosure; Figure 6 An electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] Here, exemplary embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the embodiments of the present disclosure.

[0020] The terms used in the disclosed embodiments are only for the purpose of describing specific embodiments and are not intended to limit the disclosed embodiments. The singular forms of "a", "said" and "the" used in the disclosed embodiments and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0021] It should be understood that although the terms first, second, third, etc. may be used to describe various information in the disclosed embodiments, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the disclosed embodiments, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0022] like Figure 1 The figure is a flow chart of a knowledge transfer method of an evolutionary algorithm for solving a constrained multi-objective optimization problem according to an embodiment of the present disclosure, including: S101: Projecting the input two-dimensional image data onto a spherical base by an inverse North Pole projection method to form a spherical image; S102: Mapping the spherical image into a spherical grid; S103: performing spherical convolution calculation according to the image in the spherical grid, specifically comprising: converting the image in the spherical grid from the spatial domain to the frequency domain space by generalized Fourier transform, using spherical harmonic functions and Wigner D matrix as orthogonal basis, fitting the image in the spherical grid and the convolution kernel, performing inner product summation operation in the frequency domain space, and completing the spherical convolution calculation; S104: Restore the frequency domain convolution result to the geometric space through inverse Fourier transform to obtain an output feature map.

[0023] Specific: S101: Projecting a two-dimensional image onto a spherical base, refer to Figure 2 .

[0024] Input: Input 2D image data (e.g. MNIST handwritten digit data).

[0025] Method: Use the inverse North Pole projection method to convert the two-dimensional image into a spherical coordinate system: The pixel points on the two-dimensional plane are mapped to the spherical basis. This projection process corresponds the pixel points (x1, y1) of the two-dimensional image to the spherical coordinates (x0, y0, z0) one by one, and each pixel point after projection has spherical coordinates on the sphere.

[0026] Figure 2 The green grid in the middle represents the spherical base. The projected two-dimensional image forms a three-dimensional spherical image on the sphere, retaining the geometric features of the original image.

[0027] S102: Mapping the spherical image to the spherical grid is for the purpose of standardizing the projected spherical image data and placing it into the spherical grid.

[0028] Specific operations: Constructs a uniform spherical mesh (a grid of points formed by spherical coordinates).

[0029] The data points on the spherical image are matched with the spherical grid so that the spherical image data corresponds to the grid nodes one by one.

[0030] If some grid nodes are missing data, they can be supplemented by interpolation methods.

[0031] Output: Spherical image data on a spherical grid.

[0032] S103: Spherical convolution calculation, reference Figure 3 , performing efficient convolution operations on spherical grids.

[0033] (1) Conversion from spatial domain to frequency domain: The spherical image is converted from the spatial domain to the frequency domain by generalized Fourier transform (spherical Fourier transform).

[0034] Spherical harmonics and Wigner D matrix are used as orthogonal basis to fit the spherical image data.

[0035] Spherical harmonics: The basis for frequency domain representation of spherical images.

[0036] Wigner D matrix: used to handle the transformation when the sphere is rotated to ensure rotation invariance.

[0037] (2) Convolution operation: In the frequency domain, the frequency domain representation of the spherical image and the spherical convolution kernel are summed by an inner product operation.

[0038] This process is accomplished by multiplying the frequency domain data points with the coefficients of the convolution kernel on the spherical harmonic basis and summing them.

[0039] Output: Get the frequency domain representation result after convolution.

[0040] S104: Inverse Fourier transform restores the output feature map and continues to combine Figure 3 : The convolution result in the frequency domain is restored back to the spatial domain of the spherical grid through the inverse generalized Fourier transform. This restoration ensures that the result is mapped back to the original spherical geometry space.

[0041] Output: The output spherical feature map contains the feature information after the convolution operation.

[0042] The output feature map retains the geometric structure characteristics of the 3D image and maintains rotation invariance.

[0043] Figure 4-Figure 5 : is a comparison diagram of the spherical convolutional neural network algorithm and the planar convolutional neural network algorithm of the embodiment of the present disclosure, wherein Figure 4 is a comparison chart of loss functions, Figure 5 This is a comparison chart of the accuracy function.

[0044] Among them, the performance differences between the spherical convolutional neural network (spherical CNN) and the planar convolutional neural network (planar CNN) were compared through experimental results.

[0045] Figure 4 and Figure 5 The horizontal axis represents the number of training times.

[0046] Figure 4 The vertical axis refers to the loss function value, which reflects the error during the training process.

[0047] The red curve refers to planar convolution: planar convolution has higher training loss and slower convergence speed.

[0048] The blue curve refers to spherical convolution: the training loss of spherical convolution drops rapidly, converges quickly and has a lower final error.

[0049] Figure 5 The vertical axis refers to the accuracy, which reflects the performance of the model on the test data set.

[0050] The red curve refers to planar convolution: the accuracy of planar convolution is low and the performance is poor.

[0051] The blue curve refers to spherical convolution: the accuracy of spherical convolution is higher, which means that spherical convolution can better preserve three-dimensional features and improve recognition effect.

[0052] Furthermore, the present invention provides an image processing system based on a spherical convolutional neural network, which is used to execute the above-mentioned image processing method based on a spherical convolutional neural network, comprising: A data projection module, used for converting an input two-dimensional image into a spherical image by inverse North Pole projection; A spherical grid mapping module, used for mapping a spherical image to a spherical base grid; Convolution calculation module, including: Generalized Fourier transform unit, used to transform spherical images into frequency domain space, Convolution operation unit, which performs convolution operation in frequency domain based on spherical harmonics and Wigner D matrix. Inverse Fourier transform unit, used to restore the convolution result to the spatial domain; The output module is used to output the image feature map.

[0053] Optionally, the system also includes a comparison module for performing performance comparison analysis on the results of the spherical convolutional neural network and the planar convolutional neural network.

[0054] Furthermore, the present invention provides an image processing device based on a spherical convolutional neural network, which uses the above-mentioned image processing system based on a spherical convolutional neural network, including: An input unit, used for receiving input two-dimensional image data; A processing unit, including a data projection module, a spherical grid mapping module and a convolution calculation module, for projecting the two-dimensional image data onto a spherical surface and performing a spherical convolution operation; The output unit is used to output the feature results after convolution.

[0055] Beneficial effects: The present invention performs convolution operations directly on the sphere, avoiding the edge stretching and compression phenomenon that occurs when the spherical image is unfolded into a plane, and can more accurately maintain the geometric features of the three-dimensional image. In addition, the use of spherical harmonic functions and Wigner D matrices as orthogonal bases for frequency domain convolution calculations effectively reduces the computational complexity and improves the computational efficiency. Experimental results show that the method based on spherical convolutional neural networks not only converges faster, but also has higher accuracy in image recognition and classification tasks, and is suitable for automation systems, intelligent monitoring, drone vision, robot vision and other fields.

[0056] The electronic device 600 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 600 may include, but is not limited to, a processor 601 and a memory 602. Those skilled in the art will appreciate that Figure 6 It is only an example of the electronic device 600 and does not constitute a limitation of the electronic device 600. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0057] Processor 601 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0058] The memory 602 may be an internal storage unit of the electronic device 600, for example, a hard disk or memory of the electronic device 600. The memory 602 may also be an external storage device of the electronic device 600, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 600. Further, the memory 602 may also include both an internal storage unit of the electronic device 600 and an external storage device. The memory 602 is used to store the computer program 603 and other programs and data required by the electronic device. The memory 602 may also be used to temporarily store data that has been output or is to be output.

[0059] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.

[0060] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present disclosure implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, and the computer program code may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electric carrier signals and telecommunication signals.

[0061] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.

Claims

1. An image processing method based on spherical convolutional neural network, characterized in that: The method comprises the following steps: The input two-dimensional image data is projected onto a spherical base through the inverse North Pole projection method to form a spherical image; Mapping the spherical image into a spherical grid; Performing spherical convolution calculation according to the image in the spherical grid, specifically comprising: converting the image in the spherical grid from the spatial domain to the frequency domain space by generalized Fourier transform, using spherical harmonics and Wigner D matrix as orthogonal basis, fitting the image and convolution kernel in the spherical grid, performing inner product summation operation in the frequency domain space, and completing the spherical convolution calculation; The frequency domain convolution result is restored to the geometric space through inverse Fourier transform to obtain the output feature map.

2. The method according to claim 1, characterized in that The two-dimensional image data includes handwritten digital images, and is trained and tested based on the MNIST data set.

3. The method according to claim 1, characterized in that The spherical grid is a uniformly distributed spherical base grid, and the spherical image is fitted in blocks through the grid.

4. The method according to claim 1, characterized in that: The spherical convolution operation utilizes spherical harmonic functions for fitting, and combines the Wigner D matrix to decompose and reconstruct image features in the frequency domain space.

5. The method according to claim 1, characterized in that The method is used for 3D image recognition tasks, including object detection, object classification and environment perception.

6. An image processing system based on spherical convolutional neural network, characterized in that: The system is used to execute the method according to claim 1, comprising: A data projection module, used for converting an input two-dimensional image into a spherical image by inverse North Pole projection; A spherical grid mapping module, used for mapping a spherical image to a spherical base grid; Convolution calculation module, including: Generalized Fourier transform unit, used to transform spherical images into frequency domain space, Convolution operation unit, which performs convolution operation in frequency domain based on spherical harmonics and Wigner D matrix. Inverse Fourier transform unit, used to restore the convolution result to the spatial domain; The output module is used to output the image feature map.

7. The system according to claim 6, characterized in that The system also includes a comparison module for performing performance comparison analysis on the results of the spherical convolutional neural network and the planar convolutional neural network.

8. An image processing device based on spherical convolutional neural network, characterized in that: The device uses the system described in claim 6, including: An input unit, used for receiving input two-dimensional image data; A processing unit, including a data projection module, a spherical grid mapping module and a convolution calculation module, for projecting the two-dimensional image data onto a spherical surface and performing a spherical convolution operation; The output unit is used to output the feature results after convolution.

9. An electronic device, characterized in that: include: one or more processors; A storage unit, used to store one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement the image processing method based on the spherical convolutional neural network according to any one of claims 1 to 5.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it can implement the image processing method based on a spherical convolutional neural network according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Spherical convolution-based all-round fisheye image 3D sensing method and related equipment

    CN119090738A

  • Spherical image super-resolution method and system

    CN119762346A

  • A Fully Fourier Space Spherical Convolutional Neural Network Based on Clebsch-Gordan Transforms

    US20210272233A1