Multi-modal surface data processing optimization system based on spherical Transform
The multimodal surface data processing system based on spherical Transformer solves the problems of large computational load and applicability in surface data processing in existing technologies. It achieves efficient processing of arbitrary surface data and fusion of multi-source data, and is suitable for earth simulation, 3D reconstruction and camera panoramic understanding.
Patent Information
- Application Number
- CN202511495366.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Existing technologies struggle to balance local details and global structure when processing various types of surface data, resulting in high computational costs and unsuitability for surface data in non-Euclidean space.
A multimodal surface data processing system based on spherical Transformer is adopted, including an image acquisition and preprocessing module, a spherical uniform symmetric mesh surface generation module, a spherical convolution and pooling module, a spherical Transformer module, and a spherical surface sampling and downsampling module. By generating spherical uniform symmetric mesh surfaces and cross-modal attention mechanism, a spherical Transformer model is constructed to achieve efficient processing of arbitrary curved surface data.
It achieves efficient processing of arbitrary curved surface data, can extract rich geometric features, and is suitable for applications such as earth simulation, 3D reconstruction and camera panoramic understanding. It reduces computational complexity and improves the fusion efficiency of multi-source data.
Smart Images

Figure CN120997452A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and computer vision technology, and specifically relates to a multimodal surface data processing optimization system based on spherical Transformer. Background Technology
[0002] Earth sciences, 3D graphics, remote sensing imagery, and climate simulation often require processing a type of "curved surface data," such as the Earth's surface, celestial surfaces, the outer shell of industrial parts, and 360° images captured by panoramic cameras. Traditional methods typically "flatten" these curved surfaces onto a planar grid; however, this introduces distortion at edges or in polar regions, and it struggles to capture both global and local information. Graph neural networks and spherical convolutions utilize grid adjacency relationships for local feature extraction, but they often only see information within a "small area," failing to capture larger-scale or global structural changes.
[0003] The recently developed Transformer method captures long-range dependencies using a global self-attention mechanism, but it requires the data structure to satisfy the Euclidean space assumption, making it unsuitable for non-Euclidean surfaces. To transfer the Transformer method to surface data such as spheres, current research typically first subdivides the surface into surface patches and computes attention between these patches. However, directly transferring the Transformer method from Euclidean space to non-Euclidean surfaces leads to a surge in computation due to the large number of vertices on the surface; it also encounters the problem of uneven point distribution on the surface. To address these issues, some newer methods utilize algorithms such as HEALPix to divide the surface into relatively uniform surface patches; other research references the multimodal fusion mechanism of biological vision systems to dynamically allocate attention regions. However, to date, no method can simultaneously achieve both: 1. Applicable to various curved surfaces, including Earth, human body scanning, point cloud scenes, and 3D medical images; 2. Balancing local details with overall structure; 3. The computational load is reasonable, and it can handle large-scale high-resolution data. Summary of the Invention
[0004] Purpose of the invention: The technical problem to be solved by the present invention is to address the shortcomings of the existing technology by providing a multimodal surface data processing optimization system based on spherical Transformer. This system can directly interface with arbitrary curved surface data and extract rich geometric features while ensuring speed, providing unified processing capabilities for applications such as earth simulation, 3D reconstruction, and camera panoramic understanding.
[0005] The system includes an image acquisition and preprocessing module, a spherical uniform symmetric mesh generation module, a spherical convolution and pooling module, a spherical Transformer module, a spherical sampling and downsampling module, and an application analysis module; The image acquisition and preprocessing module is used to segment regions of interest and reconstruct surface curves, complete rigid registration in a unified coordinate system, and extract features at the vertices of the surface, especially extracting cortical thickness, curvature, depth, multi-source signal values, signal gradients, and superficial white matter features for brain images; The spherical uniform symmetric mesh surface generation module transforms the surface obtained by the image acquisition and preprocessing module into a unit sphere in standard space. Then, it uses the first-order and second-order subdivision icosahedron method to resample the unit sphere into a uniform and highly symmetric graphic structure. Finally, it resamples the surface features through natural neighborhood interpolation and radial basis function interpolation. The spherical convolution and pooling module is based on the uniform symmetric mesh surface obtained by the spherical uniform symmetric mesh surface generation module. It defines spherical convolution using spherical geometric relationships and establishes a spherical pooling method based on the recursive process of subdividing a regular icosahedron. The spherical Transformer module divides the spherical graphic into uniform spherical graphic blocks based on the spherical uniform symmetric mesh surface generation module, constructs a spherical Transformer model, and constructs a spherical cross-modal Transformer model based on the cross-modal attention mechanism. The spherical sampling and downsampling module is based on the recursive process of the icosahedral subdivision method in the spherical uniform symmetric mesh generation module, which realizes the splitting and merging of spherical graphic blocks, thereby achieving upsampling and downsampling. The application analysis module integrates a spherical convolution and pooling module, a spherical Transformer module, and a spherical sampling and downsampling module to establish a spherical autoencoder model, a segmentation model, and a classification model for target recognition, risk prediction, or anomaly detection tasks.
[0006] The image acquisition and preprocessing module specifically performs the following steps: segmenting the raw data to extract the region of interest, and extracting the surface curves based on the moving cube algorithm. And use the level set algorithm to Expand inwards or outwards to obtain surfaces with one-to-one correspondence between vertices. For extracting the gray matter inner surface surface corresponding to each vertex of brain data and outer surface curved surface The surface thickness is determined by the corresponding vertex spacing, and the intermediate surface is reconstructed at 25%, 50%, and 75% of the spacing, respectively. , and The raw data mentioned above are three-dimensional medical imaging data that reflect the anatomical structure and functional information of the nervous system. For the same research object, a set of neural images with curved surface morphological features is acquired using multiple imaging techniques within a set time window. This set of neural images includes multi-source image data. Six-degree-of-freedom rigid registration is performed on the multi-source image data, aligning it to a reference space by minimizing mutual information or mean square error. Features are extracted at each curved surface node, including: exist Geometric features of thickness, depth, and curvature are extracted from the surface, and the signal intensities of T1-weighted imaging (T1WI) and fluid attenuated inversion recovery (FLAIR) images based on T1 relaxation time are mapped to the vertices of the surface. exist The surface was used as a reference to extract normalized weighted imaging T1WI and liquid attenuation inversion recovery image FLAIR, respectively. For diffusion tensor imaging (DTI) of the brain, With each vertex on the surface as the center point, and with Corresponding points on the curved surface The truncated axisymmetric Gaussian distributions are established with the direction of the line connecting the vertices on the surface as the axis of symmetry. The intersection of the truncated axisymmetric Gaussian distribution region and the white matter is taken, and the weighted average fractional anisotropy FA, average diffusivity MD, axial diffusivity AD and radial diffusivity RD are calculated. For positron emission tomography (PET) images, The normalized ingestion values corresponding to each vertex of the surface sampling are compared with the SUVR and the asymmetric exponent AI.
[0007] The spherical uniform symmetrical mesh generation module specifically performs the following steps: Based on The surface is topologically shaped into a unit sphere in standard space, while maintaining the relative spacing between the vertices of the surface and the geometry of the triangular facets within the error range and not exceeding a predetermined threshold during the deformation process; By using registration between curved surfaces to deform the sphere to standard space and scaling it down to a unit sphere, we obtain the standard sphere. (The 2 here does not mean the square of the area or radius of the sphere, but the dimension of the sphere in n-dimensional space, which is a conventional symbol for the sphere.) Then, the subdivision icosahedron method is used to resample the unit sphere into a uniform and highly symmetrical graphic structure. The described method for subdividing the icosahedron is an algorithm that approximates a sphere to a regular icosahedron through recursive geometric subdivision. It can obtain uniformly distributed and symmetrical vertices on the sphere. The icosahedron is a highly symmetrical Platonic polyhedron composed of 20 equilateral triangular faces, 12 vertices, and 30 edges. All 12 vertices are located on the circumscribing surface of the icosahedron and consist of the following three sets of symmetrical points: , in It is the golden ratio. After normalization, all vertices lie on the standard sphere. ; Reconstructing a uniform symmetric network on a standard sphere using the first-order and second-order icosahedral subdivision methods: The first-order icosahedral subdivision method includes: taking the midpoints of all edges, connecting the newly added midpoints to all faces to divide the original equilateral triangle into 4 equilateral triangles, and then normalizing the coordinates of the newly added midpoints so that they fall on a standard sphere; after a single first-order subdivision, the number of faces S increases to 4*S, and the number of edges E increases to The number of points V increases to V+E.
[0008] The second-order regular icosahedral subdivision method includes: assuming the edge length is... Take the distance from the existing vertex. The point is selected, and the center point of the equilateral triangle is chosen. Each equilateral triangle face is divided into 9 smaller equilateral triangles. The coordinates of the newly added points are normalized so that the new points fall on the standard sphere. After a single second subdivision, the number of faces S increases to 9*S, and the number of edges E increases to The number of points V increases to ; Apply the first-order regular icosahedral subdivision method and the second-order regular icosahedral subdivision method more than twice, or alternately apply the first-order regular icosahedral subdivision method and the second-order regular icosahedral subdivision method more than twice, and set the last application to be the second-order regular icosahedral subdivision method. After obtaining a spherical uniform symmetric grid, the features extracted by the image acquisition and preprocessing module are interpolated using natural neighborhood interpolation or radial basis function interpolation. The natural neighborhood interpolation includes: the original set of sampling points on the sphere. Calculate the spherical Voronoi diagram, for the i-th original sampling point. Voronoi unit Including distances on the sphere The nearest vertex, where i takes values from 1 to n. , where x represents any vertex on the sphere; This represents the geodesic distance; the spherical vertex q obtained by the icosahedral subdivision method is added to the point set P, and the Voronoi diagram is recalculated, and the Voronoi elements of the spherical vertex q are identified. All satisfy The original vertex of vertex q is called the natural neighborhood of vertex q. The contribution weights of the natural neighborhood vertices to vertex q are calculated. Where Area represents the area; the eigenvalues of vertex q Original sampling points within the natural neighborhood of vertex q eigenvalues The weighted summation is used to obtain the result. The feature value Characterizing the original sampling points The local physical or physiological properties of the location, which have been determined before interpolation; The radial basis function interpolation method includes: for any vertex q of a sphere obtained by the icosahedral subdivision method, taking the original sampling points in the neighborhood of vertex q. , making ,in, The standard deviation of the Gaussian kernel function; weights are calculated based on the Gaussian kernel function. Where exp represents the natural exponential function; and the eigenvalues of vertex q are... Original sampling points within the neighborhood eigenvalues The weighted average; in, .
[0009] The spherical convolution and pooling module specifically performs the following steps: Let u be the spherical distance between adjacent vertices. Spherical convolution will transform the current vertex... of The vertex information within the neighborhood is synthesized, among which The radius of the spherical convolution kernel; and the radius of the current vertex. spacing , and Each sphere has 6 vertices, evenly distributed on 3 concentric circles; the spherical convolution kernel passes through the current vertex. Using the meridian as a reference, convolution weights are distributed clockwise, with a distance of [missing information] from the 12 vertices of the icosahedron. The vertex has 5 vertices, and the average of the 5 vertices is used as the filler. A spherical pooling method is established based on the recursive process of subdividing a regular icosahedron: The pooling method corresponding to a first-order subdivided regular icosahedron includes: if the current subdivision recursion number is l, then the central vertex and the outer hexagonal vertices are the spherical vertices after l-1 recursions, and the remaining vertices are the newly generated vertices in this recursion; spherical pooling gathers the information of the central vertex and all vertices first-order adjacent to the central vertex to the central point, and removes the newly generated vertices in this recursion; The pooling method corresponding to the second-order subdivision method includes: pooling the information of newly generated vertices adjacent to the center point and the center points of surrounding equilateral triangles to the center point.
[0010] The spherical Transformer module specifically performs the following steps: Spherical graphic block division: Taking the vertex of the penultimate icosahedral subdivision as the center point, neighboring vertices of the second-order subdivision are taken to form graphic blocks; in addition to the 12 vertices of the icosahedron, each center point has 12 neighboring vertices, of which 6 adjacent vertices located at the center of the equilateral triangle are shared with surrounding graphic blocks to capture the contextual information between adjacent blocks; each vertex of the icosahedron has only 10 neighboring vertices, including 5 edge center points and 5 equilateral triangle face center points; the mean values of the edge and equilateral triangle face center points are respectively used to fill the vector; after flattening the spherical graphic blocks, a linear projection is performed to map each graphic block into a feature matrix. : , in Represents the real number field, where d is the vertex feature dimension, and each spherical graphic block has a total of 13 vertices.
[0011] The spherical graphic block is positionally encoded using spherical Fourier position: , in These are the zenith angle and azimuth angle of the vertex in spherical polar coordinates, respectively. PE represents the learnable parameter matrix; PE represents the positional encoding; k represents the index variable for summation; Tensor products can combine different vectors, matrices or other mathematical objects according to specific rules to generate more complex mathematical objects, so as to express richer information or perform specific operations. Establish a spherical Transformer model: For the feature vector h of the spherical graphic block, apply a sub-attention mechanism: , Here, Attention represents the attention mechanism, Q, K, and V are the query vector, key vector, and value vector, respectively, which are obtained by multiplying the feature vector h with the query matrix, key matrix, and value matrix, respectively; T represents matrix transfer; the spherical Transformer model multiplies the attention with the feature vector h, performs layer standardization, inputs it into the feedforward layer for synthesis, and then performs layer standardization again; Establish a spherical cross-modal Transformer model: The spherical cross-modal Transformer model includes a cross-modal attention module with D layers. The multi-head attention module is suitable for multi-source data. eigenvectors and multi-source data eigenvectors (in and The difference is: Provide source data features to generate key and value vectors to capture the inherent patterns in the data; (Provides target data features to generate query vectors to guide attention allocation) Cross-modal attention module Will Information input : , in query vector From the eigenvector get, key vector Sum value vector Depend on get, Let be the dimension of the key vector; , and These are the parameters to be trained; For a cross-modal attention module with D layers The multi-head attention module uses cross-layer connections in each layer to... With cross-modal attention module The output is first summed across layers, then normalized, and then input into the feedforward layer for nonlinear deformation, followed by a second summation and normalization across layers. The summation across layers refers to combining the input feature vector of the current layer with the summation of the previous layer's input feature vector. and The output feature vectors processed by the cross-modal attention module are added element-wise to preserve the original input information during feature fusion while introducing new processed features to enhance the model's expressive power and training stability.
[0012] The spherical sampling and downsampling module includes a downsampling layer and an upsampling layer; The downsampling layer uses a recursive process based on the subdivision of the icosahedron to merge spherical graphic blocks. The center point of the initial graphic block is the vertex obtained by the last first-order subdivision method. The mean of the feature vectors of all vertices in the graphic block is taken and converged at the center point of the graphic block. A weighted average is taken for all vertices in the block, and the weights are determined by the area ratio of the spherical Voronoi unit. , in, The area of the Voronoi region at vertex i is calculated recursively through subdivision levels; This represents the sum of the areas of all Voronoi regions within the graphic block; This represents the eigenvector of the i-th vertex; This represents the output feature vector obtained after weighted pooling, which serves as the aggregated representation of the image patch; N represents the total number of vertices. Using the vertices of the previous level of subdivision as the convergence target, the feature mean of vertex i and its adjacent newly sampled vertices at this level is taken, and the downsampling process is used to simulate the subdivision recursive steps in reverse.
[0013] The upsampling layer reconstructs vertices using a recursive process based on the icosahedral subdivision method. Starting from the vertex of the last icosahedral subdivision, it backtracks upwards layer by layer to the initial icosahedral structure. The parent vertex of each level is determined through the subdivision tree topology. Linear interpolation is used to restore low-resolution features to a high-resolution mesh. Perform first-order feature interpolation: For each side of an equilateral triangle, the two endpoints correspond to two original vertices A and B, respectively, and the feature vectors of the two original vertices A and B are as follows: and ,right and Take the average value to generate edge midpoint features. : , The first-order feature interpolation corresponds to the midpoint projection of the set of first-order subdivisions of a regular icosahedron; the set of subdivisions of the regular icosahedron refers to the set of two or more levels of mesh structures generated by recursively subdividing the initial regular icosahedron, with each level corresponding to the vertex and face topology formed by one subdivision operation. Perform second-order feature expansion: Insert two new vertices at each trisection point of the edge, and calculate the features of the two new vertices using linear interpolation. and : , , Then, the calculated values on each edge and That is, the characteristic mean of the six edge vertices is used as the characteristic of the center point of the surface.
[0014] The application analysis module is used in the following application scenarios: Spherical autoencoder extracts depth features: The image acquisition and preprocessing module is used to establish a surface and extract surface features. A spherical mesh is established through the spherical uniform symmetric mesh generation module. Then, an encoder is established using the spherical Transformer module and the downsampling layer. A lightweight decoder is established using spherical convolution, spherical pooling and upsampling layers. The mean square error is used as the loss function. The output of the encoder is the depth features of the data. A segmentation model integrating multi-source data: The image acquisition and preprocessing module is used to establish a surface and extract surface features. A spherical mesh is established through a spherical uniform symmetrical mesh generation module. A spherical encoder is established using a spherical convolution and pooling module. The depth spherical features of the multi-source data are extracted respectively. The features are fused using a spherical Transformer module. After stitching, a U-Net network is established using a spherical convolution and pooling module and a spherical surface sampling and downsampling module to achieve data segmentation. A classification model that integrates multi-source data: The image acquisition and preprocessing module is used to establish a surface and extract surface features. A spherical mesh is established through a spherical uniform symmetrical mesh generation module. A spherical encoder is established using a spherical convolution and pooling module. The depth spherical features of the multi-source data are extracted respectively. The features are fused using a spherical Transformer module. After stitching, a ResNet network is established using a spherical convolution and pooling module to achieve data classification.
[0015] The present invention also provides an electronic device, including a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to execute the system.
[0016] The present invention also provides a storage medium storing a computer program or instructions that execute the system when the computer program or instructions are run on a computer.
[0017] The present invention has the following beneficial effects: 1. It constructs a mapping process from raw data to a spherically uniform symmetric mesh, uses a regular icosahedral mesh to ensure that each vertex is evenly distributed, effectively preserves the surface topology, and reduces the computational complexity of subsequent data processing by utilizing the geometric consistency of the spherical surface.
[0018] 2. The attention mechanism of Transformer can model long-range dependencies and capture complex patterns that are difficult to extract by traditional CNNs. The cross-modal attention mechanism improves the complementarity of multi-source features and achieves efficient multi-source data fusion. Attached Figure Description
[0019] Figure 1 This is a structural block diagram of a multimodal surface data processing optimization system based on a spherical Transformer according to an embodiment of the present invention.
[0020] Figure 2 This is a schematic diagram of a spherical icosahedral subdivision method according to an embodiment of the present invention, including first-order and second-order subdivision methods.
[0021] Figure 3 This is a schematic diagram of a spherical convolution and spherical pooling method according to an embodiment of the present invention.
[0022] Figure 4 This is a structural diagram of a spherical transmodal Transformer according to an embodiment of the present invention.
[0023] Figure 5 This is a network diagram of a spherical autoencoder for extracting surface depth features according to an embodiment of the present invention.
[0024] Figure 6 This is a schematic diagram of a spherical segmentation model that integrates multi-source data according to an embodiment of the present invention. Detailed Implementation
[0025] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.
[0026] like Figure 1 As shown, this embodiment of the invention provides a multimodal surface data processing optimization system based on a spherical Transformer, comprising the following modules: Image acquisition and preprocessing module: used for segmenting regions of interest and reconstructing surface curves, completing rigid registration in a unified coordinate system, and extracting features at surface vertices, especially for brain images to extract cortical thickness, curvature, depth, multi-source signal values, signal gradients, superficial white matter features, etc. The specific steps are as follows: the original data is segmented to extract the region of interest, and the surface is extracted based on the moving cube algorithm. And use the level set algorithm to Expand inwards or outwards to obtain surfaces with one-to-one correspondence between vertices. For extracting the gray matter inner surface surface corresponding to each vertex of brain data and outer surface curved surface The surface thickness is determined by the corresponding vertex spacing, and the intermediate surface is reconstructed at 25%, 50%, and 75% of the spacing, respectively. , and The raw data mentioned above are three-dimensional medical imaging data that reflect the anatomical structure and functional information of the nervous system. For the same research subject (referring to a specific subject or experimental sample), within a set time window (e.g., the same or similar time windows) (the specific time range is determined according to research needs, for example, selecting data obtained within a time interval of 48 hours; a uniform standard can be specified within a study), a collection of neural images with curved morphological features is obtained using various imaging techniques (such as magnetic resonance imaging (MRI), positron emission tomography (PET), etc.; this invention does not limit the specific imaging techniques used, and researchers can choose the appropriate imaging method according to the characteristics of the required data; these imaging techniques are technical means in the data preparation stage and are not the core content of this invention). The neural image collection includes multi-source image data, which describe neural tissue characteristics from different imaging modalities (e.g., structure, function, or metabolism). Six-degree-of-freedom rigid registration is performed on the multi-source image data, aligning it to a reference space with mutual information or mean square error minimization. Features are extracted at each curved node, including: exist Geometric features of thickness, depth, and curvature are extracted from the surface, and the signal intensities of T1-weighted imaging (T1WI) and FLAIR (fluid attenuation inversion recovery image) based on T1 relaxation time (i.e. the speed of magnetic resonance signal recovery) are mapped to the vertices of the surface. exist The surface was used as a reference to extract normalized weighted imaging T1WI and liquid attenuation inversion recovery image FLAIR, respectively. For diffusion tensor imaging (DTI) images of the brain, with With each vertex on the surface as the center point, and with Corresponding points on the curved surface The truncated axisymmetric Gaussian distributions are established with the direction of the line connecting the vertices on the surface as the axis of symmetry. The intersection of the truncated axisymmetric Gaussian distribution region and the white matter is taken, and the weighted average fractional anisotropy FA, average diffusivity MD, axial diffusivity AD and radial diffusivity RD are calculated. For positron emission tomography (PET) images, in The normalized ingestion values corresponding to each vertex of the surface sampling are compared with the SUVR and the asymmetric exponent AI.
[0027] Module for generating uniform symmetrical spherical mesh surfaces: based on the image acquisition and preprocessing module. A surface is topologically shaped into a unit sphere in standard space, while maintaining the relative spacing between the vertices and the geometry of the triangular facets within a predetermined threshold (e.g., the distance between vertices does not change by more than 5%) during the deformation process. Then, by using inter-surface registration, the spherical surface is deformed to standard space and scaled to a unit sphere, which is the standard sphere. Then, the subdivision icosahedron method is used to resample the unit sphere into a uniform and highly symmetrical graphic structure.
[0028] The subdivision method for the icosahedron is an algorithm that approximates a sphere to a regular icosahedron through recursive geometric subdivision, obtaining uniformly distributed and symmetrical vertices on the sphere. The regular icosahedron is a highly symmetrical Platonic polyhedron composed of 20 equilateral triangular faces, 12 vertices, and 30 edges. All 12 vertices lie on its circumscribing sphere and are composed of the following three sets of symmetrical points: , in After normalization, all vertices lie on the standard sphere. .like Figure 2 As shown, a uniform symmetric network is reconstructed on a standard sphere using first-order and second-order icosahedral subdivision methods: First-order icosahedral subdivision method: Taking the midpoints of all edges, and connecting the newly added midpoints to all faces, the original equilateral triangle is divided into 4 equilateral triangles. Then, the coordinates of the newly added midpoints are normalized so that they lie on a standard sphere. After a single first-order subdivision, the number of faces S increases to 4*S, and the number of edges E increases to... The number of points V increases to V+E.
[0029] Second-order icosahedral subdivision method: Let the edge length be... Take the distance from the existing vertex. The point is chosen as the center point of the equilateral triangle, and each equilateral triangle face is divided into 9 smaller equilateral triangles. The coordinates of the newly added points are normalized so that they lie on a standard sphere. After a single second subdivision, the number of faces S increases to 9*S, and the number of edges E increases to [missing information]. The number of points V increases to For faster calculation, when At that time, take .
[0030] The first-order subdivision method, second-order subdivision method, or alternating methods are applied multiple times, with the second-order icosahedral subdivision method being mandatory as the final application. After obtaining a spherically uniform and symmetric mesh, natural neighborhood interpolation or radial basis function interpolation is applied to interpolate the features extracted by the image acquisition and preprocessing module.
[0031] Natural neighborhood interpolation: for the original set of sampled points on the sphere Calculate the spherical Voronoi diagram for each original vertex. The Voronoi unit contains distances on a sphere. The most recent peak, , This represents the geodesic distance. The spherical vertex q obtained by the icosahedral subdivision method is added to the point set, the Voronoi diagram is recalculated, and the Voronoi elements of point q are identified. All satisfied The original vertex of point q is called the natural neighborhood of point q, and the Sibson weights are calculated. Area represents the area. The eigenvalues of point q. From its natural neighborhood vertices eigenvalues The weighted summation is used to obtain the result. .
[0032] Radial basis function interpolation: For any vertex q of a sphere obtained by the icosahedral subdivision method, take the original sampling points in its neighborhood. , making ,in Indicates geodesic distance. Let be the standard deviation of the Gaussian kernel function. Calculate the weights based on the Gaussian kernel function. The eigenvalues of point q are the vertices in its neighborhood. eigenvalues The weighted average, .
[0033] Spherical convolution and pooling module: Let u be the spherical distance between adjacent vertices. Spherical convolution will transform the current vertex... of The vertex information within the neighborhood is synthesized, among which Let be the radius of the spherical convolution kernel. For example... Figure 3 As shown in the left figure, the distance from the current vertex p , Both 2u and 2u have 6 vertices, evenly distributed on 3 concentric circles. The purple vertex represents the current vertex, and the red, blue, and green colors represent radii. , and The convolution kernel adds a new vertex. The spherical convolution kernel passes through the current vertex. Using the meridian as a reference, convolution weights are distributed clockwise. Furthermore, due to the Gaussian curvature of the spherical surface, the distance between the convolution weights and the 12 vertices of the icosahedron is... There are only 5 vertices, and their average value is used as the filler.
[0034] The spherical pooling method is based on the recursive process of subdividing a regular icosahedron. The pooling method corresponding to a first-order subdivided regular icosahedron is as follows: Figure 3As shown in the middle diagram, if the current recursion count of the subdivision method is l, then the central red vertex and the outer hexagonal blue vertex are the spherical vertices after l-1 recursions, and the remaining vertices are the newly generated vertices in this recursion. Spherical pooling gathers the information of the vertices within the red circle (the inner hexagonal green vertex and the red center point) to the center point and removes the newly generated vertices in this recursion. The pooling method corresponding to the second-order subdivision method is as follows: Figure 3 As shown in the right figure, the information of newly generated vertices adjacent to the center point and the center points of surrounding equilateral triangles (green vertices) are converged to the center point.
[0035] Spherical Transformer module: Spherical graphic block partitioning: Using the penultimate icosahedral subdivision vertex as the center point, neighboring vertices of the second-order subdivision are combined to form graphic blocks. In addition to the 12 vertices of the icosahedron, each center point has 12 neighboring vertices, of which 6 adjacent vertices located at the center of the equilateral triangle are shared with surrounding graphic blocks to capture contextual information between adjacent blocks. Due to the Gaussian curvature on the sphere, a tessellation similar to a two-dimensional hexagon cannot be formed; each icosahedral vertex has only 10 neighboring vertices, including 5 edge center points and 5 equilateral triangle face center points. To ensure consistent vector lengths after flattening the spherical graphic blocks, the average values of the edge and equilateral triangle face center points are used to fill the vectors. After flattening the spherical graphic blocks, a linear projection is performed, mapping each graphic block to a feature matrix. : , in Represents the real number field, where d is the vertex feature dimension, and each spherical graphic block has a total of 13 vertices.
[0036] Position encoding: The spherical graphic block is positionally encoded using spherical Fourier position: , in These are the zenith angle and azimuth angle of the vertex in spherical polar coordinates, respectively. This is the learnable parameter matrix.
[0037] Spherical Transformer Model: For the feature vector h of a spherical graphic block, a sub-attention mechanism is applied: , Where Q, K, and V are the query vector, key vector, and value vector, respectively, obtained by multiplying the feature vector h with the query matrix, key matrix, and value matrix. The spherical Transformer model multiplies the attention with the feature vector h, performs layer normalization, inputs it into the feedforward layer for synthesis, and then performs layer normalization again.
[0038] Spherical cross-modal Transformer model: such as Figure 4 As shown, for representing multi-source data and eigenvectors and Cross-modal attention module Will Information input : , The query vector From the eigenvector Obtain the key vector Sum value vector Depend on get, Let be the dimension of the key vector. , and These are the parameters to be trained. Cross-modal Transformer, such as... Figure 4 As shown in the right figure, including layer The multi-head attention module uses cross-layer connections in each layer to... and After summing the output, layer standardization is performed. Then, the input to the feedforward layer undergoes nonlinear deformation, followed by a second cross-layer summation and layer standardization.
[0039] Spherical sampling and downsampling module: Downsampling layer: The downsampling layer merges spherical graphic blocks based on the recursive process of the subdivision icosahedron method. The center point of the initial graphic block is the vertex obtained by the last first-order subdivision method. The mean of the feature vectors of all vertices in the graphic block is taken and converged at the center point of the graphic block. A weighted average is taken for all vertices in the block, and the weights are determined by the area ratio of the spherical Voronoi unit. , in Let i be the area of the Voronoi region of vertex i, which is calculated recursively through subdivision levels. Taking the vertices of the previous level of subdivision as the convergence target, the feature mean of the vertices of the previous level and the adjacent newly sampled vertices of the same level is taken, and the downsampling process simulates the subdivision recursive steps in reverse.
[0040] Upsampling layer: The upsampling layer reconstructs vertices based on the recursive process of the icosahedral subdivision method. Starting from the vertex of the last icosahedral subdivision, it backtracks upwards layer by layer to the initial icosahedral structure. The parent vertex of each level is determined by the subdivision tree topology. Low-resolution features are then restored to a high-resolution mesh using linear interpolation. First-order feature interpolation: The mean value of the two original fixed-point features of each equilateral triangle side is used to generate the side midpoint feature. , This operation corresponds to the projection of the midpoint of the set of first-order subdivisions of a regular icosahedron.
[0041] Second-order feature extension: Insert two new vertices at the trisection points of each edge, and calculate the features using linear interpolation. , , Subsequently, the mean of the features of the six edge vertices is used as the feature of the face center point.
[0042] Application Analysis Module: Multiple application methods are available to address different task requirements. Below are three application scenario examples: Spherical autoencoders extract depth features: such as Figure 5 As shown, the image acquisition and preprocessing module is used to establish a surface and extract surface features, and a spherical mesh is established using a spherical uniform symmetrical mesh generation module. Then, an encoder is built using the Transformer module and a downsampling layer, and a lightweight decoder is built using spherical convolution, spherical pooling, and an upsampling layer, using mean squared error as the loss function. The encoder output is the depth feature of the data. Figure 5 In this flowchart, N is a natural number representing the number of sub-regions (patches) the spherical mesh is divided into before being input into the Transformer model for processing. The specific value of N depends on the task, i.e., the number of vertices into which the sphere is subdivided. In this flowchart, N is presented as an unknown quantity to represent the generality of the method, intended to illustrate the processing flow of the invention rather than a fixed value.
[0043] Segmentation models that integrate multi-source data: such as Figure 6 As shown, the image acquisition and preprocessing module is used to establish a surface and extract surface features, and a spherical mesh is established through the spherical uniform symmetrical mesh generation module. A spherical encoder is established using the spherical convolution and pooling module to extract depth spherical features from multi-source data. Then, the features are fused using the spherical Transformer module, and after stitching, a U-Net network is established using the spherical convolution and pooling module, upsampling module, etc., to achieve data segmentation.
[0044] A classification model integrating multi-source data: The image acquisition and preprocessing module is used to establish a surface and extract surface features, and a spherical mesh is established using a spherical uniform symmetrical mesh generation module. A spherical encoder is built using a spherical convolution and pooling module to extract depth spherical features from the multi-source data. Then, a spherical Transformer module is used to fuse the features, and after concatenation, a ResNet network is built using spherical convolution and pooling modules to achieve data classification.
[0045] In this embodiment, the method is applied to the classification task of brain imaging data. Each data sample includes T1-weighted magnetic resonance imaging (T1WI), liquid attenuated inversion recovery sequence (FLAIR), diffusion tensor imaging (DTI), functional magnetic resonance imaging (fMRI), and positron emission tomography (PET). The method uses a spherical uniform symmetric mesh generation module, a spherical convolution and pooling module, a spherical Transformer module, and spherical sampling and downsampling modules to construct a spherical Transformer network to achieve the brain imaging data classification task.
[0046] Step 1, Image Acquisition and Preprocessing; The raw 3D brain image data underwent preprocessing, including denoising, eddy current correction, brain region extraction, and gray-white matter segmentation. Rigid registration was then performed on the multimodal image data, deforming it to T1WI images. The inner surface of the gray matter (i.e., the gray-white matter interface) was used as the surface. The outer surface of the gray matter is taken as a curved surface. and on the curved surface and Reconstruct the intermediate surface at 25%, 50%, and 75% of the vertex spacing. , and The thickness is defined as the Euclidean distance between corresponding vertices, with an average thickness ranging from 2 mm to 4 mm.
[0047] Feature extraction at each surface node: Geometric features were extracted from the surface, including thickness (range 0 mm to 5 mm), depth (range 0 mm to 20 mm), and curvature (range -0.5 to 0.5); T1WI and FLAIR signal intensities were mapped to the surface vertices and normalized to the range of 0 to 1; to Between all curved surfaces, with The average signal value of the surface is used as a reference to extract normalized T1WI and FLAIR values; in Normalized T1WI and FLAIR signal gradients were extracted from the curved surface; fMRI was used to extract these gradients. Local heterogeneity (ReHo) and fractional low-frequency fluctuation amplitude (fALFF) are extracted from the surface, and the mean is taken between corresponding vertices; for PET images, the average uptake rate (SUVR) is calculated with the cerebellum or pons average uptake value as a reference. SUVR (range 0 to 2) is extracted from the surface, and the mean value is taken among the corresponding vertices. The asymmetry index AI is calculated as: AI = 2 × |left SUVR - right SUVR| / (left SUVR + right SUVR), range 0 to 1. Left SUVR represents the average uptake rate of the left side of the human brain, and right SUVR represents the average uptake rate of the right side of the human brain. For DTI images, using... A truncated axisymmetric Gaussian distribution (standard deviation σ between 1 and 3 mm) is constructed with the vertex of the surface as the center and the corresponding connecting line as the axis of symmetry. The weighted average anisotropy fraction (FA, range 0 to 1), average diffusivity (MD, range 0 to 0.002 mm² / s), axial diffusivity (AD, range 0 to 0.002 mm² / s), and radial diffusivity (RD, range 0 to 0.002 mm² / s) are calculated by taking the intersection with the white matter.
[0048] The input is the raw multimodal image, for example, the T1WI volume of a sample is 256×256×176 voxels. The output is a set of surfaces for each sample. to The dataset contains approximately 150,000 vertices, each associated with a 23-dimensional feature vector (3D geometric features + 6D T1WI + 6D FLAIR + 2D fMRI + 2D PET + 4D DTI). The registration error, calculated using root mean square error (RMSE), is less than 0.5 mm.
[0049] Step 2, surface deformation and resampling; Arbitrary curved surfaces are mapped to a standard unit sphere through topological deformation and registration, with the number of iterations between 150 and 250, while maintaining the error of relative vertex spacing and triangular face shape to less than 5%. Using a regular icosahedron as a seed, a recursive first-order subdivision method is used to generate a uniform and highly symmetric mesh (level=4, generating 2,562 vertices).
[0050] Interpolate the pre-extracted vertex features: use natural neighborhood interpolation, find the natural neighborhood of the new point q (number of neighborhoods k=6) through the spherical Voronoi diagram, and sum the neighborhood features according to the Sibson weight.
[0051] The input consists of the surface and features from step 1. The output is a uniform spherical mesh (symmetry error less than 3%), interpolated into a continuous 23-dimensional feature field. The mesh coverage reaches 99%, and the interpolation error is less than 0.03 calculated using the L2 norm. For example, the thickness feature variation (variance) of a sample with Alzheimer's disease (AD) in the spherical frontal lobe region (the front of the brain, responsible for thinking) is approximately 0.2 (compared to approximately 0.1 in normal individuals, suggesting a possible abnormality). The continuous field formed by the fractional anisotropy (DTI FA) measured using diffusion tensor imaging (like a map of white matter pathways) shows interruptions in white matter trajectories (neural signal transmission pathways), which may be a sign of white matter damage.
[0052] Step 3, spherical convolution and pooling; Vertex information is comprehensively processed using spherical convolution: The spherical distance between adjacent vertices is set to u (the average vertex spacing is approximately 0.02). Within a neighborhood of the current vertex p with a radius of u, information is aggregated from vertices (6 vertices evenly distributed on concentric circles) that are u apart from p. For example... Figure 3 As shown in the left figure, with the current vertex spacing , and Each vertex has 6 vertices, evenly distributed on 3 concentric circles. The purple vertex represents the current vertex, and the red, blue, and green vertices represent radii of 6. , and The convolution kernel adds new vertices. The convolution kernel distributes weights clockwise based on the meridian passing through p, and fills the 12 vertices of the regular icosahedron with the average of the 5 adjacent points with a spacing of u.
[0053] Spherical pooling is implemented based on a recursive process of subdividing a regular icosahedron: For example... Figure 3 As shown in the middle diagram, first-order subdivision pooling gathers all adjacency information of the center vertex and its outer hexagonal vertices (belonging to level l-1) of the current recursive level l to the center point and removes newly generated vertices. The model uses two convolutional layers (the number of channels increases from 16 to 32, and the activation function is ReLU) and one pooling layer. Furthermore, the pooling method corresponding to second-order subdivision is as follows... Figure 3 As shown in the right figure, the information of newly generated vertices adjacent to the center point and the center points of surrounding equilateral triangles (green vertices) are converged to the center point.
[0054] The input is the spherical features (2,562 vertices × 8 dimensions) obtained from convolution. The output is a compressed feature map (vertices reduced to 642, dimension 32). Feature aggregation after convolution enhances local contrast; for example, the peak curvature of AD samples is enhanced by approximately 10%–15%. Pooling achieves a compression ratio of 70%–75%, preserving key geometric information such as thickness gradients. For normal control NC samples, the standard deviation of the output feature map is approximately 0.1, while AD samples show anomalous clustering.
[0055] Step 4: Fuse multimodal features using a spherical cross-modal Transformer; Spherical graphic block division: Taking the penultimate icosahedral subdivision vertex as the center, the graphic blocks are formed by taking the nearest second-order subdivision vertices. For example... Figure 4 As shown in the left figure, each center point has 12 neighboring points (6 of which are shared with neighboring blocks), and each vertex of the icosahedron has 10 neighboring points. The mean fill ensures that the vector length is consistent.
[0056] Cross-modal Transformer models: such as Figure 4As shown in the right figure, T1WI and DTI features are fused. The cross-modal attention module inputs DTI information into the T1WI features and contains four layers of multi-head attention, each of which undergoes layer normalization and feedforward processing. The input is the compressed features from step 3, and the output is the initially fused 32-dimensional feature matrix.
[0057] Step 5: Sampling and downsampling on the sphere; Downsampling layer: Based on recursive subdivision of a regular icosahedron, it merges graphic blocks and takes the average vertex feature value to converge at the midpoint. From level=4 to level=2, the number of vertices decreases from 642 to 162. For example... Figure 5 As shown, downsampling is used for encoder feature compression.
[0058] Upsampling layer: Backtracking from level=2 to level=4, linear interpolation recovers features, and the midpoint of an edge is calculated by averaging the values at both ends. For example... Figure 5 As shown, upsampling is used for feature reconstruction by the decoder. The input is the preliminary fused features from step 4, and the output is a resolution-adjusted feature field with 642 vertices.
[0059] Step 6, Application Analysis (Task Classification); Using the aforementioned spherical uniform symmetric grid as a unified representation framework, a multimodal analysis system is constructed: a spherical convolution-pooling module extracts features, a Transformer module fuses T1WI and DTI data, and a sampling module adjusts the resolution; the classification head consists of global average pooling followed by a fully connected layer, outputting two class labels using the softmax activation function. Training is performed using the PyTorch framework, the Adam optimizer (learning rate 0.001), the cross-entropy loss function, a batch size of 32, and 50 training epochs. The dataset is split as follows: 80% for training (160 samples) and 20% for testing (40 samples). Figure 6 As shown, ResNet is constructed using spherical convolution and pooling modules after feature fusion.
[0060] The input is multimodal spherical features. The output is a classification label (AD or NC). The training loss decreases from an initial 1.2 to 0.3-0.4 (converging in rounds 30-40). For example, the output probability of a test AD sample is 0.80-0.85 for AD and 0.15-0.20 for NC; the output probability of an NC sample is 0.10-0.15 for AD and 0.85-0.90 for NC. The overall accuracy on the test set is between 80% and 85%, precision is between 79% and 84%, recall is between 81% and 86%, and F1 score is between 0.80 and 0.85. The confusion matrix shows that AD was correctly classified 15-17 / 20 times, and NC was correctly classified 16-18 / 20 times.
[0061] The proposed method addresses the issues of topological distortion and multimodal alignment in brain image processing, improving classification accuracy while maintaining rotational equivariance and geometric consistency. In practical applications, this method can be used for complex data analysis and filtering.
[0062] This invention provides a multimodal surface data processing optimization system based on a spherical Transformer. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment of the invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.
Claims
1. A multimodal surface data processing and optimization system based on a spherical Transformer, characterized in that, It includes an image acquisition and preprocessing module, a spherical uniform symmetric mesh generation module, a spherical convolution and pooling module, a spherical Transformer module, a spherical sampling and downsampling module, and an application analysis module; The image acquisition and preprocessing module is used to segment the region of interest and reconstruct the surface, complete rigid registration in a unified coordinate system, and extract features at the vertices of the surface. The spherical uniform symmetric mesh surface generation module transforms the surface obtained by the image acquisition and preprocessing module into a unit sphere in standard space. Then, it uses the first-order and second-order subdivision icosahedron method to resample the unit sphere into a uniform and highly symmetric graphic structure. Finally, it resamples the surface features through natural neighborhood interpolation and radial basis function interpolation. The spherical convolution and pooling module is based on the uniform symmetric mesh surface obtained by the spherical uniform symmetric mesh surface generation module. It defines spherical convolution using spherical geometric relationships and establishes a spherical pooling method based on the recursive process of subdividing a regular icosahedron. The spherical Transformer module divides the spherical graphic into uniform spherical graphic blocks based on the spherical uniform symmetric mesh surface generation module, constructs a spherical Transformer model, and constructs a spherical cross-modal Transformer model based on the cross-modal attention mechanism. The spherical sampling and downsampling module is based on the recursive process of the icosahedral subdivision method in the spherical uniform symmetric mesh generation module, which realizes the splitting and merging of spherical graphic blocks, thereby achieving upsampling and downsampling. The application analysis module integrates a spherical convolution and pooling module, a spherical Transformer module, and a spherical sampling and downsampling module to establish a spherical autoencoder model, a segmentation model, and a classification model for target recognition, risk prediction, or anomaly detection tasks.
2. The system according to claim 1, characterized in that, The image acquisition and preprocessing module specifically performs the following steps: segmenting the raw data to extract the region of interest, and extracting the surface curves based on the moving cube algorithm. And use the level set algorithm to Expand inwards or outwards to obtain surfaces with one-to-one correspondence between vertices. For extracting the gray matter inner surface surface corresponding to each vertex of brain data and outer surface curved surface The surface thickness is determined by the corresponding vertex spacing, and the intermediate surface is reconstructed at 25%, 50%, and 75% of the spacing, respectively. , and The raw data mentioned above are three-dimensional medical imaging data that reflect the anatomical structure and functional information of the nervous system. For the same research object, a set of neural images with curved surface morphological features is acquired using imaging techniques within a set time window. This set of neural images includes multi-source image data. Six-degree-of-freedom rigid registration is performed on the multi-source image data, aligning it to a reference space by minimizing mutual information or mean square error. Features are extracted at each curved surface node, including: exist Geometric features of thickness, depth, and curvature are extracted from the surface, and the signal intensities of T1-weighted imaging (T1WI) and FLAIR (fluid attenuation inversion recovery image) based on T1 relaxation time are mapped to the vertices of the surface. exist The surface was used as a reference to extract normalized weighted imaging T1WI and liquid attenuation inversion recovery image FLAIR, respectively. For diffusion tensor imaging (DTI) images of the brain, with With each vertex on the surface as the center point, and with Corresponding points on the curved surface The truncated axisymmetric Gaussian distributions are established with the direction of the line connecting the vertices on the surface as the axis of symmetry. The intersection of the truncated axisymmetric Gaussian distribution region and the white matter is taken, and the weighted average fractional anisotropy FA, average diffusivity MD, axial diffusivity AD and radial diffusivity RD are calculated. For positron emission tomography (PET) images, in The normalized ingestion values corresponding to each vertex of the surface sampling are compared with the SUVR and the asymmetric exponent AI.
3. The system according to claim 2, characterized in that, The spherical uniform symmetrical mesh generation module specifically performs the following steps: Based on The surface is topologically shaped into a unit sphere in standard space, while maintaining the relative spacing between the vertices of the surface and the geometry of the triangular facets within the error range and not exceeding a predetermined threshold during the deformation process; By using registration between curved surfaces to deform the sphere to standard space and scaling it down to a unit sphere, we obtain the standard sphere. Then, the subdivision icosahedron method is used to resample the unit sphere into a uniform and highly symmetrical graphic structure; The described method for subdividing a regular icosahedron is an algorithm that approximates a regular icosahedron to a sphere through recursive geometric subdivision. It can obtain uniformly distributed and symmetrical vertices on the sphere. The regular icosahedron is a highly symmetrical Platonic polyhedron composed of 20 equilateral triangular faces, 12 vertices, and 30 edges. All 12 vertices are located on the circumscribing surface of the icosahedron, and are composed of the following three sets of symmetrical points. composition: , in It is the golden ratio. ; After normalization, all vertices lie on the standard sphere. ; Reconstructing a uniform symmetric network on a standard sphere using the first-order and second-order icosahedral subdivision methods: The first-order icosahedral subdivision method includes: taking the midpoints of all edges, connecting the newly added midpoints to all faces to divide the original equilateral triangle into 4 equilateral triangles, and then normalizing the coordinates of the newly added midpoints so that they fall on a standard sphere; after a single first-order subdivision, the number of faces S increases to 4*S, and the number of edges E increases to The number of points V increases to V+E; The second-order regular icosahedral subdivision method includes: assuming the edge length is... Take the distance from the existing vertex. The point is selected, and the center point of the equilateral triangle is chosen. Each equilateral triangle face is divided into 9 smaller equilateral triangles. The coordinates of the newly added points are normalized so that the new points fall on the standard sphere. After a single second subdivision, the number of faces S increases to 9*S, and the number of edges E increases to The number of points V increases to ; Apply the first-order regular icosahedral subdivision method and the second-order regular icosahedral subdivision method more than twice, or alternately apply the first-order regular icosahedral subdivision method and the second-order regular icosahedral subdivision method more than twice, and set the last application to be the second-order regular icosahedral subdivision method. After obtaining a spherical uniform symmetric grid, the features extracted by the image acquisition and preprocessing module are interpolated using natural neighborhood interpolation or radial basis function interpolation. The natural neighborhood interpolation includes: the original set of sampling points on the sphere. Calculate the spherical Voronoi diagram, for the i-th original sample point. Voronoi unit Including distances on the sphere The nearest vertex, where i takes values from 1 to n. , where x represents any vertex on the sphere; This represents the geodesic distance; the spherical vertex q obtained by the icosahedral subdivision method is added to the point set P, and the Voronoi diagram is recalculated, and the Voronoi elements of the spherical vertex q are identified. All satisfy The original vertex of vertex q is called the natural neighborhood of vertex q. The contribution weights of the natural neighborhood vertices to vertex q are calculated. Where Area represents the area; the eigenvalues of vertex q Original sampling points within the natural neighborhood of vertex q eigenvalues The weighted summation is used to obtain the result. The feature value Characterizing the original sampling points The local physical or physiological properties of the location, which have been determined before interpolation; The radial basis function interpolation method includes: for any vertex q of a sphere obtained by the icosahedral subdivision method, taking the original sampling points in the neighborhood of vertex q. , making ,in, The standard deviation of the Gaussian kernel function; weights are calculated based on the Gaussian kernel function. Where exp represents the natural exponential function; and the eigenvalues of vertex q are... Original sampling points within the neighborhood eigenvalues The weighted average; in, .
4. The system according to claim 3, characterized in that, The spherical convolution and pooling module specifically performs the following steps: Let u be the spherical distance between adjacent vertices. Spherical convolution will transform the spherical distance between the current vertex q into u. The vertex information within the neighborhood is synthesized, among which The radius of the spherical convolution kernel; Spacing from the current vertex q , and Each spherical convolutional kernel has 6 vertices, evenly distributed on 3 concentric circles. The kernel distributes convolution weights clockwise based on the meridian passing through the current vertex q, with a distance of [missing information] from the 12 vertices of the icosahedron. The vertex has 5 vertices, and the average of the 5 vertices is used as the filler. A spherical pooling method is established based on the recursive process of subdividing a regular icosahedron: The pooling method corresponding to a first-order subdivided regular icosahedron includes: if the current subdivision recursion number is l, then the central vertex and the outer hexagonal vertices are the spherical vertices after l-1 recursions, and the remaining vertices are the newly generated vertices in this recursion; spherical pooling gathers the information of the central vertex and all vertices first-order adjacent to the central vertex to the central point, and removes the newly generated vertices in this recursion; The pooling method corresponding to the second-order subdivision method includes: pooling the information of newly generated vertices adjacent to the center point and the center points of surrounding equilateral triangles to the center point.
5. The system according to claim 4, characterized in that, The spherical Transformer module specifically performs the following steps: Spherical graphic block division: Taking the vertex of the penultimate icosahedral subdivision as the center point, neighboring vertices of the second-order subdivision are taken to form graphic blocks; in addition to the 12 vertices of the icosahedron, each center point has 12 neighboring vertices, of which 6 adjacent vertices located at the center of the equilateral triangle are shared with surrounding graphic blocks to capture the contextual information between adjacent blocks; each vertex of the icosahedron has only 10 neighboring vertices, including 5 edge center points and 5 equilateral triangle face center points; the mean values of the edge and equilateral triangle face center points are respectively used to fill the vector; after flattening the spherical graphic blocks, a linear projection is performed to map each graphic block into a feature matrix. : , in Represents the real number field, where d is the vertex feature dimension, and each spherical graphic block has a total of 13 vertices; The spherical graphic block is positionally encoded using spherical Fourier position: , in These are the zenith angle and azimuth angle of the vertex in spherical polar coordinates, respectively. PE represents the learnable parameter matrix; PE represents the positional encoding; k represents the index variable for summation; Represents the tensor product; Establish a spherical Transformer model: For the feature vector h of the spherical graphic block, apply a sub-attention mechanism: , Here, Attention represents the attention mechanism, Q, K, and V are the query vector, key vector, and value vector, respectively; T represents matrix transfer; the spherical Transformer model multiplies the attention with the feature vector h, performs layer standardization, inputs it into the feedforward layer for synthesis, and then performs layer standardization again; Establish a spherical cross-modal Transformer model: The spherical cross-modal Transformer model includes a cross-modal attention module with D layers. The multi-head attention module is suitable for multi-source data. eigenvectors and multi-source data eigenvectors Cross-modal attention module Will Information input : , in query vector From the eigenvector get, key vector Sum value vector Depend on get, Let be the dimension of the key vector; , and These are the parameters to be trained; For a cross-modal attention module with D layers The multi-head attention module uses cross-layer connections in each layer to... With cross-modal attention module The output is first summed across layers, then normalized, and then input into the feedforward layer for nonlinear deformation, followed by a second summation and normalization across layers. The summation across layers refers to combining the input feature vector of the current layer with the summation of the previous layer's input feature vector. and The output feature vectors processed by the cross-modal attention module are then summed element by element.
6. The system according to claim 5, characterized in that the spherical sampling and downsampling module includes a downsampling layer and an upsampling layer; The downsampling layer uses a recursive process based on the subdivision of the icosahedron to merge spherical graphic blocks. The center point of the initial graphic block is the vertex obtained by the last first-order subdivision method. The mean of the feature vectors of all vertices in the graphic block is taken and converged at the center point of the graphic block. A weighted average is taken for all vertices in the block, and the weights are determined by the area ratio of the spherical Voronoi unit. , in, Let be the area of the Voronoi region at vertex i; This represents the sum of the areas of all Voronoi regions within the graphic block; This represents the eigenvector of the i-th vertex; This represents the output feature vector obtained after weighted pooling, which serves as the aggregated representation of the image patch; N represents the total number of vertices. Using the vertices of the previous level of subdivision as the convergence target, the feature mean of vertex i and its adjacent newly sampled vertices at this level is taken, and the downsampling process is used to simulate the subdivision recursive steps in reverse.
7. The system according to claim 6, characterized in that, The upsampling layer reconstructs vertices using a recursive process based on the icosahedral subdivision method. Starting from the vertex of the last icosahedral subdivision, it backtracks upwards layer by layer to the initial icosahedral structure. The parent vertex of each level is determined through the subdivision tree topology. Linear interpolation is used to restore low-resolution features to a high-resolution mesh. Perform first-order feature interpolation: For each side of an equilateral triangle, the two endpoints correspond to two original vertices A and B, respectively, and the feature vectors of the two original vertices A and B are as follows: and ,right and Take the average value to generate edge midpoint features. : , The first-order feature interpolation corresponds to the midpoint projection of the set of first-order subdivisions of a regular icosahedron; the set of subdivisions of the regular icosahedron refers to the set of two or more levels of mesh structures generated by recursively subdividing the initial regular icosahedron, with each level corresponding to the vertex and face topology formed by one subdivision operation. Perform second-order feature expansion: Insert two new vertices at each trisection point of the edge, and calculate the features of the two new vertices using linear interpolation. and : , , Then, the calculated values on each edge and That is, the characteristic mean of the six edge vertices is used as the characteristic of the center point of the surface.
8. The system according to claim 7, characterized in that, The application analysis module is used in the following application scenarios: Spherical autoencoder extracts depth features: The image acquisition and preprocessing module is used to establish a surface and extract surface features. A spherical mesh is established through the spherical uniform symmetric mesh generation module. Then, an encoder is established using the spherical Transformer module and the downsampling layer. A lightweight decoder is established using spherical convolution, spherical pooling and upsampling layers. The mean square error is used as the loss function. The output of the encoder is the depth features of the data. A segmentation model integrating multi-source data: The image acquisition and preprocessing module is used to establish a surface and extract surface features. A spherical mesh is established through a spherical uniform symmetrical mesh generation module. A spherical encoder is established using a spherical convolution and pooling module. The depth spherical features of the multi-source data are extracted respectively. The features are fused using a spherical Transformer module. After stitching, a U-Net network is established using a spherical convolution and pooling module and a spherical surface sampling and downsampling module to achieve data segmentation. A classification model that integrates multi-source data: The image acquisition and preprocessing module is used to establish a surface and extract surface features. A spherical mesh is established through a spherical uniform symmetrical mesh generation module. A spherical encoder is established using a spherical convolution and pooling module. The depth spherical features of the multi-source data are extracted respectively. The features are fused using a spherical Transformer module. After stitching, a ResNet network is established using a spherical convolution and pooling module to achieve data classification.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the system as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The system contains computer programs or instructions that, when run on a computer, execute the system as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Panoramic picture saliency prediction method and device based on full convolutional graph neural network
CN113947524A
Panoramic feature matching method based on non-Euclidean space
CN117541830A
Computer-aided detection method and system based on spherical convolution
CN118887215A
Saliency prediction method and system for 360-degree image
US20230245419A1