Geometric representation and inverse transformation restoration method of image and related equipment

By converting images into standard grids and defining metric functions, geometric representation and inverse transformation restoration of images are achieved using Fourier neural operators and quasi-conformal mappings. This solves the problem of decoupling photometric appearance from geometric structure in existing technologies, and enables high-fidelity image reconstruction and generation.

CN122023716APending Publication Date: 2026-05-12XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XI AN JIAOTONG UNIV
Filing Date
2026-01-19
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing image representation methods struggle to effectively decouple photometric appearance from intrinsic geometric structure, making it difficult to guarantee structural continuity and integrity during image transformation and generation, and lacking an explicit description of the image's intrinsic geometric structure.

Method used

By processing the input image into a standard grid, defining the source measure and converting it into the target measure function, and using the optimal transport solution model of the Fourier neural operator to solve the optimal transport mapping, the quasi-conformal mapping and Beltrami coefficient field are calculated to realize the geometric representation and inverse transformation restoration of the image.

Benefits of technology

It achieves lossless and reversible conversion of image attribute information into geometric representation, maintaining the integrity and continuity of the topological structure, effectively eliminating artifacts in image generation and interpolation tasks, and providing theoretically guaranteed structural continuity and high fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023716A_ABST
    Figure CN122023716A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of digital images, and relates to an image geometric representation and inverse transformation restoration method and related equipment, and the method comprises the steps: converting an input image into a standard grid, and defining a uniform source measure; constructing a target measure function based on the attribute information of the image; transmitting the source measure to the target measure through optimal transmission mapping, and generating an optimal transmission grid; and calculating the mapped Beltrami coefficient field as the geometric representation of the images of different dimensions. During reconstruction, the coefficient field is used to solve the quasi-conformal mapping to reconstruct the grid, and the original image is recovered according to the local area proportion of the grid. The invention also provides an implementation scheme for realizing rapid mapping solution by using the Fourier neural operator and constructing an end-to-end geometric variational auto-encoder. According to the technical scheme, high-fidelity image coding and decoding are achieved, ghosting can be effectively eliminated in tasks such as image interpolation and generation, and structural coherence is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital image technology and relates to a method and related equipment for geometric representation and inverse transformation restoration of images. Background Technology

[0002] In computer vision and machine learning, a common approach to representing images is to treat them as intensity or color values ​​on a pixel grid. This discrete, pixel-based representation serves as direct input to various learning models in existing technologies, such as convolutional neural networks. However, this pixel-based representation has inherent technical limitations. First, a single pixel contains only local information, making it difficult to effectively capture the global structure and contextual relationships of an image. Second, pixel-based statistical information loses the fine textures and internal structures of the image.

[0003] To overcome the locality limitation of pixel representation, existing techniques (such as convolutional neural networks CNNs) expand the receptive field by stacking convolutional layers in an attempt to capture a wider range of context. However, this approach only learns the structure indirectly and implicitly, rather than fundamentally solving the problem, and the learned features are still highly dependent on the local neighborhood of the pixel.

[0004] Other image representation methods exist in the prior art, such as Fourier transform or wavelet transform. Although these methods can capture the global frequency information of an image, their localization ability in the spatial domain is poor, making it difficult to accurately describe and preserve the local geometric features of the image.

[0005] In recent years, although deep learning models have been able to learn powerful hierarchical features from raw pixels, these learned features are inherently still deeply entangled with the image's photometric properties (such as brightness and contrast). In these representations, the image's intrinsic geometry is treated as an implicit, uncontrollable byproduct, rather than a primary entity that can be directly analyzed, manipulated, and preserved. This entanglement of photometric and geometric properties makes it difficult to guarantee the continuity and integrity of the image's intrinsic structure when performing image transformation, editing, or generation tasks, easily leading to artifacts or structural distortions.

[0006] Furthermore, other coordinate-based image representation methods have emerged in the prior art, such as implicit neural representations or two-dimensional Gaussian splatting techniques. These methods model images as continuous functions of neural networks or Gaussian units. While they excel in image fitting and compression, they essentially "memorize" or "fit" the appearance of an image implicitly through network weights. This representation still lacks an explicit and controllable description of the image's intrinsic geometry. More importantly, when performing image transformations or interpolations, they cannot provide theoretical guarantees of preserving topological structure and geometric continuity like methods based on differential homeomorphisms. Their ability to preserve structure relies entirely on the black-box network learning results, lacking interpretability and stability.

[0007] Therefore, how to decouple the representation of an image from its luminous appearance, extract a more essential representation method that focuses on its intrinsic geometric structure, and provide theoretically guaranteed explicit structural continuity and integrity during transformation and generation is a technical problem that urgently needs to be solved in the fields of image processing and computer vision. Summary of the Invention

[0008] This invention aims to address the shortcomings of existing image representation methods by providing a representation method that decouples the appearance of an image from its intrinsic geometric structure. Specifically, this invention addresses how to losslessly and reversibly convert the attribute information of an image into a geometric representation that preserves its topological structure and can be effectively utilized by machine learning models.

[0009] This invention is achieved through the following technical solution: A geometric representation method for an image, comprising: The input image is processed into a standard grid and a source metric is defined. The image's attribute information is then converted into a target metric function on the standard grid. The attribute information includes pixel intensity or other attributes derived from the original image. Based on the source and target measures, the optimal transport map is obtained by solving the optimal transport model of the Fourier neural operator. The standard grid is converted into the initial optimal transport grid according to the optimal transport map, and the initial optimal transport grid is optimized by matching the initial optimal transport grid with the target measures to obtain the optimized optimal transport grid. The deformation field from the standard grid to the optimal transmission grid is calculated, and the deformation field is represented as a quasi-conformal mapping. The Beltrami coefficient field corresponding to the quasi-conformal mapping is calculated, and the Beltrami coefficient field serves as the geometric representation of the image in different dimensions.

[0010] Preferably, the input image is processed into a standard grid and a source metric is defined. The image's attribute information is then converted into a target metric function on the grid. Specifically, the attribute information can be selected as pixel intensity, or other attributes such as gradient, curvature, texture features, or depth information. The input image is processed into a grid, and the feature points of the image's attribute information are set as grid vertices. A standard triangular grid is then constructed using these grid vertices. The uniform probability distribution of the standard triangular mesh is defined as the source measure; The attribute information of the input image is normalized and defined as a specific function in the image domain; Constructing a target measure function on a standard grid based on a specific function in the image domain; Before constructing the target measure function on the standard grid, a positive hyperparameter is added to process the zero-intensity region of the image.

[0011] Preferably, based on the source and target measure functions, the optimal transport map is obtained by solving the optimal transport solution model of the Fourier neural operator. The standard mesh is then converted into the optimal transport mesh according to the optimal transport map, specifically including: Based on the source and target measure functions, the target measure function and the coordinates of the standard grid are concatenated in the channel dimension to form an input tensor. The tensor is then input into the optimal transport solution model of the pre-trained Fourier neural operator, which maps the tensor from the low-dimensional space to the high-dimensional latent feature space to obtain the optimal transport mapping. The optimal transport map is applied to each vertex of the standard grid, and an optimal transport grid is formed through an energy optimization method.

[0012] Preferably, the specific process for obtaining the Beltrami coefficient field is as follows: For each triangular facet in the standard grid, calculate the Beltrami coefficient of each triangular facet based on the corresponding vertex coordinates of the triangular facet in the standard grid and the optimal transport grid; For each vertex on the standard grid, the average of the Beltrami coefficients of all adjacent triangular faces of the vertex is selected as the Beltrami coefficient of that vertex; The Beltrami coefficients of each vertex on all standard grids are combined to form a Beltrami coefficient field in the image domain.

[0013] An image inverse transform restoration method based on geometric representation, comprising, Based on the geometric representation of images in different dimensions, the corresponding quasi-conformal mapping is calculated by solving the linear Beltrami equation, and the standard mesh is transformed into a reconstructed quasi-conformal mesh through the quasi-conformal mapping. The local area of ​​each vertex in the reconstructed quasi-conformal mesh is calculated. Based on the local area, the attribute information features of each vertex are calculated inversely through the inverse operation of the target measure function, and the reconstructed image is recovered.

[0014] An image processing method, implemented based on the aforementioned geometric representation method of the image and the image inverse transform restoration method based on the geometric representation, includes: The geometric representations of the images in different dimensions are used as input to the autoencoder model of the image; In the latent space of the autoencoder model of the image, interpolation or operation is performed on the geometric representations corresponding to the image in different dimensions to obtain the interpolation or operation results; The interpolation or calculation results are decoded into an image output using the geometric representation-based inverse image transformation restoration method.

[0015] An image geometric representation system for performing the geometric representation method of the image, comprising: The preprocessing module is used to process the input image into a standard grid and define the source measure, converting the image's attribute information into a target measure function on the grid; The mapping calculation module obtains the optimal transport mapping by solving the source and target measure functions using Fourier neural operators, and then converts the standard grid into the optimal transport grid based on the optimal transport mapping. The geometric feature extraction module calculates the deformation field from the standard grid to the optimal transmission grid, represents the deformation field as a quasi-conformal mapping, and calculates the Beltrami coefficient field corresponding to the quasi-conformal mapping. The Beltrami coefficient field and the attribute information of the image together constitute the geometric representation of the image in different dimensions.

[0016] An image restoration system for performing the geometric representation-based inverse image transformation restoration method, characterized in that it comprises: The mesh reconstruction module is used to calculate the corresponding quasi-conformal mapping based on the geometric representation of the input image by solving the linear Beltrami equation, and to transform the standard mesh into a reconstructed quasi-conformal mesh through the quasi-conformal mapping. The image restoration module calculates the local area of ​​each vertex in the reconstructed quasi-conformal mesh. Based on the local area, it calculates the attribute information features of each vertex through the inverse operation of the target measure function. The attribute information features at the vertex are then mapped to the image pixel grid to restore the reconstructed image.

[0017] A computer-readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements a geometric representation method of the image as described.

[0018] An image processing device, comprising: One or more processors; Memory, which stores computer program instructions; When the computer program instructions are executed by the one or more processors, the device performs a geometric representation method of the image, or implements an image inverse transformation restoration method based on geometric representation, or runs an image processing method.

[0019] Compared with the prior art, the present invention has the following beneficial technical effects: This invention discloses a geometric representation and inverse transformation restoration method and related equipment for images. The method involves transforming the image's attribute information into a Beltrami coefficient field describing grid deformation using optimal transfer theory, thereby obtaining a structure-preserving and reversible geometric representation of the image. Based on optimal transfer theory and quasi-conformal mapping theory, the entire image is reformulated as a differential homeomorphic deformation field, where the inherent geometric information itself constitutes a complete and robust representation of the image structure. The process includes converting the input image into a standard grid and defining a uniform source measure; constructing a target measure function based on the image's attribute information; transferring the source measure to the target measure through optimal transfer mapping to generate an optimal transfer grid; and calculating the Beltrami coefficient field of this mapping as the geometric representation of the image. During reconstruction, the quasi-conformal mapping is solved using this coefficient field to reconstruct the grid, and the original image is recovered based on the local area ratio of the grid. This invention also provides an implementation scheme utilizing Fourier neural operators to achieve fast mapping solution and constructing an end-to-end geometric variational autoencoder. This technical solution achieves high-fidelity image encoding and decoding, and can effectively eliminate ghosting and maintain structural coherence in tasks such as image interpolation and generation.

[0020] Furthermore, through optimal transport and quasi-conformal geometry theory, this invention encodes these physical or geometric quantities into physically meaningful grid deformations. This representation not only decouples the underlying features but also provides a topology-preserving geometric descriptor for subsequent, more abstract semantic understanding. It fundamentally decouples the image's attribute information from its intrinsic geometric structure, representing the image as a reversible, structure-preserving geometric deformation field. This representation is more fundamental and less susceptible to external factors such as illumination and contrast. Since the entire transformation process is based on differential homeomorphisms, it ensures the integrity of the image's topological structure. Therefore, in tasks such as image generation and interpolation, operations in this geometric representation space naturally maintain the continuity and smoothness of the structure, avoiding artifacts and structural breaks common in pixel space. Moreover, based on the measurable Riemannian mapping theorem, the representation method of this invention is theoretically lossless and completely reversible, ensuring high-fidelity image reconstruction. The encoding and decoding processes of this method can be seamlessly integrated into existing deep learning frameworks as independent plug-and-play modules, providing novel, structure-preserving input representations for various learning tasks. Experiments demonstrate that in tasks such as image reconstruction and interpolation, this method outperforms traditional pixel-based representation methods in both quantitative metrics and qualitative visual effects, generating smoother, more continuous, and perceptually higher quality results. Attached Figure Description

[0021] Figure 1 Flowchart for converting an image into a geometric representation; Figure 2 A flowchart for geometrically representing and restoring an image; Figure 3 To compare handwritten digit interpolation in the latent space of an autoencoder using image geometric representation; Figure 4 This is a schematic diagram of the network architecture of the end-to-end geometric variational autoencoder (GIR-AE) model in Embodiment 4 of the present invention; Figure 5 This is a schematic diagram of the specific structure of the Fourier Neural Operator (FNO) model in Embodiment 4 of the present invention, illustrating the parallel processing flow of spectral convolution and residual connection; Figure 6 This is a schematic diagram of the specific structure of the RefineUNet in Embodiment 4 of the present invention, showing the U-Net architecture with skip connections and identity residual branches; Figure 7 This is a schematic diagram showing different objective functions for the same image.

[0022] Figure 8 A comparison of the optimal transmission model predicted by the optimal transmission solution model based on Fourier neural operators and the optimal transmission grid of the traditional geometric semi-discrete optimal transmission solution for handwritten digits. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0024] This invention proposes a unified computational method for mapping multidimensional features of an image to geometric deformation. The method first treats the image's attribute information as a generalized quality distribution, obtaining a deformable mesh reflecting the deep content structure of the image through optimal transport theory. Then, using quasi-conformal geometry theory, this deformation process is uniquely encoded as a Beltrami coefficient field, serving as a universal geometric representation of the image. This representation method can decouple and capture the fine, global intrinsic structure of the image. Furthermore, this invention provides a method for accurately recovering the original image from this geometric representation. Since the entire transformation process is differentially homeomorphic, this method can well preserve the topological structure when processing and generating images, and can be embedded as a plug-and-play module into existing learning models. This invention further proposes a Fourier neural operator to replace the traditional semi-discrete optimal transport method, combining explicit partial differential equation solving with deep neural networks to achieve efficient learning and inference of the geometric representation. Experiments demonstrate that the interpolation continuity of this geometric representation in the latent space of the same autoencoder is superior to that of the traditional pixel-based representation. It should be noted that the "image attribute information" described in this invention has a broad meaning. Although the preferred embodiments and formula derivations below mainly use the "pixel intensity (grayscale value)" of the image as an example for the sake of simplicity and intuitiveness, this does not constitute a limitation on the scope of protection of this invention. In practical applications, the attribute information can be completely replaced by any multidimensional attribute derived from the original image, including but not limited to: (1) geometric attributes: such as the gradient magnitude, normal vector, local Gaussian curvature, average curvature, etc. of the image; (2) visual attributes: such as texture features (direction, contrast, etc.); (3) spatial attributes: such as depth map information. By mapping the above feature functions to the target measure function of the grid, the geometric representation of the multidimensional features of the image is realized; the geometric representation can be used as the underlying feature to support more abstract semantic analysis and processing. Those skilled in the art will understand that after normalizing any of the above attributes to the target measure function, the same processing flow as the "pixel intensity" described below can be used to obtain the geometric representation of the corresponding dimension.

[0025] The first aspect of the present invention discloses a geometric image representation method, comprising the following steps: By using optimal transport theory and auxiliary metric methods, a quasi-conformal mapping corresponding to the attribute information of the image is constructed, transforming the original image into a geometric representation, and the original image can be recovered from this geometric representation by inverse transformation.

[0026] S1. Image to Standard Mesh Conversion: The input image is converted into a standard triangular mesh with a metric. The attribute information feature points of the image are initialized as vertices of the standard triangular mesh. Specifically, a standard uniform triangular mesh is initialized in the image domain as the standard mesh, and the attribute information of the image is regarded as a quality or density. After normalization, it is defined as the target metric function on the standard mesh. S2. Calculate the optimal transport mapping: Calculate an optimal transport (OT) mapping that transforms a standard uniform triangular mesh into an intensity-aware target measure mesh such that the local area distribution of the deformed mesh matches the target measure function defined in S1. S3. Generate geometric image representation: The deformation field from the standard mesh to the optimal transport mesh is represented as a quasi-conformal mapping, and the Beltrami coefficient field corresponding to this mapping is calculated. This Beltrami coefficient field serves as the geometric representation (G-IR) of images of different dimensions in this invention. Furthermore, in S1, in order to handle zero-intensity regions, the pixel intensity value is increased by a small positive hyperparameter before being converted to the target measure function.

[0027] More preferably, after S2, an optimal transmission optimization step is added: by minimizing an energy functional designed to balance image spatial fidelity and deformation field regularity, this energy function is designed to penalize the fidelity error between the reconstructed image and the original image, as well as the regularity of the deformation field. The vertex positions of the initial optimal transmission grid are fine-tuned to improve approximation accuracy and suppress visual artifacts.

[0028] Furthermore, in step S2, the method for calculating the optimal transport mapping includes prediction using an optimal transport solution model based on Fourier neural operators; the optimal transport solution model based on Fourier neural operators adopts a Fourier neural operator (FNO) architecture, the input of the model is a tensor containing image pixel coordinates and attribute information, and the output is the vertex coordinates of the optimal transport grid; the model learns the mapping operator from image density to geometric deformation by performing a linear transformation on low-frequency patterns in the Fourier frequency domain.

[0029] In S3, the Beltrami coefficient field is constructed as follows: S3.1 For each triangular facet in the mesh, calculate the Beltrami coefficient of a piecewise constant based on its corresponding vertex coordinates in the standard mesh and the optimal transmission mesh; S3.2 For each vertex on the mesh, take the average of the Beltrami coefficients of all its adjacent triangular faces as the Beltrami coefficient of that vertex; S3.3 Combine the Beltrami coefficients of all vertices to form a complete Beltrami coefficient field defined over the entire image domain.

[0030] A second aspect of the present invention discloses an image inverse transform restoration method based on geometric representation, comprising the following steps: S1. Calculate the quasi-conformal mapping: Based on the given geometric image representation (i.e., the Beltrami coefficient field), the corresponding quasi-conformal mapping is calculated by solving the linear Beltrami equation. This mapping transforms a standard uniform triangular mesh into a reconstructed target metric mesh. S2. Reconstructing the image from the grid: Calculate the local area associated with each vertex in the reconstructed target measure grid, and restore the image attribute information by inverse operation of the density definition function in S1, thereby obtaining the reconstructed image; More preferably, after S1, a quasi-conformal mapping optimization step is added: the vertex positions of the reconstructed target measure grid are optimized by minimizing an energy functional that aims to align the computed grid Beltrami coefficient field with the target measure function, in order to correct the errors accumulated during discretization and numerical solution.

[0031] More preferably, the method for solving the quasi-conformal mapping is as follows: using an auxiliary metric corresponding to the Beltrami coefficient field, combined with a parameterization method based on Holomorphic 1-forms or a method based on Curvature Flows, the mesh after the quasi-conformal mapping is calculated.

[0032] Furthermore, this invention provides a general mapping method from image features to geometric deformations. Any continuous / discrete function capable of expressing local or global image features (such as grayscale, color components, gradient, geometric curvature, etc.) can be used as the input measure of this general mapping method. Through optimal transport and quasi-conformal geometry theory, this invention encodes these abstract functional information into physically meaningful grid deformations, thereby providing a unified mathematical representation with topology-preserving properties for image processing.

[0033] An image processing method based on geometric representation is provided, which obtains the geometric representation (Beltrami coefficient field) of an image; performs operations on the geometric representation in the latent space, such as interpolation; and then reconstructs the new image from the geometric representation by means of the image inverse transform restoration method steps.

[0034] A third aspect of the present invention discloses a geometry-based image representation and restoration system, comprising: The representation module is used to execute the image representation method described in the first aspect of the present invention, converting the input pixel image into a geometric image representation.

[0035] The recovery module is used to execute the image inverse transformation restoration method described in the second aspect of the present invention to recover the pixel image based on a given geometric image representation.

[0036] An image generation network model based on geometric representation, wherein the network is constructed as an end-to-end geometric variational autoencoder (GIR-AE) network model, comprising the following sub-modules: Encoding and Deformation Module: Used to receive input images. First, it uses the optimal transmission solution model of the pre-trained and frozen Fourier Neural Operator (FNO) to predict the coarse-grained optimal transmission grid. Then, it corrects the grid coordinates through the first refinement network (RefineUNet) and calculates the corresponding Beltrami coefficient field as the latent geometric representation of the image. Differentiable geometry bridging module: used to receive the Beltrami coefficient field, solve the Laplace system with boundary constraints through the embedded differentiable linear Beltrami solver, and generate a quasi-conformal mesh (QC Mesh). Decoding and reconstruction module: It receives the quasi-conformal mesh, fine-tunes its structure through a second refinement network, and finally recovers the reconstructed grayscale image based on the mapping relationship between the local area ratio and density of the mesh vertices; the entire network remains fully differentiable, allowing the parameters of the refinement network to be optimized through backpropagation of reconstruction loss.

[0037] A fourth aspect of the present invention discloses a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the methods described above.

[0038] A fifth aspect of the present invention discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above methods.

[0039] The technical solution of the present invention will be clearly and completely described below. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0040] This invention provides an image representation and inverse transformation restoration method based on geometric transformation. The core idea is to convert the pixel intensity information of an image into an equivalent geometric representation describing shape changes, namely the Beltrami coefficients, thereby achieving structure-preserving and reversible image encoding and decoding. The datasets used in this paper are all publicly available image datasets. Figure 1 As shown, it includes the following steps: S1. Initialize the standard mesh Image to standard grid conversion S1.1, Transfer the image Convert to a triangular mesh, its mesh vertices Defined in the image domain The structure of the grid is constructed by triangulating the points on the grid.

[0041] S1.2 Defining the Source Measure: Before performing optimal transmission, we first define the source measure in the image domain, denoted as . The source measure is defined by the standard grid. The uniform probability distribution on the source is mathematically represented as a normalized Lebesgue measure. The method of this invention aims to find a mapping that represents this uniform source measure. Transform into a target measure that reflects the distribution of image features (such as pixel intensity distribution). ; S1.3, Set the vertex positions of the mesh as follows Each vertex corresponds to the center of a pixel in the image. Standard grid. It is equipped with a uniform source measure, defined as a normalized Lebesgue measure, as the source measure of the geometric space. ; S2. Convert the image intensity of each pixel into a target measure.

[0042] S3. Define the target measure function based on the target measure. It is related to image intensity Proportional, and using hyperparameters For handling zero-intensity regions, the target measure function formula

[0043] Optimal transport mapping: Use the optimal transport mapping to change the mesh so that its corresponding measure is changed from the source measure. Transform into target measure .

[0044] Based on the source and target measure functions defined in S1, an optimal transport map is calculated, and an optimal transport mesh (OT Mesh) is generated. S3. Calculate the deformation field from the standard grid in S1 to the optimal transmission grid in S2, and represent the deformation field as a unique Beltrami coefficient field, which is the geometric representation of the image (G-IR). S4. Based on the Beltrami coefficient field obtained in S3, the standard mesh is transformed into a reconstructed quasi-conformal mesh (QC Mesh) by solving the quasi-conformal mapping. S5. Based on the local area information of each vertex in the quasi-conformal mesh obtained in S4, the intensity value of each pixel is calculated in reverse to reconstruct the original image.

[0045] S1. Image to standard grid conversion S1.1 Initialize the standard grid: Set the image... Convert to a triangular mesh, its mesh vertices Defined in the image domain The structure of the grid is constructed by triangulating the points on the grid.

[0046] S1.2 Normalize the image and define it as an intensity function. : Then, a target metric is defined based on the image intensity, and then mapped using optimal transmission based on the target metric.

[0047] Set the vertex positions of the mesh as Each vertex corresponds to the center of a pixel in the image. Standard grid. A uniform source measure is provided, defined as a normalized Lebesgue measure, as the initial uniform measure of the geometric space. ; S2, Optimal Transfer Mapping of Images S2.1 Definition of Target Measure: Define a target measure function that is proportional to the image intensity. And by introducing small hyperparameters We then process the zero-intensity region and normalize it.

[0048]

[0049] S2.2, Optimal Transfer Mapping: Using the optimal transfer mapping to measure the uniform source of the image. Converting target measures into intensity perception The new vertex position is obtained by using an optimal transport solution model based on Fourier neural operators. This yields the initial OT mesh.

[0050] S2.3 Optimize vertex positions on a discrete mesh. By minimizing the energy functional To refine the mapping, thereby improving the approximation accuracy while maintaining the topological consistency of the triangular mesh. The formula is as follows:

[0051] The energy functional comprises three parts: S2.3.1, Fidelity Term: Used to minimize the average error between the reconstructed image and the original image over the entire domain.

[0052] S2.3.2, Uniformity term: Used to penalize the maximum point-by-point deviation, thereby suppressing local artifacts and forcing uniform quality.

[0053] S2.3.3 Regularization term: Used to constrain the optimized vertex positions to be within the neighborhood of the initial optimal transmission mesh, preventing excessive mesh deformation and ensuring the stability of the topology.

[0054] S3. Generate Quasi-Conformal (QC) Maps / QC Optimization S3.1 Calculate the Beltrami coefficient for quasi-conformal mapping: The deformation from the standard mesh to the OT mesh is calculated using the Beltrami coefficient. :

[0055] Eccentricity This represents the local rotation angle.

[0056] S3.2, QC Optimization: Using the Beltrami coefficients obtained in S3.1, the mesh vertices are optimized. A piecewise constant Beltrami coefficient is calculated for each triangle face, derived from the corresponding triangles in the standard and OT meshes. The Beltrami coefficient of each vertex is obtained by averaging the Beltrami coefficients of its neighboring triangles. The Beltrami coefficient field is represented graphically as follows:

[0057] S4, by Solve the initial QC mapping Using the linear Beltrami solver, the vertex positions are obtained by solving on a standard mesh. As an initial solution, we obtain:

[0058] Due to discretization and numerical errors, direct solutions... With the goal There may be deviations; optimize and fine-tune the mesh through QC to minimize them. The optimized network vertices are obtained. .

[0059] in

[0060] S5. Calculate the vertex area measure from the final mesh. exist Above, for each vertex By summing the areas of its adjacent triangles (with weight), we obtain... Let the area measure vector of all vertices be... The final area ratio vector is obtained through normalization calculation:

[0061] Then, the grayscale is recovered by inverse transformation:

[0062] Will Interpolate back to the pixel raster to obtain the reconstructed image. .

[0063] Example 1 In one alternative embodiment, the method includes the following steps: 1. Receive the input image and normalize its intensity value into a target measure; 2. By using optimal transport mapping, the standard grid is transformed into an OT grid corresponding to the image intensity distribution. 3. Calculate the Beltrami coefficient based on the local Jacobian matrix of standard and OT meshes. ,in:

[0064] To ensure the differentiability and orientation preservation of the mapping; 4. Calculate the Beltrami coefficient. With intensity normalization parameter Together they form a geometric image representation ( -Image), used to characterize the structural and deformation information of an image; 5. Solve the problem through QC optimization, from Reconstruct the mesh vertex positions and calculate the area ratio of each vertex.

[0065] 6. Basis and Inversely calculate the grayscale value of each pixel. This enables reversible recovery from geometric space to pixel space.

[0066] Table 1 shows the loss metrics for recovery on different datasets;

[0067] Table 1 presents the quantitative experimental results, where each metric is obtained by averaging 10 random subsets of 500 test images on each dataset. Despite slight numerical errors due to discrete computation, our results still exhibit extremely high fidelity: SSIM (structural similarity) is close to 1.0, LPIPS (perceptual dissimilarity) is close to 0, and PSNR (peak signal-to-noise ratio) exceeds 30 dB. Example 2 This embodiment uses the Geometric Semi-Discrete Optimal Transport (SDOT) numerical method to solve for the optimal transport grid. Although this method has a higher computational cost than the FNO model, it possesses rigorous mathematical guarantees and convergence. The specific steps are as follows: 1. Define the discrete measure problem: Define the semi-discrete transport problem: Model mesh generation as a transport process from a continuous source measure (standard mesh) to a discrete target measure (the set of Dirac measures of image pixels).

[0068] 2. Constructing the power map: According to optimal transport theory, this mapping is entirely determined by a set of potential functions, or weight vectors. Decision. For a given weight vector. The corresponding transport map is defined by the power map. The power map divides the standard grid domain into multiple polygonal cells, each cell corresponding to an image pixel.

[0069] 3. Constructing a convex energy functional: In order to find the optimal weight vector In this embodiment, a convex energy functional is constructed. The gradient of this functional corresponds to the difference between the current cell cavity area and the target pixel intensity (target quality).

[0070] 4. Iterative Solution using Newton's Method: The energy functional above is minimized using the damped Newton's method. In each iteration: S1. Calculate the current power graph structure and its intersection with the standard grid; S2. Calculate the area of ​​each cell cavity; S3. Calculate the Hessian matrix and gradient vector; S4. Update the weight vector This continues until the ratio error between the area of ​​each cell cavity and the pixel intensity of the target image at that point is less than a preset threshold.

[0071] 5. Generating the optimal transport grid: When the energy functional converges, the final power graph cell centroid determines the optimal transport grid. Vertex position.

[0072] Example 3 This embodiment details how to train a traditional autoencoder model (G-IRAE) using image-to-geometric representations, which utilizes image-to-geometric representations (Beltrami coefficients). It replaces traditional images as input and realizes image generation and interpolation in the latent space.

[0073] The method flow of this embodiment mainly includes three stages: model input generation stage, model training stage, and inference and interpolation application stage.

[0074] 1. Model building phase: 1.1 G-IR Solution (Image to ): Used to input image Mapping to geometric features .

[0075] 1.1.1 Input Preprocessing: Preprocess the input grayscale image (size Normalize and treat as a density function .

[0076] 1.1.2 Feature Extraction: The Beltrami coefficient is obtained using this method. The real and imaginary parts (or modulus and argument). This is the input and output pair of the G-IR AE. G-IR is used to replace traditional image data for training the traditional AE.

[0077] 1.2 Differentiable Geometric Bridging Module This is a parameterless mathematical solver layer embedded in the network. It receives the encoder output. By solving the linear Beltrami equations (using a sparse matrix solver), the vertex coordinates of the corresponding quasi-conformal mesh (QC Mesh) are calculated. .

[0078] 1.3 Image Restoration Module ( (to Image) 1.3.1 Mesh Reconstruction: Receives the mesh calculated by the bridging module. .

[0079] 1.3.2 Pixel Recovery: Calculate the intensity value of each pixel based on the deformation of the grid cells (local area ratio). The denser the grid, the higher the pixel intensity.

[0080] 1.3.3 Image Output: The recovered intensity is mapped back to the pixel space of [0,255] to obtain the reconstructed image.

[0081] 2. Model Training Phase Compared to AEs that are typically trained directly with images, the difference at this stage is that the loss function is restricted to G-IR during training.

[0082] 2.1 Output Constraints: To satisfy the mathematical constraints of quasi-conformal mappings The Tanh activation function is used in the output layer of the network to restrict the output value to ( ). The modulus can be set to a 1,1) interval, or a scaling operation can be used to ensure the modulus is less than 1. In this case, the network output is the latent geometric representation of the image. .

[0083] 2.2 Since G-IR can be represented using the same data format as images, loss functions can be used directly. For example, the MSE loss and L1 loss of G-IR can be measured.

[0084] 3. Applications of Reasoning and Interpolation After training, the model is used to solve... Figure 3 The problems of "image interpolation" and "coherent generation" are shown.

[0085] S3.1 Encoding: Given two different source images (e.g., the number "1") and (For example, the number "7") convert them into their respective geometric representations. and Then, the latent representation is input into the pre-trained encoder of the AE. and .

[0086] S3.2 Perform linear interpolation in the geometric latent space. Set the interpolation coefficients. Calculate the latent representation of intermediate states :

[0087] S3.3 Decoding and Generation: The interpolated result... The input is fed into the geometry bridging module and the image restoration module. It is first restored via the decoder. The system then solves. Corresponding intermediate grid Then based on the intermediate grid Density distribution generates intermediate images .

[0088] Effect description: With The generated image sequence, from 0 to 1, shows the progression from... arrive The smooth structural evolution (such as the bending and stretching of strokes) rather than the simple superposition of pixels (fade in and fade out) effectively eliminates the ghosting effect, verifying that the geometric representation extracted in this invention successfully decouples the topological information of the image.

[0089] This scheme uses μ-Image as the encoding input for latent space modeling and generation tasks, achieving structure-preserving and continuous latent space interpolation. Table 2 shows that the G-IR AE model achieved lower neighbor-to-neighbor LPIPS values ​​in all experimental pairs. Experiments demonstrate that G-IR AE has significant advantages over pixel autoencoders (Pixel-IR AE) in interpolation smoothness and perceptual quality, effectively eliminating ghosting, noise, and abrupt changes. The interpolation results of G-IR AE maintain structural and textural coherence, resulting in a visually natural appearance. Experiments verify that the latent space induced by G-IR possesses stronger geometric consistency and interpretability. Figure 3 This demonstrates the changes in images interpolated using the latent space of the apocalypse (AE). Figure 3 Each set (two rows of identical handwritten digits) shows the image changes after interpolation and reconstruction of different handwritten digits in the latent space. The top row of images in each set is trained based on GIR, and the bottom row is trained directly from the image. The training AE architecture is completely consistent. It can be seen that GIR has better continuity.

[0090] Table 2 shows the changes in LPIPS between adjacent numbers with different latent space differences;

[0091] Example 4 This embodiment proposes an optimal transport solution model based on the Fourier Neural Operator (FNO). In step 4 of Embodiment 2, traditional optimal transport mapping solutions typically involve optimization problems, which are computationally time-consuming. To achieve real-time image geometric representation, this embodiment utilizes an FNO network to learn the mapping relationship from the image density function to the vertices of the optimal transport grid. For example... Figure 5 As shown, the model's logical structure mainly consists of three interconnected parts: a feature dimensionality enhancement module, an iterative Fourier integral layer module, and a projection decoding module. The specific functions and structures of each module are as follows: 1. Feature Upscaling Module: This module serves as the input interface for the model, used to map low-dimensional physical space data to a high-dimensional latent feature space.

[0092] 2. Iterative Fourier Integral Layer Module: Composed of multiple stacked Fourier layers, serving as the core processing unit. Each layer contains parallel spectral convolution branches and spatial residual branches, which transform features in the frequency and spatial domains respectively to capture the global geometric features and long-range dependencies of the image; 3. Image Decoding Module: This module is located at the end of the network and maps high-dimensional features back to the target physical quantity.

[0093] The specific implementation steps of this method are as follows: 1. Dataset Construction and Data Preprocessing To train the network, we first construct an image dataset containing widely distributed features. For each image in the dataset, perform the following operations: 1.1 Input Data Generation: Convert the image to grayscale and normalize it to obtain a size of [size missing]. The pixel intensity distribution is considered as the target measure function. Construct a standard grid consistent with the image resolution, containing the normalized coordinates of each pixel. The grid target measure and grid coordinates are concatenated along the channel dimension to form the model's input tensor, which has the following shape: ,in B This represents the batch size, where 2 represents... Two channels.

[0094] 1.2 Label Data Generation: Using the geometric semi-discrete optimal transport solver described in Example 2, the optimal transport grid corresponding to the image is calculated. The calculated OT grid vertex coordinates are used as truth labels, and their tensor shape is... , representing the coordinate position of each grid point after deformation.

[0095] 2. Construction of the Optimal Transfer Model (FNO) for the Fourier Neural Operator: This embodiment constructs a 2D FNO model to approximate the optimal transfer operator. Since the optimal transfer problem is mathematically closely related to partial differential equations (PDEs), FNO, as an operator capable of learning mappings in infinite-dimensional function spaces, is very suitable for this type of task. The model mainly includes the following modules: 2.1 Dimension Upgrading Layer: Through a fully connected layer, the input tensor is mapped from a low-dimensional space (2 channels) to a high-dimensional latent feature space.

[0096] 2.2 Fourier Spectral Convolutional Layer: This is the core of the model, containing multiple stacked spectral convolutional blocks. Within each block, the following parallel operations are performed: 2.2.1 Frequency Domain Transformation: Perform a two-dimensional fast Fourier transform (2D FFT) on the input features to transform the data from the spatial domain to the frequency domain.

[0097] 2.2.2 Low-frequency filtering and linear transformation: Low-frequency modes are truncated in the frequency domain, and a learnable complex weight matrix is ​​used to perform a linear transformation on the spectrum. This enables the model to capture global geometric deformation features.

[0098] 2.2.3 Inverse Transform: The two-dimensional inverse Fourier transform is used to restore the processed spectrum back to the spatial domain.

[0099] 2.2.4 Residual connection: In addition to the frequency domain branch, a 1×1 convolutional layer in the spatial domain is connected in parallel, its output is added to the result after the inverse transform, and then passed through a nonlinear activation function (such as GELU).

[0100] 2.2 Projection Layer: Through a series of fully connected layers (or 1×1 convolutions), the high-dimensional latent features are mapped back to the target output dimension (2 channels), and the output is the predicted optimal transmission grid vertex coordinates.

[0101] 3. Model Training and Inference 3.1 To ensure that the optimal transport mapping predicted by the FNO model not only approximates the true value numerically, but also satisfies the properties of differential homeomorphism and topology preservation in terms of geometric structure, this embodiment designs a composite loss function for model training.

[0102] The loss function It consists of the following six weighted components:

[0103] 3.1.1 Data Fidelity Item: Measures the error between the predicted grid and the vertex coordinates of the true OT grid.

[0104] 3.1.2 Curl Regularization Terms: Calculate the L2 norm of the curl of the predicted displacement field. This term is used to constrain the local rotation of the deformation field and suppress non-physical distortion. The formula is expressed as:

[0105] 3.1.3 Laplacian Smoothing Terms: Calculate the L2 norm of the Laplacian operator for the displacement field. This term forces the output deformation field to remain spatially smooth, avoiding sharp mesh distortions.

[0106] 3.1.4 Beltrami Coefficient Constraint: Calculates the Beltrami coefficient for the mapping from the standard mesh to the predicted mesh. According to quasi-conformal geometry theory, if and only if When the mapping is locally bijective and unfold-free, a penalty function is introduced. Strong constraints are applied to the region, thereby explicitly forcing the model to generate non-self-intersecting differential homeomorphisms during training. For each triangular facet... Calculate the Beltrami coefficient: , ; 3.1.5 Area Consistency Constraint: Calculate the area of ​​each triangular facet in the predicted mesh and compare it with the area of ​​the corresponding facet in the target OT mesh.

[0107] The constructed dataset is input into the FNO model, and the aforementioned composite loss function is minimized using the backpropagation algorithm. In the early stages of training, data fidelity terms are given higher weights to facilitate rapid convergence; in the later stages of training, the weights of Beltrami constraints and regularization terms are gradually increased to finely adjust the geometric properties of the mesh.

[0108] 3.2 Inference process: After the model is trained, for any new input image, only its target measure needs to be input. After one forward propagation, the corresponding geometric representation (OT mesh) can be directly output without iterative solution, which greatly improves the speed of generating geometric representation (G-IR).

[0109] Experimental comparisons show that this method has significant advantages: (1) Speed ​​improvement: FNO reduces the traditional method's second-level iterative optimization to millisecond-level inference, supporting video stream processing; (2) High-fidelity generalization: Test MSE < 1×10 -4 (3) Resolution independent: It has discretization invariance, supports low-resolution training and high-resolution inference, and does not require retraining. Figure 8 This paper compares the optimal transport grid predicted by the optimal transport solution model based on Fourier neural operators (red grid) and the optimal transport grid predicted by the traditional geometric semi-discrete optimal transport model (blue grid). It can be observed that the two grids largely overlap. Visualization results show that the predicted OT grid not only accurately captures the density distribution characteristics of the image but also maintains a good topological structure, without grid folding or self-intersection. This proves that FNO has successfully learned the operator rules from the density function to the differential homeomorphism.

[0110] Example 5 This embodiment proposes an end-to-end geometric variational autoencoder (GIR-AE) network model. This architecture uses the FNO module from Embodiment 3 as its foundation, combining a differentiable geometric solver with a thinning network to achieve end-to-end learnable reconstruction from image to geometric representation and back to image.

[0111] like Figure 4 As shown, the overall data flow processing procedure of the GIR-AE network is as follows: 1. Encoding stage (image to latent representation): The input is a grayscale image. (dimension) ).

[0112] 1.1 Coarse-grained deformation prediction: Predicting coarse-grained OT meshes using pre-trained and frozen FNO modules. For example... Figure 5 As shown, after the input features are concatenated with the coordinate grid, they are first mapped to a high-dimensional space through a projection layer; then, through four stacked Fourier layers, each layer performs frequency domain spectral convolution (capturing global deformation) and spatial domain 1×1 convolution (preserving local details) in parallel; finally, the output projection layer reduces the dimensionality to obtain two-channel coarse-grained grid coordinates.

[0113] 1.2 Mesh Refinement (RefineUNet1): The first refinement network (e.g., using an improved U-Net architecture) is applied to the coarse mesh input. Figure 6 This network extracts multi-scale features step-by-step through the encoder, and recovers high-frequency spatial information through upsampling and skip connections in the decoder. The key lies in the residual output mechanism: the network output is designed as "input grid + prediction correction," forcing the model to focus on learning minute geometric deviations, thereby significantly improving the stability and convergence speed of the generated data.

[0114] 1.3 Geometric Representation Generation (Module 1): Convert the refined OT mesh into a Beltrami coefficient field. This step calculates the local conformal distortion at each vertex, and the output is... It contains two channels, real and imaginary, which is the geometric representation of the image in the latent space.

[0115] 2. Geometric bridging stage (explicitly differentiable solution): 2.1 Quasi-conformal Mesh Generation (Module 2): This module is a fully differentiable geometric operator without learning parameters. It accepts Beltrami coefficients. The module generates a quasi-conformal mesh (QC Mesh) by constructing and solving a Laplacian linear system with boundary conditions. Specifically, the module efficiently constructs a sparse matrix using the scatter operation and solves for the mesh vertex coordinates using linear least squares. This process is designed to maintain the continuity of the computation graph, thus allowing gradients to be backpropagated from the reconstructed image to the latent representation. .

[0116] 3. Decoding stage (geometric representation to image): 3.1 Reconstructed Mesh Refinement (RefineUNet2): The QC mesh obtained from the solution is input into the second refinement network (UNet2). This network structure is similar to UNet1 and is used to correct small structural errors caused by numerical solution or discretization, outputting the final reconstructed mesh.

[0117] 3.2 Area-based Image Reconstruction (ImageDecoder): This module directly recovers the grayscale image from the geometry of the final grid by utilizing the physical correspondence between the area of ​​grid cells and image intensity. This module has no learnable parameters, ensuring that the reconstruction process follows the physical laws of quality transfer.

[0118] Using the above method, this embodiment implements a combined model of "deep network + explicit geometric operator + physical consistency decoding". Experiments show that this model can automatically learn the optimal Beltrami coefficient distribution through backpropagation, significantly improving inference speed compared to traditional numerical methods, and the generated geometric representation has better structure preservation and interpretability.

[0119] This application is described with reference to flowchart illustrations and / or block diagrams of geometric image representation methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.

[0120] The embodiments described above propose a geometric image representation (G-IR) method to transform traditional pixel-intensity-based image representations into geometric space representations with reversibility and structure preservation. This method is based on optimal transport and quasi-conformal mapping theory, and achieves a differentiable mapping from pixel space to geometric space by constructing a Beltrami coefficient field.

[0121] Example 6 This embodiment illustrates that the geometry-based representation method proposed in this invention is a general geometric representation computation framework, which is not limited to processing pixel intensity (grayscale values), but can also handle various derived geometric attributes or high-dimensional features of images. That is, the target measure can be extended to any scalar field function of the image.

[0122] 1. Definition of Generalized Objective Function: In step S1, the objective measure function of the grid is... Not only can it be derived from pixel grayscale The definition can also be defined as the geometric features of an image when it is a surface. For example: 1.1 Basic Geometric Features. We construct a 2.5D mesh with a geometric topological structure by mapping pixel positions to a domain and scalar values ​​(grayscale) to a height field. The image is viewed as a surface in three-dimensional space. Then the following geometric quantities can be calculated: 1.1.1 Gradient: Calculate the gradient (requires three channels) or magnitude of the scalar field of the image.

[0123] 1.1.2 Curvature feature extraction: Calculate the discrete Gaussian curvature or mean curvature of each vertex on the surface.

[0124] 1.1.3 Normal Vector: Calculate the normal vector for each vertex. Alternatively, a component of the normal vector can be chosen (e.g., ...). (Components) or their orientation information is mapped to a scalar field to describe the orientation changes of the surface.

[0125] 1.2 Texture and Statistical Features: Calculate the local texture attributes of the image, including the principal direction, size, contrast, and anisotropy of the texture. Normalize these statistical quantities and use them as a measure function so that the generated geometric representation can encode the material and microstructure information of the image.

[0126] 1.3 Spatial and Depth Features: When the input data contains depth information (such as RGB-D images), the depth value is directly used as the measurement function, so that the geometric deformation directly reflects the three-dimensional spatial structure of the scene.

[0127] 1.4 Semantic Abstraction: Extracting more abstract semantic information (such as object categories and scene segmentation). The topology-preserving property of geometric representation ensures that the inherent structural relationships of the image are not destroyed during semantic abstraction.

[0128] 1.5 Function Normalization: The calculated feature distribution is transformed to the [0,1] interval through linear or nonlinear mapping (similar to grayscale processing) to obtain the generalized feature function. For high-dimensional features such as RGB and normal vectors, it supports decomposing them into multiple independent component channels, which are then mapped to the corresponding generalized feature functions.

[0129] 2. General mapping process: Map the above feature functions... As a target measure, the generated OT grid and corresponding GIR no longer reflect the "brightness distribution", but rather the "structural feature curvature distribution" or multi-channel composite feature distribution of the image.

[0130] 3. Taking pixel grayscale value, discrete Gaussian curvature, and discrete mean curvature as examples. For example... Figure 7 As shown, for the same original input image ( Figure 7 a) This framework can extract its different physical or geometric features: photometric features: directly using pixel grayscale values ​​( Figure 7 b(1) and the corresponding G-IR7b(2)). Intrinsic geometric features: viewing the image as a three-dimensional surface. Calculate its discrete Gaussian curvature (e.g. Figure 7 c(1) and the corresponding G-IR7c(2)) or mean curvature (as shown in Figure 7d(1) and the corresponding G-IR7d(2)).

[0131] This invention provides a unified set of mathematical tools that can encode the multidimensional information of an image into topology-preserving geometric deformations, realizing a paradigm shift from "pixel intensity representation" to "generalized function geometric representation".

[0132] The aforementioned computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing equipment, causing it to execute the instructions to perform the functions described in the flowchart and / or block diagram.

[0133] The computer program instructions may also be stored in a computer-readable storage medium to generate means for performing the above functions when executed; or loaded onto a computer to complete the entire process of geometric image representation and reconstruction by performing a series of operational steps.

[0134] The contents of this specification are merely illustrative of the technical concept of this application and should not be construed as limiting the scope of protection of this application. Any equivalent substitutions or modifications made based on the technical concept and solution proposed in this application shall fall within the scope of protection of the claims of this application.

Claims

1. A geometric representation method for an image, characterized in that, include: The input image is processed into a standard grid and the source measure is defined. The attribute information of the image is then converted into the target measure function on the standard grid. Based on the source and target measure functions, the optimal transport map is obtained by solving the optimal transport solution model of the Fourier neural operator, and the standard grid is converted into the optimal transport grid according to the optimal transport map. The deformation field from the standard grid to the optimal transmission grid is calculated, and the deformation field is represented as a quasi-conformal mapping. The Beltrami coefficient field corresponding to the quasi-conformal mapping is calculated, and the Beltrami coefficient field constitutes the geometric representation of the image in different dimensions.

2. The geometric representation method of an image according to claim 1, characterized in that, The input image is processed into a standard grid, and a source measure is defined. The image's attribute information is then converted into a target measure function on the grid. Specifically: The input image is processed into a grid, and the feature points of the image's attribute information are set as grid vertices. A standard triangular grid is then constructed using these grid vertices. The uniform probability distribution of the standard triangular mesh is defined as the source measure; The attribute information of the input image is normalized and defined as a specific function in the image domain; Constructing a target measure function on a standard grid based on a specific function in the image domain; Before constructing the target measure function on the standard grid, a positive hyperparameter is added to process the zero-intensity region of the image; The image's attribute information includes at least one of the following: pixel intensity, gradient, curvature, texture, and depth.

3. The geometric representation method of an image according to claim 1, characterized in that, Based on the source and target measure functions, the optimal transport map is obtained through the optimal transport solution model of the Fourier neural operator. The standard mesh is then converted into the optimal transport mesh according to the optimal transport map, specifically including: Based on the source and target measure functions, the target measure function and the coordinates of the standard grid are concatenated in the channel dimension to form an input tensor. The tensor is then input into the optimal transport solution model of the pre-trained Fourier neural operator, which maps the tensor from the low-dimensional space to the high-dimensional latent feature space to obtain the optimal transport mapping. The optimal transport map is applied to each vertex of the standard grid, and an optimal transport grid is formed through an energy optimization method.

4. The geometric representation method of an image according to claim 1, characterized in that, The specific process for obtaining the Beltrami coefficient field is as follows: For each triangular facet in the standard grid, calculate the Beltrami coefficient of each triangular facet based on the corresponding vertex coordinates of the triangular facet in the standard grid and the optimal transport grid; For each vertex on the standard grid, the average of the Beltrami coefficients of all adjacent triangular faces of the vertex is selected as the Beltrami coefficient of that vertex; The Beltrami coefficients of each vertex on all standard grids are combined to form a Beltrami coefficient field in the image domain.

5. A method for image inverse transform restoration based on geometric representation, characterized in that, include, Based on the geometric representation of images in different dimensions, the corresponding quasi-conformal mapping is calculated by solving the linear Beltrami equation, and the standard mesh is transformed into a reconstructed quasi-conformal mesh through the quasi-conformal mapping. The local area of ​​each vertex in the reconstructed quasi-conformal mesh is calculated. Based on the local area, the attribute information of each vertex is calculated inversely through the inverse operation of the target measure function, and the reconstructed image is recovered.

6. An image processing method, characterized in that, The method is based on the geometric representation of the image as described in claim 1 or 5 and the image inverse transformation restoration method based on geometric representation, including: The geometric representations of the images in different dimensions are used as input to the image autoencoder model; In the latent space of the autoencoder model of the image, interpolation or operation is performed on the geometric representations corresponding to the image in different dimensions to obtain the interpolation or operation results; The interpolation or calculation results are decoded into an image output using the geometric representation-based inverse image transformation restoration method.

7. An image geometric representation system for performing the geometric representation method of an image according to any one of claims 1-4, characterized in that, include: The preprocessing module is used to process the input image into a standard grid and define the source measure, converting the image's attribute information into a target measure function on the grid; The mapping calculation module, based on the source and target measure functions, obtains the optimal transport mapping by solving the Fourier neural operator, and converts the standard grid into the optimal transport grid according to the optimal transport mapping; The geometric feature extraction module calculates the deformation field from the standard grid to the optimal transmission grid, represents the deformation field as a quasi-conformal mapping, and calculates the Beltrami coefficient field corresponding to the quasi-conformal mapping. The Beltrami coefficient field serves as the geometric representation of images in different dimensions.

8. An image restoration system for performing the image inverse transform restoration method based on geometric representation as described in claim 5, characterized in that, include: The mesh reconstruction module is used to calculate the corresponding quasi-conformal mapping by solving the linear Beltrami equation based on the geometric representation of the input image in different dimensions, and to transform the standard mesh into the reconstructed quasi-conformal mesh through the quasi-conformal mapping. The image restoration module calculates the local area of ​​each vertex in the reconstructed quasi-conformal mesh. Based on the local area, it calculates the attribute information features of each vertex through the inverse operation of the target measure function. The attribute information features at the vertex are then mapped to the image pixel grid to restore the reconstructed image.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the geometric representation method of the image as described in any one of claims 1-4.

10. An image processing device, characterized in that, include: One or more processors; Memory, which stores computer program instructions; When the computer program instructions are executed by the one or more processors, the device performs the geometric representation method of the image as described in any one of claims 1-4, or implements the image inverse transformation restoration method based on geometric representation as described in any one of claims 5, or runs the image processing method as described in claim 6.