Efficient image feature representation method and system based on two-dimensional Gaussian sputtering

By introducing two-dimensional Gaussian sputtering technology and Gaussian embedding module into the image word participle, the problem of limited discrete codebook space in the VQ method is solved, and more efficient image feature representation and reconstruction effect is achieved.

CN120125831APending Publication Date: 2025-06-10TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510044034.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, the discrete codebook space used by vector quantization (VQ) in image word segmentation is limited, limiting the modeling ability of the model in discrete space, resulting in a decrease in image reconstruction index and a lengthy training process.

Method used

The original image data is converted into characteristic two-dimensional Gaussian anchor points using two-dimensional Gaussian sputtering technology, and the attribute correction of these anchor points is performed through the Gaussian embedding module, and finally integrated into the image word participle to generate a reconstructed image.

Benefits of technology

By expanding the local modeling ability of the codebook space, the representation ability of the image word segmenter is improved, the quality of image reconstruction and the efficiency of the training process are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125831A_ABST
    Figure CN120125831A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient image feature representation method and system based on two-dimensional Gaussian sputtering, and the method comprises the steps: representing a coding sample as a plurality of two-dimensional gauss with flexible features of positions, rotation angles, scaling factors and feature coefficients, and then carrying out the quantification of the Gaussian features through a standard quantification method; and splicing a quantization result with other inherent parameters of Gaussian, and carrying out corresponding feature decoding and discrimination supervision of a reconstructed sample after sputtering operation. The invention presents a more flexible potential modeling strategy. Feature representation is determined by an original discrete codebook, and local feature attributes are adaptively learned through two-dimensional Gaussian distribution. Therefore, through continuous Gaussian distribution, diversified combinations are formed in the discrete space, so that the expression ability of the original discrete space is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and machine learning, and particularly to an efficient image feature representation method and system based on two-dimensional Gaussian sputtering. Background Art

[0002] Large language models (LLMs) have recently demonstrated their superior model capacity and scalability in natural language tasks and have taken a dominant position in these tasks. In addition, a series of vision and multimodal efforts have attempted to leverage the autoregressive architecture and pre-trained knowledge of LLMs to solve vision-related tasks. To adapt to the discrete input format of LLMs, images are first tokenized to obtain discrete visual tokens, and then text alignment and subsequent autoregressive predictions are performed according to the task format. Therefore, the representation ability of the image tokenizer directly determines the upper limit of the model's ability.

[0003] Vector quantization (VQ) is a currently popular image tokenization technique that has been widely applied in image perception, conditional image generation, and multimodal image understanding tasks. Specifically, the VQ-based strategy maintains a discrete codebook that contains a certain number of learnable vectors. The encoded image features are aligned with the codebook vectors through similarity calculation for nearest neighbor matching, enabling the image to be represented by discrete tokens generated by the codebook vectors. Then, a decoding module processes these discrete tokens to produce a reconstruction result in the RGB domain. In addition, researchers have introduced a discriminator module to impose constraints related to the generative adversarial network (GAN) on the reconstructed image to enhance the authenticity and visual perception effect of the image. However, compared with the continuous latent space of the vanilla variational autoencoder (VAE), the size of the codebook space greatly limits the ability to model distributions in the discrete space. Therefore, this limitation leads to a decline in image reconstruction metrics and causes a lengthy training process. Some methods may expand the discrete space by increasing the number of codebooks or constructing a lookup-free mapping, but they are still essentially limited to a finite discrete space and require an evolutionary training process to achieve model convergence. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems in the related art to some extent.

[0005] The present invention proposes an efficient image feature representation method based on two-dimensional Gaussian sputtering, which adopts two-dimensional Gaussian sputtering technology to enrich the codebook space to improve the modeling ability.

[0006] Another object of the present invention is to propose an efficient image feature representation system based on two-dimensional Gaussian sputtering.

[0007] To achieve the above object, on the one hand, the present invention proposes an efficient image feature representation method based on two-dimensional Gaussian sputtering, including:

[0008] Two-dimensional Gaussian quantization: Convert the original image data from the RGB domain to the feature two-dimensional Gaussian anchor points;

[0009] Gaussian embedding: Input the original image features of the original image data extracted by the encoder into the feature booster to flatten them into a feature sequence, and use the Gaussian booster to initialize the attributes of the feature two-dimensional Gaussian anchor points and related high-dimensional queries to output a unified representation of two entities with an embedding dimension; Perform attention processing based on the unified representation to obtain a comprehensive feature representation after self-attention and cross-attention processing, and perform attribute correction based on the comprehensive feature representation and the query results of cross-attention to finally obtain the corrected anchor point attributes;

[0010] Two-dimensional Gaussian sputtering: Integrate the models of two-dimensional Gaussian quantization and Gaussian embedding into an image tokenizer based on two-dimensional Gaussian sputtering to generate the corresponding reconstructed image.

[0011] The efficient image feature representation method based on two-dimensional Gaussian sputtering in the embodiments of the present invention may further have the following additional technical features:

[0012] In one embodiment of the present invention, the original quantization vector z i is extended to a feature two-dimensional Gaussian quantization unit g that describes local features within a specific region k , and all quantization units g k are aggregated at the position p i =(x,y) to obtain the contribution c ki :

[0013]

[0014] where K represents the number of Gaussian units; each feature two-dimensional Gaussian unit is described by its position covariance matrix and additional feature coefficients ; the contribution c ki is represented by the probability π of the two-dimensional Gaussian distribution ki as follows:

[0015]

[0016] Use the product of the rotation matrix and the scaling matrix to refine the factorization representation of the covariance matrix:

[0017] ∑=(RS)(RS) T ,

[0018] where R and S are determined by the rotation angle θ∈[0,π] and the scaling factor Exported as follows:

[0019]

[0020] Convert the original image data from the RGB domain to a set of 2D Gaussian cells:

[0021]

[0022] where d represents its total parameter dimension, including 2D position parameters, 2D scaling parameters, 1D rotation angle parameters, and D-dimensional feature coefficients.

[0023] In an embodiment of the present invention, the original image features of the original image data extracted by the encoder are input into the feature booster to be flattened into a feature sequence, and the Gaussian booster is used to initialize the attributes of the feature two-dimensional Gaussian anchors and related high-dimensional queries, so as to output a unified representation of two entities with an embedding dimension, including:

[0024] Input the original image data into the encoder ε for feature extraction to output a feature map

[0025] The feature map Input into the feature booster to be flattened into a feature sequence

[0026] Use the Gaussian booster to initialize the attributes of the feature two-dimensional Gaussian anchor G and its related high-dimensional query Q;

[0027] Obtain the embedded features through a multi-layer perceptron To obtain a unified representation of two entities with a D' embedding dimension.

[0028] In an embodiment of the present invention, for the two-dimensional Gaussian anchor g k , generate a series of offsets Δμ ki , by combining Δμ ki with the position of the anchor μ k to obtain a series of reference points R is a predefined parameter; calculate the corresponding attention weight A ki , for weighted summation values:

[0029]

[0030] where x represents the autoencoded image features, W represents the linear transformation matrix that projects them onto the values, and the scalar attention weight A ki is in the range [0,1], and is normalized by ; Δμ ki and A kiAll are obtained through the anchor point g k The corresponding query feature q k by performing linear projection;

[0031] The attribute adjustment amount Δg of the anchor point g ∈ G is decoded from the query result of the previous cross-attention module through a multi-layer perceptron:

[0032]

[0033] Directly replace the variables in g that need to be quickly adjusted with Δg, including the rotation angle θ, the scaling factor s, and the feature coefficient ζ, excluding the position μ; refine μ by adding a residual adjustment amount Δμ to the anchor point μ:

[0034] g new = {μ + Δμ, Δθ, Δs, Δζ}.

[0035] In an embodiment of the present invention, the method further includes imposing a quality constraint on the reconstructed image through an additional discriminator, and the constraint includes a reconstruction loss a commitment loss and an additional GAN loss The formula is as follows:

[0036]

[0037] where a and β are hyperparameters for balancing the three losses.

[0038] To achieve the above object, on the other hand, the present invention proposes an efficient image feature representation system based on two-dimensional Gaussian sputtering, including:

[0039] A two-dimensional Gaussian quantization module for converting the original image data from the RGB domain into feature two-dimensional Gaussian anchor points;

[0040] A Gaussian embedding module for inputting the original image features of the original image data extracted by the encoder into a feature booster to flatten them into a feature sequence, and using a Gaussian booster to initialize the attributes of the feature two-dimensional Gaussian anchor points and related high-dimensional queries, so as to output a unified representation of two entities with an embedding dimension; performing attention processing based on the unified representation to obtain a comprehensive feature representation after self-attention and cross-attention processing, and performing attribute correction based on the comprehensive feature representation and the query result of cross-attention to finally obtain the corrected anchor point attributes;

[0041] A two-dimensional Gaussian sputtering module for integrating the two-dimensional Gaussian quantization module and the Gaussian embedding module into an image tokenizer based on two-dimensional Gaussian sputtering to generate a corresponding reconstructed image.

[0042] ​The efficient image feature representation method and system based on two-dimensional Gaussian sputtering according to the embodiments of the present invention are used to parameterize encoded image features and convert them into multiple Gaussian distributions. Each Gaussian distribution is determined by its position, rotation angle, scaling factor, and feature coefficient, and these parameters can undergo an adaptive learning process. Subsequently, the present invention proposes to use a standard discrete codebook to quantify the feature coefficients through nearest neighbor matching and splice the other parameters (position, rotation angle, and scaling factor) of the Gaussian distribution with the quantization result. The present invention then employs a two-dimensional sputtering module to project these combined Gaussian parameters back into the image feature space. The final steps include a feature decoder for reconstructing the original image and a discriminator for further optimizing the image quality. Compared with traditional VQ-based methods, the present invention presents a more flexible latent modeling strategy. The feature representation is determined by the original discrete codebook, while local feature attributes (such as position) are adaptively learned through two-dimensional Gaussian distributions. Therefore, the present invention constitutes diverse combinations within the discrete space through continuous Gaussian distributions, thereby expanding the representation ability of the original discrete space.

[0043] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings

[0044] The above-mentioned and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:

[0045] Figure 1 is a flowchart of the efficient image feature representation method based on two-dimensional Gaussian sputtering according to the embodiments of the present invention;

[0046] Figure 2 is an overall schematic diagram of the efficient image feature representation method based on two-dimensional Gaussian sputtering according to the embodiments of the present invention;

[0047] Figure 3 is a schematic diagram of Gaussian embedding according to the embodiments of the present invention;

[0048] Figure 4 is a comparison schematic diagram between the present invention and other methods according to the embodiments of the present invention;

[0049] Figure 5 is a structural diagram of the efficient image feature representation system based on two-dimensional Gaussian sputtering according to the embodiments of the present invention. Detailed Description of the Embodiments

[0050] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0051] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the scope of protection of the present invention.

[0052] The following describes a high - efficiency image feature representation method and system based on two - dimensional Gaussian sputtering according to an embodiment of the present invention with reference to the accompanying drawings.

[0053] Figure 1 It is a flowchart of a high - efficiency image feature representation method based on two - dimensional Gaussian sputtering according to an embodiment of the present invention. As Figure 1 shown, the method includes:

[0054] S1, two - dimensional Gaussian quantization: converting the original image data from the RGB domain to feature two - dimensional Gaussian anchor points;

[0055] S2, Gaussian embedding: inputting the original image features of the original image data extracted by the encoder into the feature booster to flatten them into a feature sequence, and using the Gaussian booster to initialize the attributes of the feature two - dimensional Gaussian anchor points and related high - dimensional queries to output a unified representation of two entities with an embedding dimension; performing attention processing based on the unified representation to obtain a comprehensive feature representation after self - attention and cross - attention processing, and performing attribute correction based on the comprehensive feature representation and the query result of cross - attention to finally obtain the corrected anchor point attributes;

[0056] S3, two - dimensional Gaussian sputtering: integrating the two - dimensional Gaussian quantization model and the Gaussian embedding model into an image tokenizer based on two - dimensional Gaussian sputtering to generate a corresponding reconstructed image.

[0057] It is understandable that the present invention solves a series of problems existing in the methods of projecting pixels onto a discrete codebook using vector quantization (VQ) and reconstructing an image from the discrete representation. Compared with the continuous latent space, the limited discrete codebook space greatly restricts the representation ability of these image tokenizers. The present invention integrates the local modeling ability of the two-dimensional Gaussian distribution into the discrete space, thereby enhancing the representation ability of the image tokenizer. The encoded samples are represented as two-dimensional Gaussians with multiple flexible features including position, rotation angle, scaling factor, and feature coefficients, and then the standard quantization method is used to quantize the Gaussian features; the quantization results are concatenated with other inherent parameters of the Gaussian, and corresponding feature decoding and discriminative supervision for reconstructing the samples are performed after the sputtering operation; the process of the present invention is verified for the reconstruction performance on datasets such as CIFAR, Mini-ImageNet, and ImageNet-1K to comprehensively measure the application value of the present invention, and its overall process is as Figure 2 shown.

[0058] In one embodiment of the present invention, the encoded samples are represented as two-dimensional Gaussians with multiple flexible features including position, rotation angle, scaling factor, and feature coefficients, and then the standard quantization method is used to quantize the Gaussian features.

[0059] Specifically, the present invention proposes a new feature quantization paradigm, introducing the concept of the two-dimensional Gaussian distribution and expanding the original quantization vector z i , which only contains individual features, and expanding it into a feature two-dimensional Gaussian quantization unit g k that describes the local features within a specific region, rather than a fixed grid. The present invention does not directly replace the quantization of the feature vector i with the corresponding quantization vector z , but aggregates the contributions c k of all quantization units g i at the position p ki =(x,y):

[0060]

[0061] where K represents the number of Gaussian units. Specifically, each feature two-dimensional Gaussian unit is described by its position covariance matrix and additional feature coefficients . Therefore, the contribution c ki can be represented by the probability π ki of the two-dimensional Gaussian distribution as follows:

[0062]

[0063] Since the covariance matrix Σ of the Gaussian distribution must be positive definite, it is necessary to ensure the legality of the matrix values during the numerical optimization process. Therefore, the present invention selects to use the rotation matrix and the scaling matrix product to refine the factorization representation of the covariance matrix:

[0064] ∑=(RS)(RS) T ,

[0065] where R and S are derived from the rotation angle θ∈[0,π] and the scaling factor as follows:

[0066]

[0067] The present invention converts the image data from the RGB domain into a set of 2D Gaussian cells of the above features:

[0068]

[0069] where d represents its total parameter dimension (including 2D position parameters, 2D scaling parameters, 1D rotation angle parameter, and D-dimensional feature coefficients). The present invention first quantizes the feature coefficients ζ k in accordance with the method of the traditional VQ-VAE while maintaining the continuity of their positions and covariance matrices. In this case, the model can optimize the positions and scalings of the units in any feature map, which enables the two-dimensional Gaussian to adaptively allocate computational and storage resources according to the regional complexity. Subsequently, the quantization units are aggregated into the quantization feature map and super-fast rendering is performed through the two-dimensional Gaussian sputtering function efficiently implemented by CUDA.

[0070] In an embodiment of the present invention, based on the above two-dimensional Gaussian quantization, the present invention further proposes a Gaussian embedding module for learning meaningful Gaussian representations using image features. The process of the present invention starts from two main information carriers: the original image features and the two-dimensional Gaussian objects. Essentially, the present invention takes the Gaussian embedding module as a channel for information exchange and continuously refines the learned attributes through a series of operations. Below, the present invention will deeply explore the complexity and basis of the related operations. As Figure 3 shown.

[0071] Specifically, the image features are processed using the Gaussian embedding module to obtain the inherent spatial parameters (position, rotation angle, and scaling factor) and the feature coefficients. The present invention performs nearest neighbor matching on the feature coefficients to implement the quantization process and splices the quantization results with the spatial parameters.

[0072] Lifter Module: As a preparatory step for subsequent modules, the Lifter Module converts two main information carriers into a unified vector. To adapt to the attention architecture, the feature lifter flattens the feature map from the encoder ε into a feature sequence It should be noted that D′ is much larger than the channel dimension D of the required quantized feature map Z to retain more image information for Gaussian representation learning. In addition, the present invention is concatenated with the cosine position embedding sequence, enabling the model to have the ability to distinguish order. On the other hand, the Gaussian lifter initializes the attributes of the two-dimensional Gaussian anchor G and its associated high-dimensional query Q. Since each anchor g k is refined in the form of target Gaussian parameters, we maintain a multi-layer perceptron (MLP) to obtain the embedded features to ensure seamless interaction with the query. Finally, the present invention obtains a unified representation of two entities with an embedding dimension of D′.

[0073] Self-attention: The present invention uses a self-attention layer on the visual feature sequence and the query Q to further compress the image information and fuse the interactions between the two-dimensional Gaussian anchors. The present invention replaces the general attention form with differentiable attention (DA) when processing the visual feature sequence, alleviating the challenges of high computational complexity and dealing with high-resolution features.

[0074] Cross-attention: The randomly initialized two-dimensional Gaussian object extracts visual information in the cross-attention module, which is also based on differentiable attention (DA). Specifically, for the two-dimensional Gaussian anchor g k , the present invention generates a series of offsets Δμ ki . The present invention further combines Δμ ki with the position of the anchor μ k to obtain a series of reference points where R is a predefined parameter. Then the present invention calculates the corresponding attention weights A ki , which are subsequently used for weighted summation:

[0075]

[0076] where x represents the autoencoded image features and W represents the linear transformation matrix that projects them onto the values. The scalar attention weights A ki are in the range [0,1] and are normalized by . Both Δμ ki and A ki are obtained through the query feature q k corresponding to the anchor g kObtained by performing linear projection. In practical applications, the present invention embeds the features of the anchor points before module input with the query feature q k merged, which enhances the connection between them.

[0077] Refinement: The present invention uses a refinement module to correct the attributes of the anchor point G, which are guided by the query results of the previous cross-attention module Specifically, the present invention first decodes the attribute adjustment amount Δg of the anchor point g∈G from through a multi-layer perceptron:

[0078]

[0079] In a specific adjustment strategy, the present invention adopts a more reasonable method, directly replacing the variables in g that need to be quickly adjusted with Δg, including the rotation angle θ, the scaling factor s, and the feature coefficient ζ, excluding the position μ. Considering that μ determines the region in the feature map affected by the anchor point, frequently replacing μ may disrupt the optimization of all other attributes, resulting in instability during the training process. Instead, the present invention refines μ by adding a residual adjustment amount Δμ to the anchor point μ:

[0080] g new ={μ + Δμ, Δθ, Δs, Δζ}

[0081] In an embodiment of the present invention, the present invention employs a two-dimensional sputtering module and an image decoder to generate a corresponding reconstructed image. Finally, the present invention imposes a quality constraint on the reconstructed image through an additional discriminator. The overall constraints include an image reconstruction loss, a VQ loss related to quantization and a commitment loss, as well as a GAN loss for image quality.

[0082] Specifically, the preamble steps of the present invention can be easily integrated into an existing visual tokenizer with only minor modifications to the quantization process. To benchmark against the state-of-the-art methods, the present invention deploys the method by simply inserting a Gaussian embedding module into VQGAN. The overall loss of the model includes a reconstruction loss commitment loss and an additional GAN loss The formula is as follows:

[0083]

[0084] where a and β are hyperparameters for balancing the three losses.

[0085] Appendix Figure 4This is a comparison schematic diagram between the present invention and other methods. The variational autoencoder (VAE) uses a continuous space to represent image features, which cannot be aligned with discrete modalities. In addition, the vector quantization (VQ) method uses a discrete codebook for nearest neighbor matching of image features, and the representation space is limited within the codebook size. In contrast, the present invention performs local adaptive learning on discrete features based on discretization, using continuous Gaussian parameters, thereby expanding the representation space.

[0086] According to the efficient image feature representation method based on two-dimensional Gaussian sputtering of an embodiment of the present invention, the local quantization feature range is adaptively learned through Gaussian position, rotation angle, and scaling factor, thereby enhancing the performance ability of the latent discrete space. The present invention further renders all Gaussian parameters and obtains the final reconstruction result through an image decoder, and the overall process verifies the effectiveness in the image reconstruction task.

[0087] To implement the above embodiment, as Figure 5 shown, the present embodiment also provides an efficient image feature representation system 10 based on two-dimensional Gaussian sputtering, including:

[0088] A two-dimensional Gaussian quantization module 100, configured to convert the original image data from the RGB domain into feature two-dimensional Gaussian anchor points;

[0089] A Gaussian embedding module 200, configured to input the original image features of the original image data extracted by an encoder into a feature booster to flatten them into a feature sequence, and use a Gaussian booster to initialize the attributes of the feature two-dimensional Gaussian anchor points and related high-dimensional queries, so as to output a unified representation of two entities with an embedding dimension; perform attention processing based on the unified representation to obtain a comprehensive feature representation after self-attention and cross-attention processing, and perform attribute correction based on the comprehensive feature representation and the query result of cross-attention to finally obtain the corrected anchor point attributes;

[0090] A two-dimensional Gaussian sputtering module 300, configured to integrate the two-dimensional Gaussian quantization module and the Gaussian embedding module into an image tokenizer based on two-dimensional Gaussian sputtering to generate a corresponding reconstructed image.

[0091] Further, the two-dimensional Gaussian quantization module is further configured to expand the original quantization vector z i into a feature two-dimensional Gaussian quantization unit g that describes local features within a specific region k , and aggregate all quantization units g k at the position p i =(x, y) contribution c ki :

[0092]

[0093] where K represents the number of Gaussian units; each feature two-dimensional Gaussian unit is determined by its position covariance matrix and an additional feature coefficient to describe; the contribution c ki is represented by the probability π of the two-dimensional Gaussian distribution ki as follows:

[0094]

[0095] Use the product of the rotation matrix and the scaling matrix to refine the factorization representation of the covariance matrix:

[0096] ∑=(RS)(RS) T ,

[0097] where R and S are derived from the rotation angle θ∈[0,π] and the scaling factor as follows:

[0098]

[0099] Convert the original image data from the RGB domain to a set of 2D Gaussian units:

[0100]

[0101] where d represents its total parameter dimension, including 2D position parameters, 2D scaling parameters, 1D rotation angle parameters, and D-dimensional feature coefficients.

[0102] Furthermore, the Gaussian embedding module 200 is also used for:

[0103] Input the original image data into the encoder ε for feature extraction and output the feature map

[0104] Input the feature map into the feature booster to flatten it into a feature sequence

[0105] Use the Gaussian booster to initialize the attributes of the feature two-dimensional Gaussian anchor G and its associated high-dimensional query Q;

[0106] Obtain the embedded features through a multi-layer perceptron to obtain a unified representation of two entities with an embedded dimension of D'.

[0107] Furthermore, for the two-dimensional Gaussian anchor g k , generate a series of offsets Δμ ki , by adding Δμ ki to the anchor μ kCombined with the positions, a series of reference points are obtained R is a predefined parameter; calculate the corresponding attention weight A ki , for weighted summation value:

[0108]

[0109] where x represents the image features of the autoencoder, W represents the linear transformation matrix that projects them onto the values, and the scalar attention weight A ki is within the range [0, 1], and is normalized through ; Δμ ki and A ki are both obtained by linearly projecting the query feature q k corresponding to the anchor point g k ;

[0110] Decode the attribute adjustment amount Δg of the anchor point g ∈ G from the query result of the previous cross-attention module through a multi-layer perceptron:

[0111]

[0112] Directly replace the variables in g that need to be quickly adjusted with Δg, including the rotation angle θ, the scaling factor s, and the feature coefficient ζ, excluding the position μ; refine μ by adding a residual adjustment amount Δμ to the anchor point μ:

[0113] g new = {μ + Δμ, Δθ, Δs, Δζ}.

[0114] Furthermore, it also includes imposing quality constraints on the reconstructed image through an additional discriminator, and the constraints include the reconstruction loss commitment loss and the additional GAN loss The formulas are as follows:

[0115]

[0116] According to the efficient image feature representation system based on two-dimensional Gaussian sputtering of the embodiments of the present invention, the local quantization feature range is adaptively learned through Gaussian positions, rotation angles, and scaling factors, thereby enhancing the performance of the latent discrete space. The present invention further renders all Gaussian parameters and obtains the final reconstruction result through an image decoder, and the overall process verifies the effectiveness in the image reconstruction task.

[0117] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0118] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

Claims

1. An efficient image feature representation method based on two-dimensional Gaussian sputtering, characterized in that: include: 2D Gaussian quantization: convert the original image data from RGB domain to feature 2D Gaussian anchor points; Gaussian embedding: The original image features of the original image data extracted by the encoder are input into the feature booster to be flattened into a feature sequence, and the Gaussian booster is used to initialize the attributes of the feature two-dimensional Gaussian anchor point and the related high-dimensional query to output a unified representation of two entities with an embedded dimension; attention processing is performed based on the unified representation to obtain a comprehensive feature representation after self-attention and cross-attention processing, and attribute correction is performed based on the query results of the comprehensive feature representation and cross-attention to finally obtain the corrected anchor point attributes; 2D Gaussian Sputtering: The 2D Gaussian quantization model and the Gaussian embedding model are integrated into the 2D Gaussian sputtering based image segmentor to generate the corresponding reconstructed image.

2. The method according to claim 1, characterized in that: The original quantized vector z i Expanded to a characteristic two-dimensional Gaussian quantization unit g that describes the local features in a specific area k , gather all quantized units g k At position p i = contribution c of (x,y) ki : Where K represents the number of Gaussian units; each feature two-dimensional Gaussian unit is represented by its position Covariance matrix and additional characteristic coefficients To describe; contribution c ki Using the probability π of a two-dimensional Gaussian distribution ki It is expressed as follows: Using the rotation matrix and the scaling matrix The factorization of the covariance matrix is ​​refined by multiplying ∑=(RS)(RS) T , where R and S are determined by the rotation angle θ∈[0,π] and the scaling factor Export as follows: Convert the raw image data from the RGB domain to a set of 2D Gaussian cells: Where d represents the total parameter dimension, including 2-dimensional position parameters, 2-dimensional scaling parameters, 1-dimensional rotation angle parameters and D-dimensional feature coefficients.

3. The method according to claim 2, characterized in that The original image features of the original image data extracted by the encoder are input into the feature booster to be flattened into a feature sequence, and the Gaussian booster is used to initialize the attributes of the feature two-dimensional Gaussian anchor and the related high-dimensional query to output a unified representation of two entities with embedded dimensions, including: The original image data is input into the encoder ε for feature extraction and output feature map The feature map The input feature booster is flattened into a feature sequence Use Gaussian booster to initialize the attributes of feature two-dimensional Gaussian anchor G and its related high-dimensional query Q; Obtaining embedded features through multi-layer perceptron To obtain a D ′ Unified representation of two entities with embedding dimensions.

4. The method according to claim 3, characterized in that For the two-dimensional Gaussian anchor point g k , generating a series of offsets Δμ ki , by setting Δμ ki With anchor point μ k Combined with the positions of R is a predefined parameter; calculate the corresponding attention weight A ki , for weighted sum values: where x represents the self-encoded image features, W represents the linear transformation matrix that projects them onto the values, and A is the scalar attention weight. ki In the range [0,1], through Normalized; Δμ ki and A ki All through the anchor point g k The corresponding query feature q k Obtained by linear projection; The query results from the previous cross-attention module are obtained through a multi-layer perceptron. The attribute adjustment Δg of the anchor point g∈G is decoded: Directly replace the variables in g that need to be adjusted quickly with Δg, including the rotation angle θ, the scaling factor s, and the characteristic coefficient ζ, excluding the position μ; refine μ by adding a residual adjustment Δμ to the anchor point μ: g new = {μ+Δμ,Δθ,Δs,Δζ}.

5. The method according to claim 1, characterized in that The method further includes imposing a quality constraint on the reconstructed image through an additional discriminator, wherein the constraint includes a reconstruction loss commitment loss and additional GAN ​​loss The formula is as follows: Among them, α and β are hyperparameters used to balance the three losses.

6. An efficient image feature representation system based on two-dimensional Gaussian sputtering, characterized in that: include: A two-dimensional Gaussian quantization module is used to convert the original image data from the RGB domain into characteristic two-dimensional Gaussian anchor points; A Gaussian embedding module is used to input the original image features of the original image data extracted by the encoder into the feature booster to flatten them into a feature sequence, and use the Gaussian booster to initialize the attributes of the feature two-dimensional Gaussian anchor point and the related high-dimensional query to output a unified representation of two entities with embedded dimensions; perform attention processing based on the unified representation to obtain a comprehensive feature representation after self-attention and cross-attention processing, and perform attribute correction based on the query results of the comprehensive feature representation and the cross-attention to finally obtain the corrected anchor point attributes; The two-dimensional Gaussian sputtering module is used to integrate the two-dimensional Gaussian quantization module and the Gaussian embedding module into the image segmenter based on two-dimensional Gaussian sputtering to generate a corresponding reconstructed image.

7. The system according to claim 6, characterized in that The two-dimensional Gaussian quantization module is also used to convert the original quantized vector z i Expanded to a characteristic two-dimensional Gaussian quantization unit g that describes the local features in a specific area k , gather all quantized units g k At position p i = contribution c of (x,y) ki : Where K represents the number of Gaussian units; each feature two-dimensional Gaussian unit is represented by its position Covariance matrix and additional characteristic coefficients To describe; contribution c ki Using the probability π of a two-dimensional Gaussian distribution ki It is expressed as follows: Using the rotation matrix and the scaling matrix The factorization of the covariance matrix is ​​refined by multiplying ∑=(RS)(RS) T , where R and A are determined by the rotation angle θ∈[0,π] and the scaling factor Export as follows: Convert the raw image data from the RGB domain to a set of 2D Gaussian cells: Where d represents the total parameter dimension, including 2-dimensional position parameters, 2-dimensional scaling parameters, 1-dimensional rotation angle parameters and D-dimensional feature coefficients.

8. The system according to claim 6, characterized in that Gaussian Embedding module, also used for: The original image data is input into the encoder ε for feature extraction and output feature map The feature map The input feature booster is flattened into a feature sequence Use Gaussian booster to initialize the attributes of feature two-dimensional Gaussian anchor G and its related high-dimensional query Q; Obtaining embedded features through multi-layer perceptron To obtain a unified representation of the two entities with D′ embedding dimension.

9. The system according to claim 8, characterized in that For the two-dimensional Gaussian anchor point g k , generating a series of offsets Δμ ki , by setting Δμ ki With anchor point μ k Combined with the positions of R is a predefined parameter; calculate the corresponding attention weight A ki , for weighted sum values: where x represents the self-encoded image features, W represents the linear transformation matrix that projects them onto the values, and A is the scalar attention weight. ki In the range [0,1], through Normalized; Δμ ki and A ki All through the anchor point g k The corresponding query feature q k Obtained by linear projection; The query results from the previous cross-attention module are obtained through a multi-layer perceptron. The attribute adjustment Δg of the anchor point g∈G is decoded: Directly replace the variables in g that need to be adjusted quickly with Δg, including the rotation angle θ, the scaling factor s, and the characteristic coefficient ζ, excluding the position μ; refine μ by adding a residual adjustment Δμ to the anchor point μ: g new ={μ+Δμ,Δθ,Δs,Δζ} 10. The system according to claim 6, characterized in that It also includes imposing quality constraints on the reconstructed image through an additional discriminator, including reconstruction loss commitment loss and additional GAN ​​loss The formula is as follows: Among them, α and β are hyperparameters used to balance the three losses.