Global structure sensing vector quantization method based on optimal transmission

By decomposing the vector quantization framework into the optimal transmission problem and optimizing the allocation matrix using the Sinkhorn-Knopp algorithm, the problems of index crash, local optimization and data range sensitivity in the traditional vector quantization method are solved, and more stable and efficient training and higher quality data reconstruction are achieved.

CN120123628APending Publication Date: 2025-06-10TSINGHUA UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510090361.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional vector quantization methods are prone to index crashes, local minimum value problems, sensitivity to data range and training complexity during training, which limits its application potential in discrete compression scenarios of high-dimensional data such as images and videos.

Method used

A global structure-aware vector quantization method based on optimal transmission is proposed. The input data is mapped to the continuous hidden variable space through an encoder, and the distance matrix between the continuous characterization and the codebook is constructed, and the allocation matrix is ​​solved using the optimal transmission problem. The Sinkhorn-Knopp algorithm is used to perform row-to-square normalization iteration until the allocation matrix converges, thereby determining the quantization characteristics and reconstructing the data.

Benefits of technology

It significantly enhances the stability and efficiency of training, solves the problems of index crashes and local optimization, improves the compactness and accuracy of codebook utilization and data representation, and improves the fidelity of reconstructed data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123628A_ABST
    Figure CN120123628A_ABST
Patent Text Reader

Abstract

The invention provides a global structure sensing vector quantization method based on optimal transmission, which comprises the following steps of: mapping input data to a continuous hidden variable space through an encoder to generate continuous representation; constructing a distance matrix between the continuous representation and the codebook, and solving a distribution matrix by using an optimal transmission problem; carrying out normalization processing on the distance matrix, and generating an initial value of a distribution matrix by adopting a special initialization method; performing row and column normalization iteration on the distribution matrix through a Sinkhorn-Knopp algorithm until the distribution matrix is converged; and after iteration is finished, quantitative characteristics are determined based on the maximum value of the distribution matrix, the quantitative characteristics are mapped back to an original data space through a decoder, and a reconstructed data result is obtained. Based on the scheme provided by the invention, a codebook utilization rate close to 100% and an excellent data reconstruction result can be ensured, and a powerful tool is provided for discrete compression scene reconstruction of continuous data such as images and videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of discrete compression and reconstruction of image data, and particularly to a global structure-aware vector quantization method based on optimal transport. Background Art

[0002] Vector Quantization Networks (VQNs) are an important technology for mapping continuous-space data to discrete tokens. Initially mainly used for image distribution modeling, by mapping input data to a continuous latent space and quantizing it using a discrete codebook, an efficient representation of the data is achieved. However, traditional vector quantization methods usually calculate the distance between data features and the codebook using the Euclidean distance, and select the closest codebook token as the discrete representation through the nearest neighbor strategy. Due to the greedy nature of the nearest neighbor operation, the phenomenon of "index collapse" is prone to occur during the training process, that is, only a few codebooks are overused, while most codebooks are not utilized, resulting in low codebook utilization and limited model expression ability.

[0003] Traditional vector quantization methods also face the problem of local minima during training. Since the nearest neighbor operation cannot backpropagate, VQNs usually approximate the gradient of data features by gradient copying. This direct gradient estimation method is prone to falling into local optima, affecting the training effect. In the prior art, to solve these problems, researchers have proposed various methods, such as alleviating the index collapse problem through elaborate initialization and low-dimensional codebook design, or optimizing the training through a distillation strategy. However, these methods rely on complex technical implementations, pose high requirements on model design and training processes, and still cannot fundamentally solve the local optimum problem in vector quantization.

[0004] In addition, when traditional vector quantization networks process large-scale data or complex data modal distributions, their local optimization strategies usually have difficulty capturing the global structure information of the data. This limitation further affects the training stability and efficiency of the model. Conventional VQNs are similar to parameterized online K-Means algorithms, and the K-Means algorithm itself also has the inherent problem of local convergence, making it difficult to effectively utilize the global distribution characteristics of the data. Therefore, vector quantization networks perform limitedly in tasks that require efficient compression and reconstruction of data.

[0005] The vector quantization methods in the prior art are also very sensitive to the distribution range of the data. Since the value ranges of different data modalities vary greatly, this difference will directly affect the efficiency and accuracy of the quantization process. Therefore, the prior art requires complex preprocessing or data normalization schemes to alleviate this problem, but the effects and generality of these schemes are often limited, increasing the complexity and maintenance cost of the application.

[0006] In summary, traditional vector quantization techniques have many deficiencies in terms of index collapse, local optimal traps, sensitivity to data ranges, and training complexity. These drawbacks significantly limit their application potential in discrete compression scenarios of high-dimensional data such as images and videos. Summary of the Invention

[0007] This application aims to solve at least one of the technical problems in the related art to some extent.

[0008] To this end, the first objective of this application is to propose a global structure-aware vector quantization method based on optimal transport.

[0009] The second objective of this application is to propose a global structure-aware vector quantization device based on optimal transport.

[0010] The third objective of this application is to propose an electronic device.

[0011] The fourth objective of this application is to propose a computer-readable storage medium.

[0012] The fifth objective of this application is to propose a computer program product.

[0013] To achieve the above objectives, the first aspect embodiment of this application proposes a global structure-aware vector quantization method based on optimal transport, including:

[0014] Mapping the input data to a continuous latent variable space through an encoder to generate a continuous representation;

[0015] Constructing a distance matrix between the continuous representation and the codebook, and using the optimal transport problem to solve for the assignment matrix;

[0016] Normalizing the distance matrix and generating an initial value of the assignment matrix using a special initialization method;

[0017] Performing row and column normalization iterations on the assignment matrix through the Sinkhorn-Knopp algorithm until the assignment matrix converges;

[0018] After the iteration ends, determining the quantization feature based on the maximum value of the assignment matrix, and mapping the quantization feature back to the original data space through a decoder to obtain the reconstructed data result.

[0019] Optionally, the mapping the input data to a continuous latent variable space through an encoder to generate a continuous representation includes:

[0020] Representing the original input data as multiple data samples Using the encoder f to map each data sample x to a representation in the continuous latent variable space to obtain the continuous representation Z e= f(x) ∈ R m×s Denote each vector element of the continuous representation as z e ∈ R d where m is the number of samples and d represents the dimension of each vector.

[0021] Optionally, constructing the distance matrix between the continuous representation and the codebook and using the optimal transport problem to solve for the assignment matrix includes:

[0022] Construct the distance matrix D using the Euclidean distance, where D ij = ‖z i - c j ‖, z i represents each vector in the continuous representation, and c j represents the codebook vector in the codebook c = {c 1 , …, c n};

[0023] Establish an optimization model for the assignment matrix A through the objective function of the optimal transport problem. The form of the optimal transport problem is as follows:

[0024]

[0025] where H(A) = -∑ ij A ij logA ij represents the entropy function, ∈ is a hyperparameter, 1 r and 1 c are all - ones vectors of length r and c respectively, l represents the number of features, and n represents the number of codebooks.

[0026] Optionally, normalizing the distance matrix and generating an initial value of the assignment matrix using a special initialization method includes:

[0027] Perform mean normalization on the distance matrix D column - by - column. The formula is:

[0028]

[0029] where D′ ij is the element of the standardized matrix, D ij is the element in the i - th row and j - th column of the distance matrix D, is the mean of the elements in the distance matrix D, and std represents the standard - deviation calculation process;

[0030] Perform an offset adjustment on the mean - normalized distance matrix to make the matrix elements non - negative. The formula is:

[0031] D″ ij = D′ ij - min(D′ij )

[0032] where min(D′ ij ) is the minimum value of the matrix D′ after mean normalization;

[0033] Each element of the adjusted matrix D″ is operated on by an exponential function to generate the initial assignment matrix A 0 , and the formula is:

[0034] A 0 = e -∈D″

[0035] where ∈ is a hyperparameter.

[0036] Optionally, the row and column normalization iteration of the assignment matrix by the Sinkhorn-Knopp algorithm until the assignment matrix converges includes:

[0037] Taking the initial assignment matrix A 0 as the starting matrix for iteration;

[0038] In each iteration, first normalize each row of the current assignment matrix so that the sum of the elements in each row is 1, and the expression is:

[0039]

[0040] where t represents the current iteration number, represents the sum of all elements in the i-th row;

[0041] Then normalize each column of the row-normalized assignment matrix so that the sum of the elements in each column is 1, and the expression is:

[0042]

[0043] where, represents the sum of all elements in the j-th column;

[0044] After each row and column normalization iteration is completed, update the iteration number t to t + 2;

[0045] Repeat the above row and column normalization operations until the value of the assignment matrix A converges.

[0046] Optionally, after the iteration ends, determining the quantization feature based on the maximum value of the assignment matrix includes:

[0047] For each continuous representation vector z i , by finding the maximum value index position of its corresponding row in the assignment matrix A, determine the quantization result of the vector, and the expression is:

[0048]

[0049] Among them, k represents the codebook index;

[0050] According to the index position k, the continuous representation vector z i is mapped to the corresponding vector c in the codebook k , generating a quantization representation z q = h(z i ) = c k ;

[0051] All the quantization representations are rearranged into a structure consistent with the input data to obtain the quantization feature Z q .

[0052] Optionally, mapping the quantization feature back to the original data space through the decoder to obtain a reconstructed data result, including:

[0053] Inputting the quantization feature Z q into the decoder to generate reconstructed data consistent with the original data Among them, the decoder is a neural network model that accepts the quantization feature as input and outputs data with the same structure as the original data.

[0054] To achieve the above object, an embodiment of the second aspect of the present application proposes a global structure-aware vector quantization device based on optimal transport, including:

[0055] An encoding module, configured to map input data to a continuous latent variable space through an encoder to generate a continuous representation;

[0056] A construction and solution module, configured to construct a distance matrix between the continuous representation and the codebook, and use an optimal transport problem to solve the assignment matrix;

[0057] A normalization and initialization module, configured to perform a normalization process on the distance matrix and generate an initial value of the assignment matrix using a special initialization method;

[0058] A normalization iteration module, configured to perform row and column normalization iteration on the assignment matrix through the Sinkhorn-Knopp algorithm until the assignment matrix converges;

[0059] A quantization and reconstruction module, configured to, after the iteration ends, determine the quantization feature based on the maximum value of the assignment matrix, and map the quantization feature back to the original data space through the decoder to obtain a reconstructed data result.

[0060] To achieve the above object, an embodiment of the third aspect of the present application proposes an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0061] The memory stores computer-executable instructions;

[0062] The processor executes the computer-executable instructions stored in the memory to implement the method described in any one of the first aspects.

[0063] To achieve the above object, an embodiment of the fourth aspect of the present application proposes a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in any one of the first aspects.

[0064] To achieve the above object, an embodiment of the fifth aspect of the present application proposes a computer program product, and when the computer program is executed by a processor, it implements the method described in any one of the first aspects.

[0065] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:

[0066] (1) The present application formulates the vector quantization framework as an optimal transport problem and optimizes the assignment matrix through the Sinkhorn-Knopp algorithm, realizing an efficient mapping from data features to codebooks from a global perspective. Different from traditional nearest neighbor search, the present application avoids the local optimal trap caused by the greedy strategy, significantly enhancing the stability and efficiency of training. Without relying on complex initialization or distillation strategies, the training process is simplified.

[0067] (2) Through the global optimization strategy, the present application ensures the full utilization of each element in the codebook, solving the common "index collapse" problem in traditional vector quantization methods. The codebook utilization rate close to 100% greatly improves the expressive ability of the model, while enhancing the compactness and accuracy of data representation.

[0068] (3) By introducing a global structure awareness mechanism, the present application makes full use of the global distribution information of the data, making the reconstructed data results have higher fidelity. Experiments in image reconstruction tasks show that the reconstruction quality of the present application is better than that of the existing state-of-the-art vector quantization networks.

[0069] (4) The present application designs a simple but effective normalization strategy, reducing the impact of different data distribution ranges on the quantization process and enhancing the robustness of the algorithm. This normalization technique can effectively neutralize the interference of the data value range on the performance of the Sinkhorn algorithm, ensuring that the quantization process has high stability in various data modalities.

[0070] (5) This application can process various continuous data modalities such as images and videos, and demonstrates extremely high flexibility and applicability in discrete compression and reconstruction scenarios. Its generality enables this application to serve as a bridge between large-scale sequence models and data modality distributions, providing strong technical support for general data modeling tasks.

[0071] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of this application. Brief Description of the Drawings

[0072] The above and / or additional aspects and advantages of this application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:

[0073] Figure 1 is a schematic text flow diagram of a global structure-aware vector quantization method based on optimal transport provided by an embodiment of this application;

[0074] Figure 2 is a schematic image flow diagram of a global structure-aware vector quantization method based on optimal transport provided by an embodiment of this application;

[0075] Figure 3 is a schematic diagram comparing the technical effects of the method of this application and the conventional VQ method provided by an embodiment of this application;

[0076] Figure 4 is a schematic diagram comparing the technical means of the method of this application and the conventional VQ method provided by an embodiment of this application;

[0077] Figure 5 is a schematic structural diagram of a global structure-aware vector quantization device based on optimal transport provided by an embodiment of this application. Detailed Description of the Embodiments

[0078] Embodiments of this application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain this application and should not be construed as limiting this application.

[0079] In view of the problems existing in the prior art, an embodiment of this application provides a global structure-aware vector quantization method based on optimal transport. Figure 1 and Figure 2 are respectively a schematic text flow diagram and a schematic image flow diagram of a global structure-aware vector quantization method based on optimal transport provided by an embodiment of this application. As Figure 1 shown, the method includes the following steps:

[0080] Step 101, map the input data to a continuous latent variable space through an encoder to generate a continuous representation. In the embodiments of the present application, first represent the original input data as multiple data samples where N represents the total number of data samples, and each sample x i is an instance of the input data (such as an image or a video frame).

[0081] Next, use the encoder f to process each data sample x, and map it from the original data space to the continuous latent variable space. The mapping relationship is:

[0082] Z e = f(x)

[0083] where Z e ∈R m×d represents the representation in the continuous latent variable space, m is the number of samples, and d represents the dimension of each vector in the latent variable space. The continuous representation Z e is a matrix composed of multiple vectors, and each vector element represents a specific feature of the original data.

[0084] In the embodiments of the present application, each vector element of the continuous representation Z e is denoted as z e ∈R d , z e is a vector with a length of z e , representing the feature representation of the data sample in the continuous latent variable space.

[0085] As a possible implementation, the encoder f can be implemented based on a deep neural network, such as a convolutional neural network (CNN) or other architectures suitable for the input data type. The main role of the encoder is to extract the feature information in the input data and represent this information as a high-dimensional vector in the continuous latent variable space. The present application does not make specific limitations on the structure of the encoder.

[0086] Step 102, construct a distance matrix between the continuous representation and the codebook, and use the optimal transport problem to solve the assignment matrix.

[0087] It can be understood that the essence of vector quantization is to find the best matching relationship between the continuous representation z e and the codebook .

[0088] To this end, in the embodiments of the present application, first construct a distance matrix D, and each element D of this matrix ij is defined as the Euclidean distance between the vector z i in the continuous representation and the codebook vector c j . The formula is as follows:

[0089] D ij = ‖z i - c j ‖

[0090] In the formula, z i represents each vector in the continuous representation, and c j represents the codebook vector in the codebook . The distance matrix D is an l×n matrix, where l represents the number of features and n represents the number of codebooks.

[0091] To achieve the best mapping relationship from the continuous representation to the codebook, this application models the vector quantization problem as an optimal transport problem. The goal is to solve the assignment matrix A to satisfy the global optimality of the assignment while minimizing the transport cost. Its optimization objective function is:

[0092]

[0093] where Tr(A T D) represents the total transport cost, and A T D is the matrix product; H(A) = -∑ ij A ij logA ij is the entropy function, which encourages the dispersion of the assignment matrix; ∈ is a hyperparameter used to adjust the balance between the transport cost and the dispersion.

[0094] The assignment matrix A needs to satisfy the following constraints:

[0095]

[0096] where A1 r = 1 r represents the row normalization of the assignment matrix, ensuring that each feature vector can be assigned; A T 1 c = 1 c represents the column normalization of the assignment matrix, ensuring that each codebook participates in the assignment; represents that all elements of the assignment matrix are non - negative.

[0097] By optimizing the above objective function, the best assignment scheme from the feature vectors to the codebook vectors can be found globally, thus realizing the quantization with global structure awareness.

[0098] By introducing the optimal transport theory, the embodiments of the present application overcome the problem of local optimal traps caused by greedy nearest neighbor search in traditional vector quantization. The introduction of the entropy function further alleviates the "index collapse" phenomenon, enabling each codebook to be fully utilized and enhancing the expressive ability and utilization rate of the codebook. This optimization strategy provides an effective solution for the global optimization of vector quantization and significantly improves the stability and efficiency of training.

[0099] Step 103: Normalize the distance matrix and generate the initial value of the assignment matrix using a special initialization method.

[0100] The embodiments of the present application adopt an iterative approach to solve the approximate optimal solution of the above problem. Before that, first normalize the distance matrix D to eliminate the influence of its value range on the result, and at the same time use a special initialization method to generate the initial assignment matrix A 0 , laying the foundation for subsequent iterative optimization.

[0101] The specific steps are as follows:

[0102] First, perform mean normalization on the distance matrix D column by column. The formula is:

[0103]

[0104] where D′ ij is the element of the standardized matrix, D ij is the element in the i-th row and j-th column of the distance matrix D, is the mean of the elements in the distance matrix D, and std represents the standard deviation calculation process.

[0105] The purpose of mean normalization is to eliminate the influence of its value range on the result, thereby improving the stability and robustness of subsequent calculations.

[0106] Next, perform offset adjustment on the mean-normalized distance matrix to make the matrix elements non-negative. The formula is:

[0107] D″ ij = D′ ij - min(D′ ij )

[0108] where min(D′ ij ) is the minimum value of the mean-normalized matrix D′.

[0109] This step ensures that all values in the normalized matrix D″ are non-negative to meet the mathematical definition and physical meaning of the optimal transport problem.

[0110] Finally, perform an operation on each element of the adjusted matrix D″ through the exponential function to generate the initial assignment matrix A0 , the formula is:

[0111] A 0 = e -∈D″

[0112] where ∈ is a hyperparameter.

[0113] The assignment matrix A generated by this special initialization method 0 provides an optimization starting point for subsequent iterations. While considering distance information, this matrix effectively reduces the impact of large distances through an exponential decay mechanism, thereby improving the rationality of the initial assignment.

[0114] Step 104: Perform row and column normalization iterations on the assignment matrix through the Sinkhorn-Knopp algorithm until the assignment matrix converges.

[0115] In the embodiment of this application, the Sinkhorn-Knopp algorithm is used to perform multiple row and column normalization iterations on the initial assignment matrix A 0 to meet the row and column normalization constraints of the assignment matrix and gradually approximate the optimal assignment result.

[0116] First, use the initial assignment matrix A 0 as the starting matrix for iteration. In each iteration process, perform the following steps to normalize the assignment matrix:

[0117] (1) Normalize each row of the current assignment matrix so that the sum of the elements in each row is 1. The expression is:

[0118]

[0119] where t represents the current iteration number, represents the sum of all elements in the i-th row.

[0120] (2) Normalize each column of the assignment matrix after row normalization so that the sum of the elements in each column is 1. The expression is:

[0121]

[0122] where represents the sum of all elements in the j-th column.

[0123] (3) After each row normalization and column normalization operation, update the iteration number t to t + 2, indicating a complete row and column normalization iteration.

[0124] (4) Repeat the above row and column normalization operations until the change in the value of the assignment matrix between two iterations is less than a preset threshold or the preset maximum number of iterations is reached.

[0125] In actual operation, due to the fast convergence of the Sinkhorn-Knopp algorithm, the assignment matrix can generally meet the row and column normalization requirements within 5 iterations, so there will be no high computational overhead.

[0126] Through the above steps, while meeting the row and column normalization constraints, the assignment matrix A gradually approaches the optimal assignment result. This iterative optimization method not only ensures the global convergence of the algorithm but also effectively avoids the local optimal trap, laying a solid foundation for the determination of subsequent quantization features. The high efficiency and stability of the Sinkhorn-Knopp algorithm make it an ideal choice for solving the optimal transport problem.

[0127] Step 105, after the iteration ends, determine the quantization features based on the maximum value of the assignment matrix, and map the quantization features back to the original data space through the decoder to obtain the reconstructed data result.

[0128] After the iterative optimization of the Sinkhorn-Knopp algorithm is completed, the embodiments of the present application determine the quantization features according to the maximum value of the assignment matrix A, and map the quantization features back to the original data space through the decoder, and finally generate the reconstructed data result.

[0129] The specific steps are as follows:

[0130] First, for each continuous representation vector z i , determine the quantization result of the vector by finding the index position of the maximum value in the corresponding row of the assignment matrix A. The calculation formula for the quantization result is:

[0131]

[0132] where k represents the codebook index.

[0133] Next, according to the codebook index k, map the continuous representation vector z i to the corresponding vector c k in the codebook. The generation formula for the quantization representation is:

[0134] z q = h(z i ) = c k

[0135] In the formula, h represents the quantization function, and c k represents the k-th codebook vector in the codebook C.

[0136] Then, rearrange all the quantization representations z q according to the structure of the continuous representation Z e to generate the quantization feature matrix Z qThis process ensures that the arrangement order of the quantized features is consistent with the order of the input data.

[0137] Finally, the quantized feature Z q is input into the decoder g, and the quantized feature is mapped back to the original data space through the decoder to generate the reconstructed data result. The expression of the decoding process is:

[0138]

[0139] where represents the reconstructed data, which has the same structure as the original data x; g represents the decoder, which is a neural network-based model that accepts the quantized feature Z q as input and generates the reconstructed data.

[0140] It can be understood that the design goal of the decoder is to restore the original information of the input data as much as possible. Therefore, its structure is usually symmetric with the encoder, and the input data is restored by decoding and reconstructing the quantized feature. The output of the decoder retains most of the information of the input data, and at the same time realizes efficient representation and compression through the quantized feature.

[0141] Through this step, the input data completes the conversion from the original data to the discrete tokens and then to the reconstructed data during the processes of encoding, quantization, and decoding.

[0142] To demonstrate the advantages of the method of this application, Figure 3 and Figure 4 compare the technical effects and means of the method of this application with the conventional VQ method. OptVQ in the figure is the method of this application.

[0143] By using the optimal transport method to optimize and solve the matching relationship between the data and the codebook, this application achieves extremely high codebook utilization rate and significantly reduces the reconstruction loss. Therefore, compared with the traditional method, the method of this application can obtain higher-quality data compression and reconstruction effects. The method of this application overcomes the limitation of relying on local optimality in the traditional method through distribution perception and global optimization, and thus demonstrates excellent performance advantages in the discrete reconstruction scenarios of continuous data such as images and videos.

[0144] To implement the above embodiments, this application also proposes a global structure-aware vector quantization device based on optimal transport. Figure 5 is a schematic structural diagram of a global structure-aware vector quantization device based on optimal transport provided by an embodiment of this application. As Figure 5 shown, the device includes:

[0145] An encoding module 100, configured to map input data to a continuous latent variable space through an encoder to generate a continuous representation;

[0146] The construction and solution module 200 is used to construct the distance matrix between the continuous representation and the codebook, and solve the assignment matrix using the optimal transport problem;

[0147] The normalization and initialization module 300 is used to normalize the distance matrix and generate the initial value of the assignment matrix using a special initialization method;

[0148] The normalization iteration module 400 is used to perform row and column normalization iteration on the assignment matrix through the Sinkhorn-Knopp algorithm until the assignment matrix converges;

[0149] The quantization and reconstruction module 500 is used, after the iteration ends, to determine the quantization features based on the maximum value of the assignment matrix, and map the quantization features back to the original data space through a decoder to obtain the reconstructed data result.

[0150] To implement the above embodiments, the present application also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0151] To implement the above embodiments, the present application also proposes a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are used to implement the method provided in the foregoing embodiments when executed by a processor.

[0152] To implement the above embodiments, the present application also proposes a computer program product including a computer program, and the computer program implements the method provided in the foregoing embodiments when executed by a processor.

[0153] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present application and other processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0154] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing the relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0155] This application is expected to provide an implementation scheme for users to selectively prevent the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of users.

[0156] In the description of the foregoing embodiments, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0157] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0158] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred implementation of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.

[0159] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definable list of executable instructions for implementing logical functions, and can be embodied specifically in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in connection with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0160] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0161] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0162] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0163] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

[0164] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present application can be achieved, and no limitation is imposed herein.

[0165] The above specific implementation manners do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A global structure-aware vector quantization method based on optimal transmission, characterized in that: The following steps are involved: The encoder maps the input data to a continuous latent variable space to generate a continuous representation. Constructing a distance matrix between the continuous representation and the codebook, and solving the allocation matrix using the optimal transmission problem; Normalizing the distance matrix and generating initial values ​​of the allocation matrix using a special initialization method; Iteratively normalizing the allocation matrix by using a Sinkhorn-Knopp algorithm until the allocation matrix converges; After the iteration is completed, the quantization feature is determined based on the maximum value of the allocation matrix, and the quantization feature is mapped back to the original data space through the decoder to obtain a reconstructed data result.

2. The method according to claim 1, characterized in that The encoder maps the input data to a continuous latent variable space to generate a continuous representation, including: Represent the original input data as multiple data samples Use the encoder f to map each data sample x into a representation in the continuous latent variable space to obtain a continuous representation Z e =f(x)∈R m×d , let each vector element of the continuous representation be z e ∈R d , where m is the number of samples and d is the dimension of each vector.

3. The method according to claim 2, characterized in that: The step of constructing a distance matrix between the continuous representation and the codebook and solving the allocation matrix using the optimal transmission problem includes: Use Euclidean distance to construct the distance matrix D, where D ij =‖z i -c j ‖, z i represents each vector in the continuous representation, c j Representation Codebook The codebook vector in ; The optimization model of the allocation matrix A is established through the objective function of the optimal transmission problem, and the form of the optimal transmission problem is as follows: Where H(A) = -∑ ij A ij logA ij represents the entropy function, ∈ is a hyperparameter, 1 r and 1 c are all-1 vectors of length r and c respectively, l represents the number of features, and n represents the number of codebooks.

4. The method according to claim 3, characterized in that The normalizing process of the distance matrix and using a special initialization method to generate an initial value of the allocation matrix includes: The distance matrix D is mean-normalized by column, and the formula is: Among them, D′ ij is the standardized matrix element, D ij is the element in the i-th row and j-th column of the distance matrix D, is the mean of the elements in the distance matrix D, and std represents the standard deviation calculation process; The distance matrix after mean standardization is offset adjusted to make the matrix elements non-negative. The formula is: D ij =D′ ij -min(D′ ij ) Among them, min(D′ ij ) is the minimum value of the matrix D after mean normalization; The initial allocation matrix A is generated by operating each element of the adjusted matrix D″ through the exponential function 0 , the formula is: TO 0 =and -∈D″ Among them, ∈ is a hyperparameter.

5. The method according to claim 4, characterized in that The iterative process of normalizing the allocation matrix by using the Sinkhorn-Knopp algorithm until the allocation matrix converges includes: The initial allocation matrix A 0 As the starting matrix for iteration; In each iteration, each row of the current allocation matrix is ​​first normalized so that the sum of the elements in each row is 1. The expression is: Where t represents the current iteration number, Represents the sum of all elements in the i-th row; Then normalize each column of the row-normalized distribution matrix so that the sum of the elements in each column is 1. The expression is: in, represents the sum of all elements in the jth column; After each row and column normalization iteration is completed, the number of iterations t is updated to t+2; The above row and column normalization operation is repeated until the value of the allocation matrix A converges.

6. The method according to claim 5, characterized in that After the iteration is completed, determining the quantitative feature based on the maximum value of the allocation matrix includes: For each continuous representation vector z i , by finding the maximum value index position of the corresponding row in the allocation matrix A, the quantization result of the vector is determined, and the expression is: Wherein, k represents the codebook index; The continuous representation vector z is converted according to the index position k. i Mapped to the corresponding vector c in the codebook k , generate the quantized representation z q =h(z i )=c k ; Rearrange all quantized representations into a structure consistent with the input data to obtain the quantized feature Z q .

7. The method according to claim 6, characterized in that Mapping the quantized features back to the original data space through a decoder to obtain a reconstructed data result includes: The quantized feature Z q Input decoder to generate reconstructed data consistent with the original data The decoder is a neural network model that accepts quantized features as input and outputs data with the same structure as the original data.

8. A global structure-aware vector quantization device based on optimal transmission, characterized in that: include: The encoding module is used to map the input data into a continuous latent variable space through an encoder to generate a continuous representation; A construction and solution module, used to construct a distance matrix between the continuous representation and the codebook, and solve the allocation matrix using the optimal transmission problem; A normalization and initialization module, used for normalizing the distance matrix and generating initial values ​​of the distribution matrix using a special initialization method; A normalization iteration module, used for performing row and column normalization iteration on the allocation matrix by using a Sinkhorn-Knopp algorithm until the allocation matrix converges; The quantization and reconstruction module is used to determine the quantization feature based on the maximum value of the allocation matrix after the iteration, and map the quantization feature back to the original data space through the decoder to obtain the reconstructed data result.

9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Cited By

  • Adaptive feature fusion method and device based on optimal transmission

    CN120510143A

  • An adaptive feature fusion method and apparatus based on optimal transmission

    CN120510143B