Lossless coding model fine tuning method, device, equipment, medium and product

Through low-rank decomposition and progressive fine-tuning methods, the problem of poor performance of lossless coding models on the extradomain image test set is solved, and efficient and adaptable lossless coding is achieved, which improves compression rate and calculation efficiency.

CN120411265APending Publication Date: 2025-08-01HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510468014.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing lossless coding fine-tuning methods have poor performance on extra-domain image test sets and are computationally large, and traditional lossless coding methods are difficult to capture complex image distribution, resulting in limited compression performance.

Method used

Using low-rank decomposition and progressive fine-tuning methods, the pre-trained lossless coding model is fine-tuned on the training image, the number of tiles is gradually increased, the incremental weight is optimized, and the bit rate-guided progressive fine-tuning strategy is used to capture image features and adapt to extra-domain images.

Benefits of technology

It realizes efficient and adaptable lossless coding on extra-domain images, reduces the calculation amount, improves the compression rate, and improves the performance of the lossless coding model in extra-domain images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411265A_ABST
    Figure CN120411265A_ABST
Patent Text Reader

Abstract

The invention discloses a lossless coding model fine tuning method and device, equipment, a medium and a product, and relates to the field of image processing, and the method comprises the steps: decomposing the weight of a linear layer and each network layer in a lossless coding model in a fine tuning stage, and determining the increment weight of the linear layer and each network layer; on the basis of a progressive fine tuning method guided by a code rate, performing fine tuning on the pre-trained lossless coding model on the training image, quantifying an increment weight, and determining a loss function of the pre-trained lossless coding model; on the basis of the loss function, according to the quantized increment weight and a pre-trained weight, determining a weight after fine tuning, and reconstructing the pre-trained lossless coding model; according to the reconstructed lossless coding model, the test image is coded into the bit stream, the decoder decodes the increment weight of the reconstructed lossless coding model, and the decoded image is determined, the calculation amount is reduced, and the performance of the pre-training model on the test set of the out-of-domain image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and particularly to a method, device, equipment, medium and product for fine-tuning a lossless coding model. Background Art

[0002] Currently, parameter-efficient fine-tuning (PEFT) strategies can be divided into three main categories.

[0003] Adapter modules: Introduced by Houlsby et al., small trainable networks added to existing models. These modules adjust the internal representation of the model to adapt to new tasks while keeping most of the original parameters unchanged.

[0004] Prompt tuning: Using task-specific prompts added to the input data. This method guides the pre-trained model to generate task-related outputs without changing its parameters.

[0005] Low-rank adaptation: Modifying the weights of the pre-trained network through low-rank matrices to achieve task adaptation with fewer additional parameters.

[0006] Some existing studies have combined learning-based lossy image compression with PETL, inserting adapters in the transformation network, mainly to address the performance degradation problem of out-of-domain images.

[0007] In the field of lossy coding, work has been done to introduce the adapter module to improve the performance degradation problem that occurs when the lossy encoder processes out-of-domain (OOD) data. First, given a pre-trained lossy encoder, all its parameters are frozen, i.e., not trained. At this time, an adapter is introduced into the decoder of the lossy coding framework, and the parameters of the adapter are optimized for the current input image, and these parameters are quantized and encoded into the bitstream. So finally, the bitstream of the image includes the bitstream of the image itself and the bitstream required for encoding the adapter.

[0008] The outline of adapter training is: optimizing the parameters of the adapter as part of the latent representation optimization. Let w be the quantization interval, be the quantized adapter parameters, using an approximation method to quantize θ, and optimizing θ in terms of rate distortion. A hybrid quantization method is used, that is, uniformly quantizing θ in the decoder using the straight-through estimator and adding uniform noise U(-w / 2, w / 2) to θ for the entropy model; let be the adapter parameters with added uniform noise.

[0009] Among them, the loss function used to optimize θ is expressed as is the entropy model of θ. The first term is the loss function of the bit rate, which is used to calculate the entropy of. A logical distribution with a scale of is used as p.

[0010] After this optimization, the updated adapter parameter θ * is obtained and is defined as follows:

[0011] Finally, θ * is quantized, and is encoded through entropy coding to obtain the corresponding bitstream.

[0012] Existing fine-tuning methods in lossy coding are based on the Adapter module, which increases the time complexity during inference. At the same time, the application scenarios of lossless coding are not specially considered. If such methods are directly applied to the lossless field, that is, directly fine-tuning the model on the full image throughout the process, it will lead to a relatively large amount of computation.

[0013] Traditional lossless image compression methods, such as PNG, JPEG-LS, and JPEG-XL, use manually designed algorithms to utilize the statistical characteristics of images. However, these methods usually have difficulty capturing the complex and diverse distributions in the original images, which limits their compression performance. In recent years, deep learning-based methods have achieved optimal results by learning the complex distributions of original images. These methods use likelihood-based generative models, such as autoregressive models, flow models, and variational autoencoders, to convert data into a compact bitstream by predicting the probabilities of entropy coding. Although these methods have been successful, they usually rely on large datasets for training, which may lead to suboptimal probability estimates for specific test images. The unique details of each image, such as texture, lighting, or structure, pose challenges that need to be addressed to improve compression performance.

[0014] It can be seen that existing lossless coding does not consider the problem of out-of-domain image processing. Therefore, there is a gap between the training images and the test images, and the specific manifestation is that the performance of the pre-trained model will deteriorate on the test set of out-of-domain images. Summary of the Invention

[0015] The purpose of this application is to provide a method, device, equipment, medium, and product for fine-tuning a lossless coding model to solve the problems that the fine-tuning method for lossy coding directly fine-tunes the model on the full image throughout the process, resulting in a large amount of computation, and the performance of the pre-trained model in the lossless coding fine-tuning method is poor on the test set of out-of-domain images.

[0016] To achieve the above purpose, this application provides the following solutions:

[0017] In the first aspect, this application provides a method for fine-tuning a lossless coding model, including:

[0018] In the fine-tuning stage, decompose the weights of the linear layer and each network layer in the lossless coding model to determine the incremental weights of the linear layer and each network layer; the network layer includes a masked convolutional layer and a masked depth convolutional layer;

[0019] Based on the rate-guided progressive fine-tuning method, fine-tune the pre-trained lossless coding model on the training images, and quantize the incremental weights to determine the loss function of the pre-trained lossless coding model;

[0020] Based on the loss function, determine the fine-tuned weights according to the quantized incremental weights and the pre-trained weights, and reconstruct the pre-trained lossless coding model;

[0021] According to the reconstructed lossless coding model, encode the test image into a bitstream, and decode the incremental weights of the reconstructed lossless coding model by the decoder to determine the decoded image.

[0022] In a second aspect, the present application provides a device for fine-tuning a lossless coding model, including:

[0023] An incremental weight determination module, configured to decompose the weights of the linear layer and each network layer in the lossless coding model in the fine-tuning stage to determine the incremental weights of the linear layer and each network layer; the network layer includes a masked convolutional layer and a masked depth convolutional layer;

[0024] A loss function determination module, configured to fine-tune the pre-trained lossless coding model on the training images based on the rate-guided progressive fine-tuning method, and quantize the incremental weights to determine the loss function of the pre-trained lossless coding model;

[0025] A reconstruction module, configured to determine the fine-tuned weights according to the quantized incremental weights and the pre-trained weights based on the loss function, and reconstruct the pre-trained lossless coding model;

[0026] An encoding and decoding module, configured to encode the test image into a bitstream according to the reconstructed lossless coding model, and decode the incremental weights of the reconstructed lossless coding model by the decoder to determine the decoded image.

[0027] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the lossless coding model fine-tuning method described in any one of the above.

[0028] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the lossless coding model fine-tuning method described in any one of the above.

[0029] In a fifth aspect, the present application provides a computer program product, including a computer program which, when executed by a processor, implements the lossless coding model fine-tuning method described in any one of the above.

[0030] According to the specific embodiments provided by the present application, the following technical effects are disclosed:

[0031] The present application fine-tunes a pre-trained lossless coding model on training images to achieve single-instance adaptation on out-of-domain test images. Both low-rank decomposition and progressive fine-tuning in the present application are strategies for fine-tuning a pre-trained model using the Low-Rank Adaptation (LoRA) method. Among them, the low-rank decomposition part describes how to introduce low-rank incremental parameters, and the Rate-guided Parameter-efficient Finetuning (RPFT) describes how to efficiently train the introduced incremental parameters. The present application makes the model more adaptable to the image features of out-of-domain images by fine-tuning the additional parameters introduced by the LoRA method, and obtains a higher compression rate.

[0032] Compared with the existing technical solution of directly fine-tuning the model on the entire image throughout the process, the present application adopts a progressive fine-tuning scheme, that is, at the beginning, a part of the patches are used to fine-tune the incremental parameters, and then the number of patches used for fine-tuning is gradually increased until finally all patches are used for fine-tuning. Compared with fine-tuning the model using the entire image throughout the process, the computational amount is less. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 It is a flowchart of the lossless coding model fine-tuning method provided by an embodiment of the present application;

[0035] Figure 2 It is a schematic diagram of a progressive curve provided by an embodiment of the present application;

[0036] Figure 3 It is a schematic diagram of the overall fine-tuning process provided by an embodiment of the present application;

[0037] Figure 4 It is a schematic diagram of low-rank decomposition of different network layers in lossless coding provided by an embodiment of the present application;

[0038] Figure 5 Schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0040] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0041] An embodiment of the present application provides a method for fine-tuning a lossless coding model. This method is executed by a computer device, specifically, it can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiment of the present application, as Figure 1 shown, this method includes the following steps.

[0042] S1: In the fine-tuning stage, decompose the weights of the linear layer and each network layer in the lossless coding model to determine the incremental weights of the linear layer and each network layer; the network layer includes a masked convolutional layer and a masked depth convolutional layer.

[0043] S2: Based on the code rate-guided progressive fine-tuning method, fine-tune the pre-trained lossless coding model on the training images, and quantize the incremental weights to determine the loss function of the pre-trained lossless coding model.

[0044] S3: Based on the loss function, determine the fine-tuned weights according to the quantized incremental weights and the pre-trained weights, and reconstruct the pre-trained lossless coding model.

[0045] S4: According to the reconstructed lossless coding model, encode the test image into a bitstream, and decode the incremental weights of the reconstructed lossless coding model by a decoder to determine the decoded image.

[0046] In an exemplary embodiment, S1 can be replaced by the following steps.

[0047] S11: Based on the number of output channels and input channels of the linear layer, for the pre-trained weight matrix of the linear layer, update the pre-trained weights in the linear layer according to the product of two low-rank matrices to determine the incremental weights of the linear layer.

[0048] S12: Using Tucker decomposition, determine a core identity tensor and four mode matrices according to the convolution kernels of the pre-trained masked convolutional layer and the number of intermediate channels, and determine the incremental weights of the pre-trained masked convolutional layer according to the core identity tensor and the four mode matrices.

[0049] S13: Using Tucker decomposition, determine a core identity tensor and three mode matrices according to the convolution kernels of the pre-trained masked depthwise convolutional layer and the number of intermediate channels, and determine the incremental weights of the pre-trained masked depthwise convolutional layer according to the core identity tensor and the three mode matrices.

[0050] In an exemplary embodiment, S2 can be replaced by the following steps.

[0051] S21: Increase the training sample ratio according to the number of tiles participating in the model fine-tuning process.

[0052] S22: Determine the number of training tiles based on the training sample ratio, fine-tune the pre-trained lossless coding model on the selected training tiles, and optimize the incremental weights and the bit rate of the image pixels during the fine-tuning process to determine the loss function of the pre-trained lossless coding model.

[0053] In an exemplary embodiment, S21 can be replaced by the following steps.

[0054] Use the formula F(t) = b + (1 - b)·[s(t′)] e The number of tiles used in the fine-tuning process to increase the training sample ratio; where F(t) is the number of tiles used in the fine-tuning process; b is the minimum value at the start of the tile growth curve; s(t′) is a smoothing function such that the curve is differentiable at the start and stop of growth; e is to adjust the curve shape and growth rate.

[0055] Furthermore, to further accelerate the fine-tuning adaptation process, the RPFT method is used to increase the training sample ratio. The RPFT method is based on the following assumptions: During the entire fine-tuning process, it is not necessary to use all tiles. Two key considerations support this assumption: 1) There is inherent redundancy among tiles due to content similarity; 2) The distribution complexity of different tiles varies, which will affect the model's learning process in different ways. These considerations prompt us to focus on two aspects in the design: First, gradually increase the number of tiles used for fine-tuning; second, the fine-tuning process should focus on training the blocks containing more information structures or details and perform more training steps on these blocks.

[0056] Therefore, the estimated bitrate is selected as an indicator of the information content of each tile, and a strategy of gradually increasing the number of training tiles is adopted. These tiles are sorted from high to low according to the bitrate. First, the bitrates of all tiles are estimated, arranged from high to low, and gradually incorporated into our training samples for fine-tuning. Initially, only b% of the tiles are selected as training samples. As the training progresses, more tiles are systematically added, and finally the complete set of tiles is used in the last d% of the training steps. Fine-tuning on the complete set ensures the consistency of the training objective and the test objective.

[0057] The process of increasing the ratio of training samples F(t) through the following function:

[0058]

[0059] F(t) = b + (1 - b)·[s(t′)] e ,

[0060] where t represents the current step, T represents the total number of training steps, and s(x) is a smoothing function used to ensure that the start and end positions of the final tile growth curve are differentiable. In this application, x is t′, that is, s(t′) = s(x).

[0061] The growth process of F(t) is as Figure 2 shown. By adjusting b, d, and e, the number of tiles participating in the fine-tuning process can be controlled. Among them, d controls the position where the curve reaches the maximum value. Specifically, as b and d increase and e decreases, the number of tiles participating in the training will increase. Through the RPFT method, the focus is placed on learning the distribution of tiles with higher information content while effectively reducing the number of tiles participating in the training. This strategy not only optimizes the learning process but also effectively shortens the time required for adaptation.

[0062] In an exemplary embodiment, during the fine-tuning process, the incremental weight and the bitrate of the image pixels are optimized, and the loss function of the pre-trained lossless coding model is determined, specifically including:

[0063] During the fine-tuning process, the loss function of the pre-trained lossless coding model is determined using the formula ; where is the loss function; is the probability distribution of the incremental weight; is the weight after adding noise; q is the entropy model; x is the training image; θ is the pre-trained parameter; is the quantized incremental weight.

[0064] Furthermore, the complete lossless image compression process is as Figure 3As shown, the first stage is to optimize the added LoRA parameters through rate-guided progressive fine-tuning, and the second stage is to encode the LoRA parameters. Before encoding the test image into a bitstream, the incremental weights are first fine-tuned for the current image as a model prompt, while the pre-trained parameters remain unchanged.

[0065] During the encoding process, these incremental parameters are quantized into discrete values and then incorporated into the bitstream. Given the relatively narrow numerical range of the incremental weights, a smaller quantization step, w < 1, is adopted. To perform a differentiable estimate of quantization during the optimization process, this application uses a hybrid quantization method.

[0066] Let φ denote the incremental weights and θ denote the pre-trained weights. This process uses the straight-through estimator (STE) to quantize the incremental weights for network inference, where sg represents the stop-gradient operation: In addition, uniform noise is added to estimate the entropy:

[0067] This application jointly optimizes the bit rate of the incremental weights and the image pixels to obtain the following loss function:

[0068]

[0069] where, represents the probability distribution of the incremental weights, modeled using a static logistic distribution with zero mean and constant scale , q is the entropy model, and the model has pre-trained parameters θ and quantized incremental weights The entropy model predicts the probability distribution of the image x through these two parts of parameters.

[0070] After fine-tuning the incremental parameters using the above loss, this application explicitly calculates and combines the weights and then performs network inference as usual, ensuring no additional time is required at this stage. Finally, the image is encoded into a bitstream using the final model.

[0071] The total bit rate of the compressed image includes the bit rate of the quantized incremental weights and the image pixels. This bitstream is then transmitted to the decoder, which first decodes the incremental weights and uses them for image decoding.

[0072] In an exemplary embodiment, S3 can be replaced by the following steps.

[0073] S31: Determine the fine-tuned weight of the linear layer using the formula W′ = W + ΔW = W + AB; where, W′ is the fine-tuned weight of the linear layer; W is the pre-trained weight of the linear layer; △W is the quantized incremental weight of the linear layer; A and B are low-rank matrices.

[0074] S32: Use the formula W′ mc = M⊙(Wmc +ΔW mc ) Determine the weights after fine-tuning of the masked convolutional layer; where, W′ mc is the weight after fine-tuning of the masked convolutional layer; M is the mask in the masked convolutional layer; ⊙ is the Hadamard product operation; W mc is the pre-trained weight of the masked convolutional layer; ΔW mc is the increment weight after quantization of the masked convolutional layer.

[0075] S33: Use the formula W′ mdc = M⊙(W mdc +ΔW mdc ) to determine the weights after fine-tuning of the masked depth convolutional layer; where, W′ mdc is the weight after fine-tuning of the masked depth convolutional layer; W mdc is the pre-trained weight of the masked depth convolutional layer; ΔW mdc is the increment weight after quantization of the masked depth convolutional layer.

[0076] S34: Reconstruct the pre-trained lossless coding model according to the weights after fine-tuning of the linear layer, the weights after fine-tuning of the masked convolutional layer, and the weights after fine-tuning of the masked depth convolutional layer.

[0077] Furthermore, this application uses a parameter-efficient fine-tuning method to adjust the pre-trained weights to capture the unique features of each image. Based on the assumption that the change of the weight matrix has a low-rank property, this application incrementally updates the pre-trained weights in the linear layer through the product of two low-rank matrices. For a pre-trained linear layer weight matrix where, m and n respectively represent the output and input channel numbers of the linear layer, and the weight update is modeled as

[0078] The weight W′ after fine-tuning of the linear layer is expressed as:

[0079] W′ = W + ΔW = W + AB.

[0080] where, is the intermediate channel number of the LoRA decomposition, r << min(n, m).

[0081] In addition, in order to fine-tune the lossless coding model, low-rank decomposition is performed on the network layers that may appear in the lossless coding, and the concept of this low-rank decomposition is extended from the two-dimensional linear layer to the three-dimensional masked convolution and depth convolution, as Figure 4 shown.

[0082] For a masked convolutional layer or a depth convolutional layer, the pre-trained convolutional kernel W mc is expressed as or where k represents the convolutional kernel size.

[0083] Using Tucker decomposition, the incremental weights of the pre-trained masked convolution are decomposed into five parts: a core identity tensor where r1, r2, r3, r4 are the numbers of four intermediate channels respectively, r1, r2, r3, r4 << min(n, m), and four mode matrices and The incremental update is defined as ΔW mc = I × 1A × 2B × 3C × 4D, where × i represents the modulo-i product, that is, the multiplication of a matrix and a tensor along the i-th dimension.

[0084] For the pre-trained weights W mc of the above masked convolution layer, the weights W' mc after fine-tuning of the masked convolution layer are

[0085] W' mc = M ⊙ (W mc + ΔW mc ).

[0086] Using Tucker decomposition, the incremental weights of the pre-trained masked depth convolution are decomposed as follows: a core identity tensor and three mode matrices and The incremental update is defined as ΔW mdc = I × 1E × 3F × 4G.

[0087] For the pre-trained weights W mdc of the above masked depth convolution layer, the weights W' mdc after fine-tuning of the masked depth convolution layer are

[0088] W' mdc = M ⊙ (W mdc + ΔW mdc ).

[0089] For ordinary convolution layers, the masks in the above expressions can be deleted. Therefore, the above low-rank decomposition process is applicable to masked convolution layers and masked depth convolution layers except linear layers, covering all common network layers in the lossless coding network.

[0090] It should be noted that the incremental weights of linear layers and convolution layers can be merged with the pre-trained weights, so that no additional computational time is generated during the inference process. We introduce low-rank decomposition into masked convolution gated blocks and feed-forward networks, especially adding incremental weights to the weight matrices W A and W v in the convolutional gating mechanism, and the first linear layer W up in the feed-forward network.

[0091] Based on the same inventive concept, an embodiment of the present application further provides a lossless coding model fine-tuning device for implementing the lossless coding model fine-tuning method involved above. The implementation solutions provided by this device to solve problems are similar to those recorded in the above method. Therefore, the specific limitations in one or more embodiments of the lossless coding model fine-tuning device provided below can refer to the limitations on the lossless coding model fine-tuning method in the above text, and will not be repeated here.

[0092] In an exemplary embodiment, a lossless coding model fine-tuning device is provided, including:

[0093] An incremental weight determination module, configured to decompose the weights of the linear layer and each network layer in the lossless coding model during the fine-tuning stage, and determine the incremental weights of the linear layer and each network layer; the network layer includes a masked convolutional layer and a masked depth convolutional layer.

[0094] A loss function determination module, configured to fine-tune the pre-trained lossless coding model on training images based on the rate-guided progressive fine-tuning method, and quantize the incremental weights to determine the loss function of the pre-trained lossless coding model.

[0095] A reconstruction module, configured to determine the fine-tuned weights based on the loss function, according to the quantized incremental weights and the pre-trained weights, and reconstruct the pre-trained lossless coding model.

[0096] An encoding and decoding module, configured to encode a test image into a bitstream according to the reconstructed lossless coding model, and decode the incremental weights of the reconstructed lossless coding model by a decoder to determine the decoded image.

[0097] In an exemplary embodiment, a computer device is provided, as Figure 5 shown. This computer device can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of this computer device is used to store lossless coding model fine-tuning data. The input / output interface of this computer device is used to exchange information between the processor and external devices. The communication interface of this computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a lossless coding model fine-tuning method.

[0098] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the above method is implemented.

[0099] In an exemplary embodiment, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the above method is implemented.

[0100] In an exemplary embodiment, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the above method is implemented.

[0101] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0102] In this application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining authorization from the owner of the corresponding device.

[0103] In each of the embodiments provided in the present application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., and is not limited thereto. In each of the embodiments provided in the present application, the processor may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., and is not limited thereto.

[0104] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0105] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for fine-tuning a lossless coding model, characterized in that, The method for fine-tuning the lossless coding model includes: In the fine-tuning stage, decompose the weights of the linear layer and each network layer in the lossless coding model to determine the incremental weights of the linear layer and each network layer; the network layer includes a masked convolutional layer and a masked depth convolutional layer; Based on the rate-guided progressive fine-tuning method, fine-tune the pre-trained lossless coding model on the training images, and quantize the incremental weights to determine the loss function of the pre-trained lossless coding model; Based on the loss function, determine the fine-tuned weights according to the quantized incremental weights and the pre-trained weights, and reconstruct the pre-trained lossless coding model; According to the reconstructed lossless coding model, encode the test image into a bitstream, and decode the incremental weights of the reconstructed lossless coding model by the decoder to determine the decoded image.

2. The method for fine-tuning a lossless coding model according to claim 1, wherein In the fine-tuning stage, decompose the weights of the linear layer and each network layer in the lossless coding model to determine the incremental weights of the linear layer and each network layer, specifically including: Based on the number of output channels and input channels of the linear layer, for the pre-trained linear layer weight matrix, update the pre-trained weights in the linear layer according to the product of two low-rank matrices to determine the incremental weights of the linear layer; Use Tucker decomposition to determine a core identity tensor and four mode matrices according to the convolution kernel and the number of intermediate channels of the pre-trained masked convolutional layer, and determine the incremental weights of the pre-trained masked convolutional layer according to a core identity tensor and four mode matrices; Use Tucker decomposition to determine a core identity tensor and three mode matrices according to the convolution kernel and the number of intermediate channels of the pre-trained masked depth convolutional layer, and determine the incremental weights of the pre-trained masked depth convolutional layer according to a core identity tensor and three mode matrices.

3. The method for fine-tuning a lossless coding model according to claim 1, wherein Based on the rate-guided progressive fine-tuning method, fine-tune the pre-trained lossless coding model on the training images, and quantize the incremental weights to determine the loss function of the pre-trained lossless coding model, specifically including: Increase the training sample ratio according to the number of patches participating in the model fine-tuning process; Determine the number of training patches based on the training sample ratio, fine-tune the pre-trained lossless coding model on the selected training images, and optimize the bit rate of the incremental weights and image pixels during the fine-tuning process to determine the loss function of the pre-trained lossless coding model.

4. The method for fine-tuning the lossless coding model according to claim 3, wherein Increase the training sample ratio according to the number of patches participating in the model fine-tuning process, specifically including: Using the formula F(t) = b + (1 - b)·[s(t′)] e The number of tiles used in the fine-tuning process, increasing the training sample ratio; where F(t) is the number of tiles used in the fine-tuning process; b is the minimum value at the start of the tile growth curve; s(t′) is a smoothing function to enable differentiability at the start and stop growth curves; e is to adjust the curve shape and growth rate.

5. The method for fine-tuning a lossless coding model according to claim 3, wherein During the fine-tuning process, optimize the bit rate of the incremental weights and image pixels to determine the loss function of the pre-trained lossless coding model, specifically including: During the fine-tuning process, the loss function of the pre-trained lossless coding model is determined using the formula ; where is the loss function; is the probability distribution of the incremental weights; is the weight after adding noise; q is the entropy model; x is the training image; θ is the pre-trained parameter; is the quantized incremental weight.

6. The method for fine-tuning a lossless coding model according to claim 1, wherein Based on the loss function, determine the fine-tuned weights according to the quantized incremental weights and the pre-trained weights, and reconstruct the pre-trained lossless coding model, specifically including: Use the formula W′ = W + ΔW = W + AB to determine the fine-tuned weights of the linear layer; where W′ is the fine-tuned weights of the linear layer; W is the pre-trained weights of the linear layer; △W is the quantized incremental weights of the linear layer; A and B are low-rank matrices; Using the formula W' mc = M⊙(W mc + ΔW mc ), determine the weight after fine-tuning of the masked convolutional layer; where W' mc is the weight after fine-tuning of the masked convolutional layer; M is the mask in the masked convolutional layer; ⊙ is the Hadamard product operation; W mc is the pre-trained weight of the masked convolutional layer; ΔW mc is the increment weight after quantization of the masked convolutional layer; Using the formula W' mdc = M⊙(W mdc + ΔW mdc ), determine the weights after fine-tuning of the masked depth convolutional layer; where W' mdc is the weight after fine-tuning of the masked depth convolutional layer; W mdc is the pre-trained weight of the masked depth convolutional layer; ΔW mdc is the incremental weight after quantization of the masked depth convolutional layer; Reconstruct the pre-trained lossless coding model according to the weights after fine-tuning of the linear layer, the weights after fine-tuning of the masked convolutional layer, and the weights after fine-tuning of the masked depth convolutional layer.

7. A lossless coding model fine-tuning device, characterized in that The lossless coding model fine-tuning device includes: An incremental weight determination module, configured to decompose the weights of the linear layer and each network layer in the lossless coding model during the fine-tuning stage, and determine the incremental weights of the linear layer and each network layer; the network layer includes a masked convolutional layer and a masked depth convolutional layer; A loss function determination module, configured to fine-tune the pre-trained lossless coding model on the training images based on the rate-guided progressive fine-tuning method, and quantize the incremental weights to determine the loss function of the pre-trained lossless coding model; A reconstruction module, configured to determine the weights after fine-tuning based on the loss function, the quantized incremental weights, and the pre-trained weights, and reconstruct the pre-trained lossless coding model; An encoding and decoding module, configured to encode the test image into a bitstream according to the reconstructed lossless coding model, and decode the incremental weights of the reconstructed lossless coding model by a decoder to determine the decoded image.

8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the lossless coding model fine-tuning method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the lossless coding model fine-tuning method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the lossless coding model fine-tuning method according to any one of claims 1-6.