Model training method and device, storage medium and program product
By obtaining multiple hyperparameters and building multiple loss functions, the cumbersome problem of training compressed models at different bit rates is solved, and the effect of simplifying the training process and improving performance is achieved.
Patent Information
- Application Number
- CN202510130464.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-06
AI Technical Summary
When training compressed models at different bit rates, the prior art requires repeated training, resulting in cumbersome process.
By obtaining multiple hyperparameters, building multiple loss functions, and training the compressed model based on these loss functions, the trained variable-code rate compression model is obtained.
The training process of variable bit rate compression model is simplified, the number of training times is reduced, and the performance of the model in the low bit rate interval is improved.
Smart Images

Figure CN120106149A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a model training method, device, storage medium and program product. Background Art
[0002] With the development of information technology, image compression models can be used to compress images. The performance of a compression model is usually related to the compression bit rate (which can be referred to as the bit rate, which reflects the size of the compressed data) and the degree of distortion of the compressed image. That is, for the same compression bit rate, the lower the degree of distortion of the compressed image, the higher the performance of the compression model.
[0003] At present, in order to improve the performance of the compression model, a fixed hyperparameter λ can be used to train a compression model of one bit rate. Therefore, if it is necessary to obtain compression models at different bit rates, it is necessary to repeatedly train the compression model of each bit rate using different λ values, which makes the training process more cumbersome. Summary of the invention
[0004] The present application provides a model training method, device, storage medium and program product, which can solve the problem that the model training process is relatively complicated.
[0005] In order to achieve the above objectives, this application adopts the following technical solutions:
[0006] In a first aspect, the present application provides a model training method. In the method, multiple hyperparameters are obtained; wherein the multiple hyperparameters correspond to multiple compression bit rates in a target bit rate interval, the target bit rate interval includes multiple sub-bit rate intervals, the difference between the lengths of each sub-bit rate interval is less than a preset length threshold, and the difference between the number of hyperparameters corresponding to each sub-bit rate interval is less than a preset number threshold. Based on the multiple hyperparameters, multiple loss functions are constructed, and one hyperparameter corresponds to at least one loss function. The compression model is trained based on the multiple loss functions to obtain a trained compression model.
[0007] Based on the above technical solution, multiple hyperparameters can be obtained, and the multiple hyperparameters correspond to multiple compression bit rates in the target bit rate interval. The target bit rate interval includes multiple sub-bit rate intervals, and the difference between the lengths of each sub-bit rate interval is less than the preset length threshold. In other words, the length of each sub-bit rate interval is roughly similar. In addition, the difference between the number of hyperparameters corresponding to each sub-bit rate interval is less than the preset number threshold, that is, the number of hyperparameters corresponding to each sub-bit rate interval is roughly the same. In addition, since one hyperparameter corresponds to one compression bit rate, it means that the number of compression bit rates included in each sub-bit rate interval is roughly the same, that is, the compression bit rates can be evenly distributed in the target bit rate interval. Afterwards, multiple loss functions can be constructed based on multiple hyperparameters, and one hyperparameter corresponds to at least one loss function. Then, the compression model can be trained based on multiple loss functions to obtain a trained compression model. In this way, there is no need to train the compression model under different compression bit rates, which can reduce the number of compression model trainings, thereby simplifying the steps of obtaining a variable bit rate compression model. Furthermore, since the compression bit rates can be evenly distributed in the target bit rate range, the number of available compression bit rates in different bit rate ranges is similar, thereby improving the performance of the variable bit rate model in the low bit rate range.
[0008] In a possible design, a first hyperparameter and a second hyperparameter are obtained, wherein the first hyperparameter is the largest hyperparameter among multiple hyperparameters, and the second hyperparameter is the smallest hyperparameter among multiple hyperparameters. The first hyperparameter and the second hyperparameter are processed in an exponential interpolation manner to obtain multiple hyperparameters.
[0009] In a possible design, based on the sizes of multiple hyperparameters, the sampling probabilities of the multiple hyperparameters are determined, the sampling probability of each hyperparameter is positively correlated with the size of the hyperparameter, and the size of the hyperparameter is positively correlated with the size of the compression bitrate corresponding to the hyperparameter. Based on the multiple hyperparameters and the sampling probability of each hyperparameter, multiple loss functions are constructed.
[0010] In a possible design, multiple hyperparameters are divided into multiple parameter sets according to their sizes, and each parameter set corresponds to a sampling probability. Based on the parameter set to which each hyperparameter belongs, the sampling probability of each hyperparameter is determined.
[0011] In one possible design, for each loss function, a loss function is constructed based on a target operation to construct multiple loss functions, and the target operation includes: determining a target hyperparameter based on multiple hyperparameters and a sampling probability of each hyperparameter, where the target hyperparameter is any one of the multiple hyperparameters; obtaining sample data in a training set, and inputting the sample data into a compression model to obtain distortion loss and bit rate loss; and constructing a loss function based on the target hyperparameter, bit rate loss, and distortion loss.
[0012] In one possible design, the compression model is trained based on multiple loss functions to obtain a target loss value. The target loss value is stopped until it meets a preset loss condition, and the trained compression model is obtained. The target loss value is the sum of the loss values corresponding to the multiple loss functions.
[0013] In one possible design, for each loss function, a compression model is set based on a bit rate control parameter corresponding to a hyperparameter in the loss function to obtain a compression model corresponding to the hyperparameter. Based on the loss function, the compression model corresponding to the hyperparameter is trained.
[0014] In a second aspect, the present application provides a model training device, which may include:
[0015] The acquisition module is used to acquire multiple hyperparameters; wherein the multiple hyperparameters correspond to multiple compression bit rates in the target bit rate interval, the target bit rate interval includes multiple sub-bit rate intervals, the difference between the lengths of each sub-bit rate interval is less than a preset length threshold, and the difference between the number of hyperparameters corresponding to each sub-bit rate interval is less than a preset number threshold. The processing module is used to construct multiple loss functions based on the multiple hyperparameters, and one hyperparameter corresponds to at least one loss function. The processing module is also used to train the compression model based on the multiple loss functions to obtain the trained compression model.
[0016] In a possible design, the acquisition module is further used to acquire a first hyperparameter and a second hyperparameter, wherein the first hyperparameter is the largest hyperparameter among multiple hyperparameters, and the second hyperparameter is the smallest hyperparameter among multiple hyperparameters. The processing module is further used to process the first hyperparameter and the second hyperparameter in an exponential interpolation manner to obtain multiple hyperparameters.
[0017] In a possible design, the processing module is further used to determine the sampling probability of each of the multiple hyperparameters based on the sizes of the multiple hyperparameters, the sampling probability of each hyperparameter is positively correlated with the size of the hyperparameter, and the size of the hyperparameter is positively correlated with the size of the compression code rate corresponding to the hyperparameter. The processing module is also used to construct multiple loss functions based on the multiple hyperparameters and the sampling probability of each hyperparameter.
[0018] In a possible design, the processing module is further used to divide the multiple hyperparameters into multiple parameter sets according to the sizes of the multiple hyperparameters, and each parameter set corresponds to a sampling probability. The processing module is also used to determine the sampling probability of each hyperparameter based on the parameter set to which each hyperparameter belongs.
[0019] In one possible design, the processing module is further used to construct a loss function based on a target operation for each loss function to construct multiple loss functions, wherein the target operation includes: determining a target hyperparameter based on multiple hyperparameters and a sampling probability of each hyperparameter, wherein the target hyperparameter is any hyperparameter among the multiple hyperparameters. Obtain sample data in a training set, and input the sample data into a compression model to obtain a distortion loss and a bit rate loss. Construct a loss function based on the target hyperparameter, the bit rate loss, and the distortion loss.
[0020] In one possible design, the processing module is also used to train the compression model based on multiple loss functions to obtain a target loss value until the target loss value meets a preset loss condition, stop training, and obtain a trained compression model, where the target loss value is the sum of the loss values corresponding to the multiple loss functions.
[0021] In one possible design, a hyperparameter corresponds to a set of rate control parameters. The processing module is further used to set a compression model for each loss function based on the rate control parameter corresponding to the hyperparameter in the loss function to obtain a compression model corresponding to the hyperparameter. The processing module is further used to train the compression model corresponding to the hyperparameter based on the loss function.
[0022] In a third aspect, the present application provides a model training device, which includes: a processor and a memory; the processor and the memory are coupled; the memory is used to store one or more programs, and the one or more programs include computer execution instructions. When the model training device is running, the processor executes the computer execution instructions stored in the memory to implement the method described in the first aspect and any possible implementation of the first aspect.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed on a computer, the computer executes the method described in the above-mentioned first aspect and any possible implementation of the first aspect.
[0024] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed by a computer, enables the computer to execute the method described in the above-mentioned first aspect and any possible implementation manner of the first aspect.
[0025] In the above scheme, the technical problems that can be solved and the technical effects achieved by the model training device, computer storage medium or computer program product can refer to the technical problems and technical effects solved by the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1A schematic diagram of a compression performance example of a compression model provided in an embodiment of the present application;
[0027] Figure 2 A flowchart of a model training method provided in an embodiment of the present application;
[0028] Figure 3 A schematic diagram of another compression performance example of a compression model provided in an embodiment of the present application;
[0029] Figure 4 A flowchart of another model training method provided in an embodiment of the present application;
[0030] Figure 5 A schematic diagram of another compression performance example of a compression model provided in an embodiment of the present application;
[0031] Figure 6 A schematic diagram of an example of compressed data provided in an embodiment of the present application;
[0032] Figure 7 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;
[0033] Figure 8 A schematic diagram of the structure of another model training device provided in an embodiment of the present application;
[0034] Fig. 9 A conceptual partial view of a computer program product provided for an embodiment of the present application. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0036] In this article, the character " / " generally indicates that the objects before and after are in an "or" relationship. For example, A / B can be understood as A or B.
[0037] The terms “first” and “second” in the description and claims of the present application are used to distinguish different objects rather than to describe a specific order of the objects.
[0038] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, but may optionally include other steps or modules that are not listed, or may optionally include other steps or modules that are inherent to these processes, methods, products or devices.
[0039] In addition, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present concepts in a specific way.
[0040] At present, compression models based on neural networks (such as image compression models, video compression models, etc.) usually adopt an encoder-decoder structure, which has good performance. Exemplarily, the compression model of the encoder-decoder structure can be expressed by formula 1, formula 2, and formula 3.
[0041] y=E(x,θ e )Formula 1.
[0042]
[0043] Among them, x is used to represent the data to be compressed, θ e is used to represent the encoder parameters, y is used to represent the latent variable to be compressed, Q(·) is used to represent the quantizer, Used to represent the quantized latent variables, θ d Used to represent decoder parameters, Used to represent decompressed data.
[0044] In order to improve the performance of the compression model, an objective function can usually be used to optimize the compression model. Exemplarily, the objective function can be expressed by formula 4.
[0045] loss=R+λD Formula 4
[0046] in, Used to indicate bit rate loss, It is used to represent the distortion loss, and the hyperparameter λ is used to represent the trade-off factor. The larger the λ, the smaller the distortion loss and the larger the bitrate loss. Conversely, the smaller the λ, the larger the distortion loss and the smaller the bitrate loss.
[0047] It should be noted that code rate loss can also be called bit rate loss, which is used to reflect the difference in data volume between the data to be compressed and the compressed data. Distortion loss is used to reflect the difference between the data to be compressed and the compressed data.
[0048] However, the above compression model usually uses a fixed hyperparameter λ to train a compression model of one bit rate. Therefore, if you need to obtain compression models at different bit rates, you need to repeatedly train the compression model of each bit rate using different λ values, which leads to a long training process. At the same time, in order to implement the use of compression models with different bit rates, you need to save the compression models corresponding to different bit rates in actual use, which takes up a lot of storage space.
[0049] Currently, a variable bit rate compression model can be used to compress data using different bit rates. For example, the variable bit rate compression model can be expressed by Formula 5, Formula 6, and Formula 7.
[0050]
[0051] in, Used to represent the rate control parameter of the encoder corresponding to the i-th hyperparameter, is used to represent the rate control parameter of the decoder corresponding to the i-th hyperparameter.
[0052] In order to improve the performance of the variable bit rate compression model, the compression model can usually be optimized by using the objective function corresponding to the variable bit rate compression model. Exemplarily, the objective function corresponding to the variable bit rate compression model can be expressed by Formula 8.
[0053]
[0054] Among them, N is used to represent the number of hyperparameters, and N is a positive integer.
[0055] However, in the process of training the variable bit rate compression model, the selection of the λ value will affect the bit rate that can be actually used by the variable bit rate compression model (also called the code point of the variable bit rate compression model). At present, multiple λ values can usually be obtained by linear interpolation. However, the λ obtained by linear interpolation often causes the code points to be concentrated at high code points, that is, the actually usable bit rates are mostly distributed in the high bit rate range, and the usable bit rates in the low bit rate range are relatively small. In this way, the usable bit rates are unevenly distributed, which affects the performance of the variable bit rate model in the low bit rate range.
[0056] For example, Figure 1As shown, the bit rate can be represented by the percentage of each pixel. The bit rate interval can be divided into bit rate interval 1 [0.05, 0.45], bit rate interval 2 (0.45, 0.58], and bit rate interval 3 (0.58, 0.65]. The multiple bit rate intervals can correspond to multiple hyperparameters, and each hyperparameter has an index value. If the multiple index values include 0-63, that is, 64 index values correspond to 64 hyperparameters. The 64 hyperparameters are obtained by linear interpolation, then the index value corresponding to bit rate interval 1 is [0, 21], the index value corresponding to bit rate interval 2 is [22, 42], and the index value corresponding to bit rate interval 3 is [43, 63]. It can be seen that the bit rates included in bit rate interval 1 are greater than those in bit rate interval 2 and bit rate interval 3, but the number of available bit rates in the intervals is the same. In this way, some bit rates in bit rate interval 1 may not belong to the bit rates corresponding to the hyperparameters, resulting in the inability to use the bit rates for compression, affecting the compression performance.
[0057] In order to solve the above problems, an embodiment of the present application provides a model training method. In the method, multiple hyperparameters can be obtained, and the multiple hyperparameters correspond to multiple compression bit rates in the target bit rate interval. The target bit rate interval includes multiple sub-bit rate intervals, and the difference between the lengths of each sub-bit rate interval is less than the preset length threshold. In other words, the length of each sub-bit rate interval is roughly similar. In addition, the difference between the number of hyperparameters corresponding to each sub-bit rate interval is less than the preset number threshold, that is, the number of hyperparameters corresponding to each sub-bit rate interval is roughly the same. In addition, since one hyperparameter corresponds to one compression bit rate, it means that the number of compression bit rates included in each sub-bit rate interval is roughly the same, that is, the compression bit rates can be evenly distributed in the target bit rate interval. Afterwards, multiple loss functions can be constructed based on multiple hyperparameters, and one hyperparameter corresponds to at least one loss function. Then, the compression model can be trained based on multiple loss functions to obtain the trained compression model. In this way, there is no need to train the compression model under different compression bit rates, which can reduce the number of compression model trainings, thereby simplifying the steps of obtaining the variable bit rate compression model. Furthermore, since the compression bit rates can be evenly distributed in the target bit rate range, the number of available compression bit rates in different bit rate ranges is similar, thereby improving the performance of the variable bit rate model in the low bit rate range.
[0058] The execution subject of the training method of the model provided in the present application can be a training device of the model, and the training device can be a server. At the same time, the training device can also be a central processing unit (CPU) of the server, or a training module for training the model in the server. In the embodiment of the present application, the training method of the model executed by the server is taken as an example to illustrate the training method of the model provided in the embodiment of the present application.
[0059] It should be noted that the server may be a single physical server, or a server cluster consisting of multiple servers. Alternatively, the server cluster may also be a distributed cluster. Alternatively, the server may be a cloud server. The embodiments of the present application do not limit the specific implementation of the server.
[0060] like Figure 2 As shown, a model training method provided in an embodiment of the present application includes:
[0061] S201. Obtain multiple hyperparameters.
[0062] Among them, multiple hyperparameters correspond to multiple compression bit rates in the target bit rate range. One hyperparameter corresponds to one compression bit rate, and the compression bit rate is used to reflect the size of the compressed data.
[0063] Exemplarily, the multiple hyperparameters include: hyperparameter 1 to hyperparameter 64. Then the target bit rate range includes the compression bit rate corresponding to each hyperparameter in hyperparameter 1 to hyperparameter 64.
[0064] In an embodiment of the present application, the multiple hyperparameters include: a first hyperparameter and a second hyperparameter, the first hyperparameter is the largest hyperparameter among the multiple hyperparameters, and the second hyperparameter is the smallest hyperparameter among the multiple hyperparameters. The target bit rate interval includes: a maximum bit rate and a minimum bit rate. The maximum bit rate of the target bit rate interval can be determined based on the first hyperparameter, and the minimum bit rate of the target bit rate interval can be determined based on the second hyperparameter.
[0065] In a possible design, the target bitrate interval includes multiple sub-bitrate intervals, and the difference between the lengths of the sub-bitrate intervals is less than a preset length threshold. The difference between the number of hyperparameters corresponding to the sub-bitrate intervals is less than a preset number threshold.
[0066] That is to say, the lengths of the sub-rate intervals are relatively close (or the same), and the number of hyperparameters corresponding to the sub-rate intervals is relatively close (or the same). Moreover, one hyperparameter corresponds to one compression bit rate, indicating that the number of compression bit rates included in each sub-rate interval is relatively close (or the same).
[0067] It should be noted that the length of the sub-rate interval can be determined based on the two endpoints of the sub-rate interval. The embodiment of the present application does not limit the preset length threshold. For example, the preset length threshold may be 0.1, 0.2, 0.3, etc. Moreover, the embodiment of the present application does not limit the preset quantity threshold. For example, the preset quantity threshold may be 1, 2, 4, etc.
[0068] For example, if the target bitrate interval is [0.05-0.55], the sub-bitrate interval 1 is [0.05-0.15], the sub-bitrate interval 2 is (0.15-0.33], and the sub-bitrate interval 3 is (0.33-0.55]. Among them, the number of hyperparameters corresponding to the sub-bitrate interval 1 is 20, the number of hyperparameters corresponding to the sub-bitrate interval 2 is 21, and the number of hyperparameters corresponding to the sub-bitrate interval 3 is 22.
[0069] In a possible implementation, a first hyperparameter and a second hyperparameter are obtained, and then the first hyperparameter and the second hyperparameter are processed in an exponential interpolation manner to obtain a plurality of hyperparameters.
[0070] In a possible design, the first hyperparameter and the second hyperparameter may be stored in a training device of the model. Each of the multiple hyperparameters may satisfy Formula 9.
[0071]
[0072] Among them, λ max is used to denote the first hyperparameter, λ min is used to represent the second hyperparameter, N is used to represent the number of hyperparameters, and i is used to represent the index value of the hyperparameter.
[0073] It can be understood that by processing the first hyperparameter and the second hyperparameter in an exponential interpolation manner, the obtained multiple hyperparameters are distributed more evenly, thereby making the compression code rate of the compression model evenly distributed in each sub-code rate interval.
[0074] S202. Construct multiple loss functions based on multiple hyperparameters.
[0075] Among them, one hyperparameter corresponds to at least one loss function.
[0076] In a possible implementation, for multiple hyperparameters, based on the sampling probability, a hyperparameter is determined from the multiple hyperparameters, and a loss function is constructed based on the hyperparameter to obtain multiple loss functions.
[0077] In a possible design, the sampling probabilities of the various hyperparameters may be the same, or the sampling probabilities of the various hyperparameters may be different.
[0078] In one possible design, the loss function may satisfy Formula 4.
[0079] In another possible implementation, for each loss function, a loss function can be constructed based on a target operation to construct multiple loss functions, wherein the target operation includes: determining a target hyperparameter based on multiple hyperparameters and a sampling probability of each hyperparameter, wherein the target hyperparameter is any one of the multiple hyperparameters. Afterwards, sample data in the training set is obtained, and the sample data is input into a compression model to obtain a distortion loss and a bit rate loss. Then, a loss function is constructed based on the target hyperparameter, the bit rate loss, and the distortion loss.
[0080] It should be noted that for each loss function, different bit rate losses and distortion losses can be obtained by training a compression model once. Different bit rate losses and distortion losses can be used together with any hyperparameter to construct a loss function. In other words, a loss function corresponds to a hyperparameter, bit rate loss, and distortion loss. Different loss functions may have the same hyperparameters, but different bit rate losses and distortion losses.
[0081] It is understandable that the target hyperparameter can be determined from multiple hyperparameters through sampling probability, and then a loss function can be constructed based on the target hyperparameter, bit rate loss and distortion loss. In this way, the probability of training compression models with different compression bit rates can be controlled, thereby improving the performance of the variable bit rate compression model.
[0082] S203: Train the compression model based on multiple loss functions to obtain a trained compression model.
[0083] Among them, the trained compression model is a variable bit rate compression model.
[0084] In one possible implementation, the compression model can be trained based on multiple loss functions to obtain a target loss value. The target loss value is stopped until it meets a preset loss condition, and the trained compression model is obtained. The target loss value is the sum of the loss values corresponding to the multiple loss functions.
[0085] In one possible design, the multiple loss functions may satisfy Formula 8, and the target loss value may satisfy Formula 8.
[0086] It should be noted that the specific setting of the preset loss condition can refer to the setting of conditions when optimizing and training models using loss functions in conventional technologies, and the present application embodiment does not limit this. For example, the preset loss condition can be that the target loss value is less than a preset loss threshold.
[0087] It is understandable that the target loss value is the sum of the loss values corresponding to multiple loss functions, and the compression model trained by multiple loss functions can be optimized by presetting the loss condition, so as to improve the performance of the variable bit rate compression model.
[0088] For example, Figure 3 As shown, the horizontal axis is the bit rate, which can be expressed by the percentage of each pixel; the vertical axis is the peak signal-to-noise ratio, which is used to reflect the distortion loss. The code points of the trained compression model (that is, the usable bit rates) can be evenly distributed in the sub-bit rate interval 1 [0.05-0.15], the sub-bit rate interval 2 (0.15-0.33], and the sub-bit rate interval 3 (0.33-0.55). Among them, the number of usable compression bit rates corresponding to sub-bit rate interval 1 is 20, the number of compression bit rates corresponding to sub-bit rate interval 2 is 21, and the number of compression bit rates corresponding to sub-bit rate interval 3 is 22.
[0089] Based on the above technical solution, it can be known that multiple hyperparameters correspond to multiple compression bit rates in the target bit rate interval, and the target bit rate interval includes multiple sub-bit rate intervals, and the difference between the lengths of each sub-bit rate interval is less than the preset length threshold. In other words, the length of each sub-bit rate interval is roughly similar. In addition, the difference between the number of hyperparameters corresponding to each sub-bit rate interval is less than the preset number threshold, that is, the number of hyperparameters corresponding to each sub-bit rate interval is roughly the same. In addition, since one hyperparameter corresponds to one compression bit rate, it means that the number of compression bit rates included in each sub-bit rate interval is roughly the same, that is, the compression bit rates can be evenly distributed in the target bit rate interval. Afterwards, multiple loss functions can be constructed based on multiple hyperparameters, and one hyperparameter corresponds to at least one loss function. Then, the compression model can be trained based on multiple loss functions to obtain a trained compression model. In this way, there is no need to train the compression model under different compression bit rates, which can reduce the number of compression model trainings, thereby simplifying the steps of obtaining a variable bit rate compression model. Furthermore, since the compression bit rates can be evenly distributed in the target bit rate range, the number of available compression bit rates in different bit rate ranges is similar, thereby improving the performance of the variable bit rate model in the low bit rate range.
[0090] It should be noted that, generally, the greater the compression bit rate, the smaller the improvement in performance of the compression model obtained through training. In the embodiment of the present application, the number of training times for the high bit rate compression model can be increased to improve the performance of the high bit rate compression model.
[0091] like Figure 4 As shown, a model training method provided in an embodiment of the present application is provided. In this method, S202 may include:
[0092] S401. Determine the sampling probability of each of the multiple hyperparameters based on the sizes of the multiple hyperparameters.
[0093] Among them, the sampling probability of each hyperparameter is positively correlated with the size of the hyperparameter, and the size of the hyperparameter is positively correlated with the size of the compression code rate corresponding to the hyperparameter.
[0094] In other words, the higher the compression bit rate corresponding to the hyperparameter, the greater the sampling probability of the hyperparameter. The lower the compression bit rate corresponding to the hyperparameter, the smaller the sampling probability of the hyperparameter. That is, the sampling probability of the hyperparameter corresponding to a high bit rate is greater than the sampling probability of the hyperparameter corresponding to a low bit rate.
[0095] In a possible implementation, a sampling probability is set for each hyperparameter according to the size of the multiple hyperparameters, wherein the sampling probabilities of the hyperparameters are different, or the sampling probabilities of some hyperparameters are the same.
[0096] In a possible implementation, multiple hyperparameters are divided into multiple parameter sets according to their sizes, and each parameter set corresponds to a sampling probability. Then, based on the parameter set to which each hyperparameter belongs, the sampling probability of each hyperparameter is determined.
[0097] That is to say, the sampling probabilities corresponding to the hyperparameters in a parameter set are the same.
[0098] In a possible design, multiple hyperparameters are sorted according to their magnitude and divided into multiple parameter sets, wherein the average values of the hyperparameters corresponding to the multiple parameter sets are increased in sequence.
[0099] Exemplarily, multiple hyperparameters include 0.1-0.64, parameter set 1 includes 0.1-0.16, parameter set 2 includes 0.17-0.32, parameter set 3 includes 0.33-0.48, and parameter set 1 includes 0.49-0.64. The sampling probability corresponding to parameter set 1 is 10%, the sampling probability corresponding to parameter set 2 is 20%, the sampling probability corresponding to parameter set 3 is 30%, and the sampling probability corresponding to parameter set 4 is 40%.
[0100] For another example, when selecting the bitrate index value of each batch, stratified sampling is used instead of uniform sampling. The bitrate index value i∈[0,N-1] can be evenly divided into multiple sets (such as 4 sets), and the sampling probability value is kept consistent within each set, but increases between sets. For example, the sampling probability corresponding to set 1 is 10%, the sampling probability corresponding to set 2 is 20%, the sampling probability corresponding to set 3 is 30%, and the sampling probability corresponding to set 4 is 50%.
[0101] It should be noted that the present embodiment does not limit the sampling probability between the various sets. For example, the sampling probability between the sets may increase exponentially.
[0102] It can be understood that by dividing multiple hyperparameters into multiple parameter sets and setting a sampling probability for each parameter set, the sampling probability of each hyperparameter can be determined based on the parameter set in which each hyperparameter is located, thereby controlling the probability of training compression models with different compression bit rates, thereby improving the performance of the variable bit rate compression model.
[0103] For example, Figure 5 As shown, the dotted line 501 is used to represent the performance of the compression model when uniform sampling is used, and the solid line 502 is used to represent the performance of the compression model when stratified sampling is used (i.e., in this application). Figure 5 It can be seen that the performance of the compression model obtained through stratified sampling training is greatly improved at high bit rates (such as bit rates greater than 0.6).
[0104] S402: Construct multiple loss functions based on multiple hyperparameters and the sampling probability of each hyperparameter.
[0105] In a possible implementation, based on the sampling probability of each hyperparameter, a hyperparameter is determined from multiple hyperparameters, and a loss function is constructed based on the hyperparameter to obtain multiple loss functions.
[0106] Based on the above technical solution, it can be known that the sampling probability of each hyperparameter is positively correlated with the size of the hyperparameter, and the size of the hyperparameter is positively correlated with the size of the compression bit rate corresponding to the hyperparameter. In other words, the sampling probability of the hyperparameter corresponding to the high bit rate is higher. Afterwards, based on multiple hyperparameters and the sampling probability of each hyperparameter, multiple loss functions are constructed, and there may be more hyperparameters corresponding to the high bit rate in the multiple loss functions. In this way, the number of training times of the variable bit rate compression model for high bit rate compression can be increased, thereby improving the performance of the variable bit rate compression model in compressing data at high bit rates.
[0107] It should be noted that, in the process of adjusting the compression model, the bit rate and performance of the compression model can be adjusted by adjusting the bit rate control parameters of the compression model.
[0108] In some embodiments, a hyperparameter corresponds to a set of rate control parameters. Training the compression model based on the loss function may include: for each loss function, setting the compression model based on the rate control parameter corresponding to the hyperparameter in the loss function to obtain the compression model corresponding to the hyperparameter. Then, based on the loss function, training the compression model corresponding to the hyperparameter.
[0109] A set of rate control parameters includes encoding control parameters and decoding control parameters.
[0110] For example, during the training process, an optimizable rate control parameter table (also referred to as codec control parameter) may be maintained. in, It is used to represent the encoding control parameter corresponding to the hyperparameter with index value i. It is used to indicate the decoding control parameter corresponding to the hyperparameter with index value i. The conditional codec can query the bit rate index value i according to the input bit rate index value i. and And used as input for the compression model condition.
[0111] It should be noted that, in the process of training the compression model, the bit rate control parameters corresponding to each hyperparameter can be continuously adjusted so that the trained compression model meets the preset loss conditions.
[0112] It can be understood that by setting a set of rate control parameters for a hyperparameter, when training the compression model, the compression model can be set based on the rate control parameters corresponding to the hyperparameters in the loss function to obtain a compression model corresponding to the hyperparameters. After that, the compression model can be trained based on the loss function to obtain a variable rate compression model.
[0113] The following is an introduction to the process of constructing a loss function with specific examples. Figure 6 As shown, the data to be compressed x is input into the conditional encoder E to obtain the latent variable y to be compressed. At the same time, the conditional encoder uses q e The condition controls the output of the latent variable y to be compressed. After that, the latent variable y to be compressed is decompressed by the quantizer Q and the autoregressive module A, and the hyper-prior part on the right. Among them, in the super prior part, the latent variable y to be compressed can be input into the conditional encoder E to obtain a new latent variable z to be compressed, and the new latent variable z to be compressed is input from the conditional encoder E to the conditional decoder D through the quantizer Q to obtain the decompressed latent variable corresponding to the new latent variable z to be compressed Then, unpack the latent variables After the conditional decoder D, the compressed data is obtained The conditional decoder uses q d Conditional Control The bit rate loss can be achieved by compressing the latent variable y and decompressing the latent variable The distortion loss can be obtained by the data to be compressed x and the compressed data get.
[0114] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. It is understandable that in order to realize the above functions, the training device of the model includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the training method steps of the models of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0115] The embodiment of the present application can divide the training device of the model into functional modules or functional units according to the above method example. For example, each functional module or functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of software functional modules or functional units. Among them, the division of modules or units in the embodiment of the present application is schematic, which is only a logical functional division, and there may be other division methods in actual implementation.
[0116] like Figure 7 FIG. 1 is a schematic diagram of a structure of a model training device provided in an embodiment of the present application. The model training device is used to execute Figure 2 and Figure 4 The training method of the model shown in FIG.
[0117] The acquisition module 701 is used to acquire multiple hyperparameters; wherein the multiple hyperparameters correspond to multiple compression bit rates in the target bit rate interval, the target bit rate interval includes multiple sub-bit rate intervals, the difference between the lengths of each sub-bit rate interval is less than a preset length threshold, and the difference between the number of hyperparameters corresponding to each sub-bit rate interval is less than a preset number threshold. The processing module 702 is used to construct multiple loss functions based on the multiple hyperparameters, and one hyperparameter corresponds to at least one loss function. The processing module 702 is also used to train the compression model based on the multiple loss functions to obtain the trained compression model.
[0118] In a possible design, the acquisition module 701 is further used to obtain a first hyperparameter and a second hyperparameter, where the first hyperparameter is the largest hyperparameter among multiple hyperparameters, and the second hyperparameter is the smallest hyperparameter among multiple hyperparameters. The processing module 702 is further used to process the first hyperparameter and the second hyperparameter in an exponential interpolation manner to obtain multiple hyperparameters.
[0119] In a possible design, processing module 702 is further used to determine the sampling probability of each of the multiple hyperparameters based on the sizes of the multiple hyperparameters, the sampling probability of each hyperparameter is positively correlated with the size of the hyperparameter, and the size of the hyperparameter is positively correlated with the size of the compression code rate corresponding to the hyperparameter. Processing module 702 is also used to construct multiple loss functions based on the multiple hyperparameters and the sampling probability of each hyperparameter.
[0120] In one possible design, the processing module 702 is further used to divide the multiple hyperparameters into multiple parameter sets according to the sizes of the multiple hyperparameters, and each parameter set corresponds to a sampling probability. The processing module 702 is also used to determine the sampling probability of each hyperparameter based on the parameter set to which each hyperparameter belongs.
[0121] In one possible design, the processing module 702 is further used to construct a loss function based on a target operation for each loss function to construct multiple loss functions, wherein the target operation includes: determining a target hyperparameter based on multiple hyperparameters and a sampling probability of each hyperparameter, wherein the target hyperparameter is any hyperparameter among the multiple hyperparameters. Obtain sample data in a training set, and input the sample data into a compression model to obtain a distortion loss and a bit rate loss. Construct a loss function based on the target hyperparameter, the bit rate loss, and the distortion loss.
[0122] In one possible design, the processing module 702 is also used to train the compression model based on multiple loss functions to obtain a target loss value until the target loss value meets a preset loss condition, stop training, and obtain a trained compression model, where the target loss value is the sum of the loss values corresponding to the multiple loss functions.
[0123] In one possible design, a hyperparameter corresponds to a set of rate control parameters. Processing module 702 is further used to set a compression model for each loss function based on the rate control parameter corresponding to the hyperparameter in the loss function to obtain a compression model corresponding to the hyperparameter. Processing module 702 is further used to train the compression model corresponding to the hyperparameter based on the loss function.
[0124] Figure 8 8 is a schematic diagram of a structure of a model training device according to an exemplary embodiment. The model training device may include a processor 802, and the processor 802 is used to execute application code to implement the model training method in this application.
[0125] The processor 802 may be a central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application.
[0126] like Figure 8 As shown, the model training device may further include a memory 803. The memory 803 is used to store application program codes for executing the solution of the present application, and the execution is controlled by the processor 802.
[0127] The memory 803 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 803 may exist independently and be connected to the processor 802 via the bus 804. The memory 803 may also be integrated with the processor 802.
[0128] like Figure 8 As shown, the training device of the model may further include a communication interface 801, wherein the communication interface 801, the processor 802, and the memory 803 may be coupled to each other, for example, via a bus 804. The communication interface 801 is used to perform information interaction with other devices, for example, to support information interaction between the training device of the model and other devices.
[0129] It should be pointed out that Figure 8 The equipment structure shown in the figure does not constitute a limitation on the training device of the model, except Figure 8 In addition to the components shown, the training device of the model may include more or fewer components than shown, or combine certain components, or have a different arrangement of components.
[0130] In actual implementation, the functions implemented by the processing module 702 can be Figure 8 The processor 802 shown calls the program code in the memory 803 to implement.
[0131] The present application also provides a computer-readable storage medium having instructions stored thereon, and when the instructions in the computer-readable storage medium are executed by a processor of a computer device, the computer is enabled to execute the training method of the model provided in the above-mentioned embodiment. For example, the computer-readable storage medium may be a memory 803 including instructions, and the above instructions may be executed by a processor 802 of a computer device to complete the above method. Optionally, the computer-readable storage medium may be a non-temporary computer-readable storage medium, for example, a non-temporary computer-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0132] Fig. 9 A conceptual partial view of a computer program product provided by an embodiment of the present application is schematically shown, where the computer program product includes a computer program for executing a computer process on a computing device.
[0133] In one embodiment, the computer program product is provided using a signal bearing medium 900. The signal bearing medium 900 may include one or more program instructions that, when executed by one or more processors, may provide the above-described Figure 2 Thus, for example, reference to Figure 2 In the embodiment shown in , one or more features of S201-S203 may be undertaken by one or more instructions associated with the signal bearing medium 900. In addition, Fig. 9 The program instructions in also describe example instructions.
[0134] In some examples, the signal bearing medium 900 may include a computer readable medium 901, such as, but not limited to, a hard drive, a compact disk (CD), a digital video disk (DVD), a digital tape, a memory, a read-only memory (ROM), or a random access memory (RAM), and the like.
[0135] In some implementations, signal bearing medium 900 may include computer recordable medium 902 such as, but not limited to, memory, read / write (R / W) CD, R / W DVD, and the like.
[0136] In some implementations, signal bearing medium 900 may include communication medium 903 such as, but not limited to, digital and / or analog communication media (eg, fiber optic cables, waveguides, wired communication links, wireless communication links, etc.).
[0137] The signal bearing medium 900 may be communicated by a wireless form of communication medium 903. The one or more program instructions may be, for example, computer executable instructions or logic implemented instructions.
[0138] In some examples, such as for Fig. 9 The described model training apparatus may be configured to provide various operations, functions, or actions in response to one or more program instructions via computer-readable medium 901 , computer-recordable medium 902 , and / or communication medium 903 .
[0139] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0140] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0141] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0142] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0143] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or the full classification part or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions for a device or processor to execute the full classification part or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk.
[0144] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto, and any changes or substitutions within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A model training method, characterized in that: The method comprises: Acquire multiple hyperparameters; wherein the multiple hyperparameters correspond to multiple compression bit rates in a target bit rate interval, the target bit rate interval includes multiple sub-bit rate intervals, the difference between the lengths of the sub-bit rate intervals is less than a preset length threshold, and the difference between the number of hyperparameters corresponding to the sub-bit rate intervals is less than a preset number threshold; Based on the multiple hyperparameters, construct multiple loss functions, where one hyperparameter corresponds to at least one loss function; The compression model is trained based on the multiple loss functions to obtain a trained compression model.
2. The method according to claim 1, characterized in that The obtaining of multiple hyper parameters includes: Obtaining a first hyperparameter and a second hyperparameter, wherein the first hyperparameter is the largest hyperparameter among the multiple hyperparameters, and the second hyperparameter is the smallest hyperparameter among the multiple hyperparameters; The first hyperparameter and the second hyperparameter are processed in an exponential interpolation manner to obtain the multiple hyperparameters.
3. The method according to claim 1, characterized in that The constructing of multiple loss functions based on the multiple hyperparameters includes: Based on the sizes of the multiple hyperparameters, determining the sampling probability of each of the multiple hyperparameters, the sampling probability of each of the hyperparameters is positively correlated with the size of the hyperparameter, and the size of the hyperparameter is positively correlated with the size of the compression code rate corresponding to the hyperparameter; The multiple loss functions are constructed based on the multiple hyperparameters and the sampling probability of each hyperparameter.
4. The method according to claim 3, characterized in that The step of determining the sampling probability of each hyperparameter based on the magnitudes of the multiple hyperparameters includes: According to the sizes of the multiple hyperparameters, the multiple hyperparameters are divided into multiple parameter sets, and each parameter set corresponds to one sampling probability; Based on the parameter set in which each hyperparameter belongs, a sampling probability of each hyperparameter is determined.
5. The method according to claim 3, characterized in that: The constructing the multiple loss functions based on the multiple hyperparameters and the sampling probability of each hyperparameter includes: For each loss function, constructing the loss function based on a target operation to construct the plurality of loss functions, the target operation comprising: Determine a target hyperparameter based on the multiple hyperparameters and the sampling probability of each hyperparameter, wherein the target hyperparameter is any hyperparameter among the multiple hyperparameters; Obtaining sample data in a training set, and inputting the sample data into the compression model to obtain distortion loss and bit rate loss; The loss function is constructed based on the target hyperparameter, the bit rate loss and the distortion loss.
6. The method according to claim 1, characterized in that The step of training the compression model based on the multiple loss functions to obtain the trained compression model includes: The compression model is trained based on the multiple loss functions to obtain a target loss value. The target loss value is stopped until it meets a preset loss condition to obtain the trained compression model, and the target loss value is the sum of the loss values corresponding to the multiple loss functions.
7. The method according to claim 1, characterized in that A hyperparameter corresponds to a set of rate control parameters; the compression model is trained based on the loss function, including: For each loss function, setting the compression model based on a rate control parameter corresponding to a hyperparameter in the loss function to obtain a compression model corresponding to the hyperparameter; Based on the loss function, the compression model corresponding to the hyperparameter is trained.
8. A model training device, characterized in that: The device comprises: An acquisition module, configured to acquire a plurality of hyperparameters; wherein the plurality of hyperparameters correspond to a plurality of compression bit rates in a target bit rate interval, the target bit rate interval includes a plurality of sub-bit rate intervals, the difference between the lengths of the sub-bit rate intervals is less than a preset length threshold, and the difference between the number of hyperparameters corresponding to the sub-bit rate intervals is less than a preset number threshold; A processing module, configured to construct a plurality of loss functions based on the plurality of hyperparameters, wherein one hyperparameter corresponds to at least one loss function; The processing module is also used to train the compression model based on the multiple loss functions to obtain a trained compression model.
9. A model training device, characterized in that: include: Memory and processor; Memory and processor coupling; The memory is used to store instructions executable by the processor; When the processor executes the instructions, the method according to any one of claims 1 to 7 is performed.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 7.
11. A computer program product comprising instructions, characterized in that When the instructions are executed by a computer, the computer is caused to perform the method according to any one of claims 1 to 7.