Model training method and device

By inputting the training image into the pretrained autoencoder and the target visual basic model in image/video model training, computing the comparison loss value and optimizing the autoencoder, the problem of poor generation effect after spatial compression coding is solved, and the generation effect of the model is improved.

CN120047765APending Publication Date: 2025-05-27SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202411946235.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

During the training process of image/video model, after spatial compression coding, the generation effect becomes worse, and the higher the encoding dimension, the worse the effect.

Method used

By inputting the training image to the pretrained autoencoder and the target visual basic model, the first and second compressed image encodings are obtained, respectively, the comparison loss value is calculated, and combined with other loss values ​​as the target loss value, for optimizing the pretrained autoencoder.

Benefits of technology

It improves the encoding and codec effect of the autoencoder, obtains more and more accurate original image information, and thus improves the generation effect of the application model containing the autoencoder.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047765A_ABST
    Figure CN120047765A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model training method and device. The method comprises the following steps: inputting a training image into an encoder in a pre-training auto-encoder to obtain a first compressed image code, and inputting the training image into a target visual basic model to obtain a second compressed image code; wherein the pre-training auto-encoder comprises an encoder and a decoder; determining a comparison loss value according to the first compressed image code and the second compressed image code, and taking the comparison loss value and at least one other loss value as a target loss value; wherein at least one other loss value is a loss value reflecting the performance of the pre-trained auto-encoder; and performing training optimization on the pre-training auto-encoder based on the target loss value. According to the scheme, the pre-trained auto-encoder can be trained and optimized by taking the image compression result of the target visual basic model as a reference, so that the encoding and decoding effects of the auto-encoder are improved, and more and more accurate original image information can be obtained through the auto-encoder.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model training, and in particular, to a model training method and apparatus. Background Art

[0002] During the process of image / video model training, in order to save computing resources and facilitate model learning, it is generally necessary to perform spatial compression encoding on the training images used for training the model. For example, for an image with a height * width of 256 * 256 pixels, each pixel is composed of the colors of 3 channels of RGB (red, green, and blue). If spatial compression encoding is not performed and the image is directly input into the model for training, then an image needs to be converted into [256 * 256] encoding vectors, and each vector includes 3 dimensions of RGB, resulting in a huge number of encodings and requiring a large amount of computing power resources to support model training. If spatial compression encoding is performed on the image, the image space can be compressed by 16 times, for example, compressed into [16 * 16] compressed encoding vectors, where each compressed encoding vector can be multiple dimensions such as 16-dimensional, 32-dimensional, 64-dimensional, etc. Compared with before spatial compression encoding, the number of encodings is significantly reduced, reducing the computing power requirement for model training.

[0003] However, after performing spatial compression encoding on the image, the generation effect after training of the image / video model will become worse. After the image undergoes spatial compression encoding, the higher the dimension of the encoding, although it contains more image information, the generation effect of the image / video model will become worse. Summary of the Invention

[0004] Embodiments of this application provide a model training method and apparatus to improve the encoding and decoding effect of the autoencoder, and further improve the generation effect of the actual application model including the autoencoder.

[0005] According to one aspect of this application, a model training method is provided. The method includes:

[0006] Inputting a training image into an encoder in a pre-trained autoencoder to obtain a first compressed image encoding, and inputting the training image into a target visual foundation model to obtain a second compressed image encoding; wherein, the pre-trained autoencoder includes an encoder and a decoder;

[0007] Determining a contrast loss value according to the first compressed image encoding and the second compressed image encoding, and using the contrast loss value and at least one other loss value as a target loss value; wherein, at least one other loss value is a loss value reflecting the performance of the pre-trained autoencoder;

[0008] Training and optimizing the pre-trained autoencoder based on the target loss value.

[0009] According to one aspect of the present application, there is provided a model training device, the device comprising:

[0010] A training image input module, configured to input a training image into an encoder in a pre-trained autoencoder to obtain a first compressed image code, and input the training image into a target vision base model to obtain a second compressed image code; wherein, the pre-trained autoencoder comprises an encoder and a decoder;

[0011] A target loss value determination module, configured to determine a contrast loss value according to the first compressed image code and the second compressed image code, and use the contrast loss value and at least one other loss value as a target loss value; wherein, the at least one other loss value is a loss value reflecting the performance of the pre-trained autoencoder;

[0012] A training optimization module, configured to perform training optimization on the pre-trained autoencoder based on the target loss value.

[0013] According to another aspect of the present application, there is provided an electronic device, the electronic device comprising:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the model training method of any embodiment of the present application.

[0017] According to another aspect of the present application, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the model training method of any embodiment of the present application when executed.

[0018] In the technical solution of the embodiment of the present application, the training image is input into the encoder in the pre-trained autoencoder to obtain the first compressed image code, and the training image is input into the target vision base model to obtain the second compressed image code; wherein, the pre-trained autoencoder includes an encoder and a decoder; according to the first compressed image code and the second compressed image code, a contrast loss value is determined, and the contrast loss value and at least one other loss value are used as the target loss value; wherein, at least one other loss value is a loss value reflecting the performance of the pre-trained autoencoder; the pre-trained autoencoder is trained and optimized based on the target loss value. The above solution can determine the contrast loss value according to the first compressed image code obtained by the autoencoder processing the training image and the second compressed image code obtained by the target vision base model processing the training image, and participate in the training and optimization of the pre-trained autoencoder, so as to train and optimize the autoencoder with reference to the result of the target vision base model for image space compression, improve the encoding and decoding effect of the autoencoder, and be able to obtain more and more accurate original image information through the autoencoder, thereby improving the generation effect of the application model including the autoencoder.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is a flowchart of a model training method provided by an embodiment of the present application;

[0022] Figure 2 It is a flowchart of a model training method provided by another embodiment of the present application;

[0023] Figure 3 It is a schematic structural diagram of a model training device provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0026] It should be noted that the terms "first", "second", "third", "fourth", "actual", "preset", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] Figure 1 It is a flowchart of a model training method provided by an embodiment of this application. The embodiment of this application is applicable to the situation of training an autoencoder. This method can be executed by a model training device, which can be implemented in the form of hardware and / or software, and the model training device can be configured in an electronic device. As Figure 1 shown, this method includes:

[0028] S110. Input the training image into the encoder in the pre-trained autoencoder to obtain the first compressed image code, and input the training image into the target visual foundation model to obtain the second compressed image code; wherein, the pre-trained autoencoder includes an encoder and a decoder.

[0029] Among them, the training images are the images in the training dataset, which can be the images obtained from the server or the images captured in real time by an image collector. The content in the training images is not limited. The number of training images can be determined according to the actual situation, for example, determined according to factors such as the computing power of the model training device, the model training performance requirements, the model structure, and the model size. Both the first compressed image encoding and the second compressed image encoding are spatial compression encodings. That is, for example, a picture with a height * width: [256 * 256] pixels is spatially compressed 16 times to be compressed into [16 * 16] first compressed image encodings and / or second compressed image encodings. The dimension of each first compressed image encoding and / or second compressed image encoding is not limited and can be multiple dimensions such as 16 dimensions, 32 dimensions, 64 dimensions, etc. The pre-trained autoencoder is an autoencoder that needs to be trained and optimized. The pre-trained autoencoder includes an encoder and a decoder. The encoder is used to compress the input data space into a latent space representation, and the decoder is used to reconstruct the original data from the latent space representation. In the embodiments of the present application, it is mainly applied to images. The encoder in the pre-trained autoencoder is used to perform spatial compression encoding on the images, and the decoder in the pre-trained autoencoder is used to recover the images from the spatial compression encoding. The vision foundation model is a model that can understand the content of images through training on a large amount of data and perform fine-grained representation learning in various vision tasks. The vision foundation model is used to extract the image features of the images for spatial compression encoding. Generally, it cannot perform decoding, or although it can perform decoding, since it classifies the image content to extract features, similar to extracting the semantics of a piece of text, it cannot reversibly decode the spatial compression encoding into the original target image, which means it cannot accurately decode the spatial compression encoding into the target image to be generated. Therefore, the vision foundation model generally cannot be directly applied to image / video models. The pre-trained autoencoder has encoding and decoding functions and is generally widely applied to image / video models. The vision foundation model is a model that has been trained well and can be one or more models that can process general images, that is, it can perform spatial compression encoding on images containing any content and achieve ideal results. The target vision foundation can be arbitrarily selected from the trained vision foundation models, or selected according to the performance test results. The vision foundation model can also be a model that has been trained for spatial compression encoding of specific types of images. In this case, the target vision foundation model is the vision foundation model selected according to the actual situation, for example, selected according to the image content type and the image content type applicable to the vision foundation model.

[0030] In an embodiment of the present application, the training image is input into the encoder of the pre-trained autoencoder to obtain the first compressed image encoding, and the training image is input into the target vision foundation model to obtain the second compressed image encoding, so as to obtain two representations of the compressed image encoding of the training image, and the pre-trained autoencoder is trained based on the effect of the image compression encoding reflected by the comparison between the first compressed image encoding and the second compressed image encoding. The beneficial effect of using the target vision foundation model to process the training image to obtain the second compressed image encoding is that the target vision foundation model is trained based on a large amount of training data, has an ideal compression encoding effect in image compression encoding, and can obtain the second compressed image encoding that more accurately reflects the features of the training image, which has a positive optimization effect on the training and optimization of the pre-trained autoencoder.

[0031] In an embodiment of the present application, the process of inputting the training image into the encoder of the pre-trained autoencoder and the process of inputting the training image into the target vision foundation model can be executed separately and sequentially, or can be executed simultaneously in parallel. Regardless of the order of inputting into the encoder of the pre-trained autoencoder and inputting into the target vision foundation model, the first compressed image encoding and the second compressed image encoding both refer to the first compressed image encoding and the second compressed image encoding corresponding to the same training image after processing.

[0032] In an embodiment of the present application, the spatial compression rate of the pre-trained autoencoder and the spatial compression rate of the target vision foundation model, and / or the encoding dimension of the pre-trained autoencoder and the encoding dimension of the target vision foundation model may be different, and the number of encoding vectors and / or the dimension of the encoding vectors in the finally obtained first compressed image encoding and second compressed image encoding may be different. Before performing the subsequent steps, it is necessary to perform consistency processing on the first compressed encoded image and the second compressed image encoding so that the number of encoding vectors and the dimension of the encoding vectors of the first compressed encoded image and the second compressed encoded image are the same.

[0033] In an embodiment of the present application, there may be at least one target vision foundation model, corresponding to generating at least one second compressed image encoding. The subsequent process can be executed respectively for at least one second compressed image encoding and the first compressed image encoding, or the second compressed image encodings can be merged first to obtain one second compressed image encoding, and then the subsequent process is executed with the first image compressed encoding.

[0034] S120. Determine a comparison loss value according to the first compressed image encoding and the second compressed image encoding, and use the comparison loss value and at least one other loss value as the target loss value; wherein, the at least one other loss value is a loss value reflecting the performance of the pre-trained autoencoder.

[0035] Exemplarily, the first compressed image encoding can reflect the effect of the encoder in the pre-trained autoencoder for spatially compressing and encoding the training images, and the second compressed image encoding can reflect the effect of the target vision base model for spatially compressing and encoding the training images. The comparison of the effects of the pre-trained autoencoder and the target vision base model for spatially compressing and encoding the training images respectively can reflect the direction for optimizing the pre-trained autoencoder. Specifically, the loss value can be determined according to the first compressed image encoding and the second compressed image encoding for training and optimizing the pre-trained autoencoder, so that the pre-trained autoencoder can refer to the image encoding effect of the target vision base model for training and optimization, and improve the encoding effect of the pre-trained autoencoder.

[0036] Exemplarily, the contrast loss value can be determined according to the first compressed image encoding and the second compressed image encoding for participating in the training and optimization of the pre-trained autoencoder. Specifically, since the target vision base model has more superior performance in image compression encoding and can understand the content of the image and compress it to obtain an encoding that more accurately reflects the image information, the contrast loss value can be determined based on the differential data between the first compressed image encoding and the second compressed encoding, so that the autoencoder tends to learn the more superior image compression encoding ability in the target vision base model during the training and optimization process.

[0037] In addition, there may be other loss functions that can reflect the performance of the pre-trained autoencoder. Other loss values can be calculated based on the first compressed image encoding and / or the image generated after decoding by the decoder. At least one other loss value can be calculated based on at least one other loss function, and the contrast loss value and at least one other loss value are jointly used as the target loss value. The contrast loss value and at least one other loss value can be separately used as the target loss value to participate in the subsequent training and optimization of the pre-trained autoencoder, or the contrast loss value and at least one other loss value can be operated to obtain the target loss value to participate in the subsequent training and optimization of the pre-trained autoencoder. This operation can be a mutually enhancing operation, that is, the contrast loss value and at least one other loss value are larger than any single loss value after the operation.

[0038] S130. Train and optimize the pre-trained autoencoder based on the target loss value.

[0039] Exemplarily, the pre-trained autoencoder can be trained and optimized based on the target loss value. During the training and optimization process, both the encoder in the autoencoder and the decoder in the autoencoder need to be trained and optimized based on the target loss value. An iteration termination condition can be set. For example, at least one loss value among the comparison loss value and at least one other loss value satisfies a preset loss value threshold, or the performance of the autoencoder reaches a preset performance, or the number of iterations reaches a preset number, etc. The beneficial effect of the above solution is that the target loss value includes the comparison loss value, and the comparison loss value is determined based on the first compressed image encoding and the second compressed image encoding. That is, the pre-trained autoencoder is trained and optimized by referring to the image compression encoding effect of the target visual base model, so as to improve the performance of image compression encoding, and further improve the generation effect of the application model composed of the autoencoder and the generation model that meet the performance requirements.

[0040] In the embodiment of the present application, taking the comparison loss value and at least one other loss value as the target loss value includes:

[0041] Weighting the comparison loss value and at least one other loss value as the target loss value to feedback to the pre-trained autoencoder for training and optimization; wherein, at least one other loss value includes at least one of the original reconstruction loss value, the generative adversarial network loss value, and the divergence loss value.

[0042] Exemplarily, weights can be set for the comparison loss value and at least one other loss value respectively to reflect the participation degrees of the comparison loss value and at least one other loss value in the training and optimization process of the pre-trained autoencoder. The specific values of the weights can be adaptively selected according to the actual situation. The sum of the weight of the comparison loss value and the weights of at least one other loss value can be 1 or not 1. The comparison loss value and at least one other loss value can be weighted, and the weighted comparison loss value and at least one other loss value are feedback to the pre-trained autoencoder for training and optimization.

[0043] In the embodiment of the present application, at least one other loss value can include at least one of the original reconstruction loss value, the generative adversarial network loss value, and the divergence loss value. The original reconstruction loss value is an index used to measure the difference between the generated image and the real image in the image generation or reconstruction task. Common original reconstruction loss functions include mean square error (MSE), mean absolute error (MAE), structural similarity loss (SSIM Loss), etc. The loss value of the generative adversarial network (GAN) is an index used to measure the difference between the image generated by the generator and the real image, and the ability of the discriminator to distinguish between real and fake images. The divergence loss value is used to measure the difference between the approximate posterior distribution and the real posterior distribution of the autoencoder.

[0044] In the technical solution of the embodiment of the present application, the training image is input into the encoder in the pre-trained autoencoder to obtain the first compressed image code, and the training image is input into the target vision foundation model to obtain the second compressed image code; wherein, the pre-trained autoencoder includes an encoder and a decoder; according to the first compressed image code and the second compressed image code, a contrast loss value is determined, and the contrast loss value and at least one other loss value are used as the target loss value; wherein, at least one other loss value is a loss value reflecting the performance of the pre-trained autoencoder; the pre-trained autoencoder is trained and optimized based on the target loss value. The above solution can determine the contrast loss value according to the first compressed image code obtained by the autoencoder processing the training image and the second compressed image code obtained by the target vision foundation model processing the training image, and participate in the training and optimization of the pre-trained autoencoder, so as to train and optimize the autoencoder with reference to the result of the target vision foundation model for image space compression, improve the encoding and decoding effect of the autoencoder, and can obtain more and more accurate original image information through the autoencoder, thereby improving the generation effect of the application model including the autoencoder.

[0045] As a non-limiting implementation manner, determining the contrast loss value according to the first compressed image code and the second compressed image code includes:

[0046] Determining an absolute loss value according to the similarity between the first compressed image code and the second compressed image code; and / or,

[0047] Determining a relative loss value according to the first self-similarity data of the first compressed image code and the second self-similarity data of the second compressed image code;

[0048] Determining the contrast loss value according to the absolute loss value and / or the relative loss value.

[0049] Exemplarily, the similarity between the first compressed image code and the second compressed image code can be determined to reflect the difference between the first compressed image code and the second compressed image code, and the similarity is used as the absolute loss value. The calculation method of the similarity is not limited, and it can be any similarity calculation method between vectors, or the result of the operation after calculating the similarity calculation methods between multiple vectors. The similarity calculation method between vectors can be the Pearson correlation coefficient calculation method, the cosine similarity calculation method, the Euclidean distance calculation method, the Manhattan distance calculation method, etc.

[0050] Exemplarily, the first self-similarity data may also be determined according to the first compressed image encoding, which reflects the similarity of each position element in the first compressed image encoding relative to the global of the first compressed image encoding. The first similarity data may be determined according to the vector similarity between the encoding vectors of each position element in the first compressed image encoding and the encoding vectors of other position elements. The second self-similarity data is determined according to the second compressed image encoding, which reflects the similarity of each position element in the second compressed image encoding relative to the global of the second compressed image encoding. The second self-similarity data may be determined according to the vector similarity between the encoding vectors of each position element in the second compressed image encoding and the encoding vectors of other position elements. The calculation method of the similarity is the same as that described in the above embodiments.

[0051] In the embodiments of the present application, the absolute loss value may be determined according to the similarity between the first compressed image encoding and the second compressed image encoding, and the contrast loss value may be determined according to the absolute loss value. The relative loss value may also be determined according to the first self-similarity data of the first compressed image encoding and the second self-similarity data of the second compressed image encoding, and the contrast loss value may be determined according to the relative loss value. The absolute loss value and the relative loss value may also be determined, and the contrast loss value may be determined according to the absolute loss value and the relative loss value. The beneficial effect of adding the relative loss value to determine the contrast loss value is as follows: The absolute loss value reflects the absolute difference value between the first compressed image encoding and the second compressed image encoding. Only when the two vectors corresponding to each position element in the first compressed image encoding and the second compressed image encoding are mostly similar can it reflect that the first compressed image encoding and the second compressed image encoding are similar. The similarity determination criterion is relatively high, which may lead to the problem of difficult to converge quickly. The relative loss value reflects the difference value of the self-similarity data between the first compressed image encoding and the second compressed image encoding. The self-similarity is the similarity of the position element relative to the global. The first self-similarity data represents the similarity structure inside the first compressed image encoding, and the second self-similarity represents the similarity structure inside the second compressed image encoding. The relative loss value reflects the difference of the similarity structures between the first compressed image encoding and the second compressed image encoding, rather than the absolute difference between the vectors, thereby adaptively reducing the similarity judgment scale and accelerating the convergence speed of the model to a certain extent without losing the training effect.

[0052] As a non-limiting implementation manner, determining the absolute loss value according to the similarity between the first compressed image encoding and the second compressed image encoding includes:

[0053] For each corresponding position element in the first compressed image encoding and the second compressed image encoding, determine the similarity between the first vector corresponding to the same position element in the first compressed image encoding and the second vector corresponding to the second compressed image encoding;

[0054] Determine whether the similarity is lower than a preset similarity threshold;

[0055] Based on the target similarity lower than the preset similarity threshold, obtain the absolute loss value.

[0056] Exemplarily, for each corresponding position element in the first compressed image encoding and the second compressed image encoding, determine the first vector corresponding to the same position element in the first compressed image encoding and the second vector corresponding to the same position element in the second compressed image encoding. For example, assume that the first compressed image encoding is [16*16] encoding vectors and the second compressed image encoding is [16*16] encoding vectors. Then, take the encoding vector corresponding to the first position element in the first compressed image encoding, that is, the (1,1) position element, as the first vector, and take the encoding vector corresponding to the first position element in the second compressed image encoding, that is, the (1,1) position element, as the second vector, and calculate the similarity between the first vector and the second vector. Similarly, execute the above scheme to calculate the similarity for each position element. For example, take the encoding vector corresponding to the (3,4) position element in the first compressed image encoding as the first vector, and take the encoding vector corresponding to the (3,4) position element in the second compressed image encoding as the second vector, and calculate the similarity between the first vector and the second vector.

[0057] The preset similarity threshold can be set in advance to reflect the magnitude of the similarity between the first vector and the second vector. If the similarity between the first vector and the second vector is lower than the preset similarity threshold, it reflects that the similarity between the first vector and the second vector is small and the difference is large, and thus it should be more involved in the training and optimization of the pre-trained autoencoder. The absolute loss value can be determined based on the target similarity lower than the preset similarity threshold and participate in the training and optimization of the pre-trained autoencoder. There may be multiple target similarities corresponding to multiple position elements that are lower than the preset similarity threshold, and then the absolute loss value can be obtained by performing operations on multiple target similarities, such as summation, average of summation, multiplication, average after multiplication, etc.

[0058] As a non-limiting implementation manner, obtaining the absolute loss value based on the target similarity less than the preset similarity threshold includes:

[0059] Determine the first difference between the preset similarity threshold and the similarity;

[0060] Determine the rectified linear unit function value of the first difference corresponding to the same position element, and obtain the absolute loss value based on the average value of the rectified linear unit function values corresponding to each position element.

[0061] Exemplarily, for a similarity less than a preset similarity threshold, the degree of difference between the similarity and the preset similarity threshold is different, and the degree of influence on the absolute loss value is different. The greater the difference between the similarity and the preset similarity threshold, the greater the influence on the absolute loss value. The first difference between the preset similarity threshold and the similarity can be determined, and thus the absolute loss value can be determined based on the first difference. In addition, it is necessary to consider the similarity between the encoding vectors corresponding to the elements at each position in the first compressed image encoding and the second compressed image encoding. The rectified linear function value of the first difference corresponding to the same position element can be determined, and based on the average value of the rectified linear function values corresponding to each position element, the absolute loss value is obtained. The meaning of the rectified linear function is that if the independent variable is greater than 0, the rectified linear function value is the value of the independent variable; if the independent variable is less than or equal to 0, the rectified linear function value is 0. The calculation process of the rectified linear function value of the first difference is that if the first difference corresponding to the position element is greater than 0, the rectified linear function value is the first difference; otherwise, the rectified linear function value is 0, which reflects the contribution of the first difference to the absolute loss value. Only when the first difference is greater than 0, that is, the similarity is less than the preset similarity threshold, will it affect the absolute loss value.

[0062] Exemplarily, the absolute loss value can be calculated based on the following formula:

[0063]

[0064] where l mcos is the absolute loss function, and substituting specific values represents the absolute loss value. a ij is the similarity corresponding to the element at position (i, j). a 0 is the preset similarity threshold. a 0 -a ij is the first difference. ReLU is the rectified linear function. h is the height of the first compressed image encoding and the second compressed image encoding, w is the width of the first compressed image encoding and the second compressed image encoding, and h×w is the number of encoding vectors of the first compressed image encoding and the second compressed image encoding. Assuming that the first compressed image encoding and the second compressed image encoding are [16*16] encoding vectors, then h is 16 and w is 16.

[0065] Specifically, assuming that the cosine similarity is used to calculate the similarity between the first vector and the second vector, the absolute loss value is calculated based on the following formula:

[0066]

[0067] where represents the cosine similarity between the first vector and the second vector, and m1 represents the preset cosine similarity threshold, which is negatively correlated with the preset similarity threshold.

[0068] As a non-limiting implementation, determining a relative loss value based on the first self-similarity data of the first compressed image encoding and the second self-similarity data of the second compressed image encoding includes:

[0069] For each corresponding position element in the first compressed image encoding and the second compressed image encoding, determining a self-similarity deviation value of the first self-similarity data and the second self-similarity data of the same position element;

[0070] Determining whether the self-similarity deviation value exceeds a preset deviation threshold;

[0071] Based on the self-similarity deviation values exceeding the preset similarity threshold, obtaining the relative loss value.

[0072] Exemplarily, for each position element in the first compressed image encoding, the first self-similarity data corresponding to each position element can be determined. For each position element in the second compressed image encoding, the second self-similarity data corresponding to each position element can be determined. The self-similarity deviation value of the first self-similarity data and the second self-similarity data corresponding to the same position element can be calculated, reflecting the deviation of the similarity structure of the position element in the first compressed image encoding in the global context relative to the similarity structure of the position element in the second compressed image encoding in the global context. For example, assume that the first compressed image encoding is [16*16] encoding vectors and the second compressed image encoding is [16*16] encoding vectors. Then, calculate the first self-similarity data of the encoding vector corresponding to the first, i.e., the (1,1) position element in the first compressed image encoding, calculate the second self-similarity data of the encoding vector corresponding to the first, i.e., the (1,1) position element in the second compressed image encoding, and calculate the self-similarity deviation value of the first self-similarity data and the second self-similarity data. Similarly, the above scheme is executed for each position element to calculate the self-similarity deviation value. For example, calculate the first self-similarity data of the encoding vector corresponding to the (3,4) position element in the first compressed image encoding, calculate the second self-similarity data of the encoding vector corresponding to the (3,4) position element in the second compressed image encoding, and calculate the self-similarity deviation value of the first self-similarity data and the second self-similarity data. Each position element corresponds to a self-similarity deviation value.

[0073] If the self-similarity deviation is large, it reflects that the similarity structure of elements at each position in the first compressed image encoding is quite different from that of elements at each position in the second compressed image encoding, and the greater the impact on the relative loss value. A preset deviation threshold can be set to reflect the magnitude of the self-similarity deviation value. If the self-similarity deviation exceeds the preset deviation threshold, it reflects that the self-similarity deviation is large; otherwise, it is determined that the self-similarity deviation is small. Based on the self-similarity deviation value exceeding the preset deviation threshold, the relative loss value can be obtained and participate in the training optimization of the pre-trained autoencoder. The specific process of obtaining the relative loss value based on the self-similarity deviation value exceeding the preset deviation threshold can be to perform operations such as summation, average of summation, multiplication, and average of multiplication on the self-similarity deviation values exceeding the preset similarity threshold to obtain the relative loss value.

[0074] As a non-limiting implementation, obtaining the relative loss value based on the self-similarity deviation value exceeding the preset similarity threshold includes:

[0075] Determine the second difference between the self-similarity deviation value and the preset deviation threshold;

[0076] Determine the rectified linear unit function value of the second difference corresponding to the element at the same position, and obtain the relative loss value based on the average value of the rectified linear unit function values corresponding to the elements at each position.

[0077] Exemplarily, the second difference between the self-similarity deviation value and the preset deviation threshold can be calculated to reflect the degree to which the self-similarity deviation value exceeds the preset deviation threshold. For the second difference corresponding to each position element, calculate the rectified linear unit function value of the second difference, that is, when the second difference corresponding to the position element is greater than 0, it is adopted to participate in the calculation of the relative loss value; otherwise, it is not adopted to participate in the calculation of the loss value. Multiple second differences may be obtained for multiple position elements in the first compressed image encoding and the second compressed image encoding. The average value of the rectified linear unit function values corresponding to each position element can be calculated to obtain the relative loss value.

[0078] The relative loss value can be calculated based on the following formula:

[0079]

[0080] where l mdms is the relative loss function, and substituting specific values gives the relative loss value, b ij is the first self-similarity data corresponding to the element at the (i, j) position in the first compressed image encoding, b' ij is the second self-similarity data corresponding to the element at the (i, j) position in the second compressed image encoding, b 0 is the preset deviation threshold. |b ij - b' ij| is the absolute value of the difference between the first self-similarity data and the second self-similarity data, that is, the second difference.

[0081] Specifically, assuming that the cosine similarity is used to calculate the first self-similarity data and the second self-similarity data, the calculation formula for the relative loss value is:

[0082]

[0083] where is the first self-similarity data of the coding vector corresponding to the element at the (i, j) position of the first compressed image coding, is the second self-similarity data of the coding vector corresponding to the element at the (i, j) position of the second compressed image coding.

[0084] As a non-limiting implementation, determining the contrast loss value according to the absolute loss value and the relative loss value includes:

[0085] Determining the value of any one or more of the sum, product, power function, exponential function, logarithmic function, and trigonometric function of the absolute loss value and the relative loss value;

[0086] Taking the product of the hyperparameter coefficient and / or the optimization weight and the value of any one or more of the sum, product, power function, exponential function, logarithmic function, and trigonometric function of the absolute loss value and the relative loss value as the contrast loss value.

[0087] Exemplarily, in the process of calculating the contrast loss value according to the absolute loss value and the relative loss value, certain operations can be performed on the absolute loss value and the contrast loss value to obtain the contrast loss value. Generally, the operations need to be operations that play a mutually enhancing role, that is, the loss value obtained after the operation is larger than any single absolute loss value or relative loss value. For example, it can be an operation of any one or more of the sum, product, power function, exponential function, logarithmic function, and trigonometric function to obtain the value after the operation of the absolute loss value and the relative loss value. The hyperparameter coefficient and / or the optimization weight can be preset, and on the basis of the value obtained after the operation of the absolute loss value and the relative loss value, multiply by the hyperparameter coefficient and / or the optimization weight to obtain the contrast loss value.

[0088] Exemplarily, an example of calculating the contrast loss value is as follows:

[0089] l vf = w 1 × w 2 × (l mcos + l mdms ) ;

[0090] where, l vfis the contrast loss function. Substituting specific values gives the contrast loss value, w 1 is the optimization weight, w 2 is the hyperparameter coefficient. The optimization weight is used to represent the influence degree of the contrast loss value on the training and optimization of the pre-trained autoencoder. The specific value can be determined according to the actual situation, such as a preset value or a value obtained through training and optimization. The hyperparameter coefficient can be a preset value or a value adaptively adjusted and optimized according to the actual situation.

[0091] As a non-limiting implementation manner, the determination process of the optimization weight includes:

[0092] Determine the optimization weight according to at least one of taking a preset value, training and adjusting values within a preset range, the bisection method, and the squeeze theorem; or,

[0093] Determine the gradient of the contrast loss function and the gradients of other loss functions, and determine the optimization weight according to the ratio of the gradient of the contrast loss function to the gradients of the other loss functions, or the ratio of the gradients of the other loss functions to the gradient of the contrast loss function; wherein, the other loss function is a loss function reflecting the performance of the pre-trained autoencoder.

[0094] In the embodiments of the present application, the optimization weight can be the absolute weight value of the contrast loss value, and specifically can be determined according to at least one of taking a preset value, training and adjusting values within a preset range, the bisection method, and the squeeze theorem. For example, the integers within [1, 10000] can be sequentially trained and adjusted to determine the optimal optimization weight, or the bisection method, the squeeze theorem, etc. can also be used to test and obtain the optimal optimization weight from the preset range.

[0095] In the embodiments of the present application, if there are other loss functions in addition to the contrast loss function corresponding to the contrast loss value for the training and optimization of the pre-trained autoencoder, the optimization weight can be determined by combining the contrast loss function and the other loss functions to reflect the influence degree of the contrast loss function in the training and optimization of the pre-trained autoencoder. The gradient of the contrast loss function and the gradients of the other loss functions can be determined, and the optimization weight can be determined according to the ratio of the gradient of the contrast loss function to the gradients of the other loss functions, or according to the ratio of the gradients of the other loss functions to the gradient of the contrast loss function. The effect of the contrast loss function on the training and optimization of the pre-trained autoencoder can be reflected by the gradient of the contrast loss function, the effect of the other loss functions on the training and optimization of the pre-trained autoencoder can be reflected by the gradients of the other loss functions, and the relative difference in the effects of the loss functions in the training and optimization of the pre-trained autoencoder can be reflected by the ratio of the gradients of the two loss functions, thereby determining the optimization weight.

[0096] As a non-limiting implementation, determining the gradient of the contrast loss function and the gradients of other loss functions, and determining the optimization weight according to the ratio of the gradient of the contrast loss function to the gradients of other loss functions, or the ratio of the gradients of other loss functions to the gradient of the contrast loss function, includes:

[0097] Determine the first norm of the gradient of the contrast loss function and the second norm of the gradients of other loss functions;

[0098] Determine the optimization weight according to the ratio of the first norm to the second norm, or the ratio of the second norm to the first norm.

[0099] In the embodiments of the present application, the first norm of the gradient of the contrast loss function and the second norm of the gradients of other loss functions can be determined to represent the magnitude of the absoluteness of the extraction. The optimization weight is determined according to the ratio of the first norm to the second norm, or the ratio of the second norm to the first norm.

[0100] Exemplarily, the optimization weight can be determined based on the following formula:

[0101]

[0102] where w 1 is the optimization weight, is the second norm of the gradients of other loss functions, is the first norm of the contrast loss function.

[0103] The optimization weight can also be determined by the ratio of the first norm to the second norm, but certain arithmetic operations need to be performed to make the optimization weight negatively correlated with the ratio of the first norm to the second norm. Through the above solution, the optimization weight can be determined by comparing the training optimization effects of the contrast loss function and other loss functions, and the influence degree of the contrast loss value in the training optimization can be adaptively adjusted to improve the training optimization effect.

[0104] Figure 2 The following is a flowchart of a model training method provided by another embodiment of the present application. The embodiments of the present application are optimized based on the above embodiments. For the solutions not described in detail in the embodiments of the present application, please refer to the above embodiments. As Figure 2 shown, the method of the embodiments of the present application specifically includes the following steps:

[0105] S210. Input the training image into the encoder in the pre-trained autoencoder to obtain the first compressed image encoding, and input the training image into the target visual foundation model to obtain the second compressed image encoding; wherein, the pre-trained autoencoder includes an encoder and a decoder.

[0106] In the embodiment of the present application, the process of determining the target vision base model includes:

[0107] Determine a target vision base model from candidate vision base models according to at least one of the content type of the training image, the spatial compression rate of the pre-trained autoencoder, and the encoding dimension of the pre-trained autoencoder;

[0108] Determining a target vision base model from candidate vision base models according to the content type of the training image includes:

[0109] Determine the target vision base model according to a candidate vision base model whose content type of the application is consistent with the content type of the training image;

[0110] Determining a target vision base model from candidate vision base models according to the spatial compression rate of the autoencoder includes:

[0111] Determine the target vision base model according to a candidate vision base model whose spatial compression rate is in a first preset ratio to the spatial compression rate of the autoencoder;

[0112] Determining a target vision base model from candidate vision base models according to the encoding dimension of the autoencoder includes:

[0113] Determine the target vision base model according to a candidate vision base model whose encoding dimension is in a second preset ratio to the encoding dimension of the autoencoder.

[0114] Exemplarily, the target vision base model can be selected from candidate vision base models, and can be adaptively selected according to the actual situation. For example, a target vision base model can be determined from candidate vision base models according to at least one of the content type of the training image, the spatial compression rate of the pre-trained autoencoder, and the encoding dimension of the pre-trained autoencoder. Selecting according to the content of the training image aims to make the selected target vision base model applicable to spatially compressing and encoding the training image of this content type to obtain better results. Selecting according to the spatial compression rate and / or the encoding dimension aims to make the selected target vision base model have the same number dimension as the compressed image encoding obtained by spatially compressing and encoding the pre-trained autoencoder, or be in a certain ratio to facilitate processing to the same number dimension, so as to facilitate subsequent processing of the compressed image encoding.

[0115] Specifically, when determining the target visual base model from the candidate visual base models according to the content type of the training image, the image content types applicable to each candidate visual base model can be determined. For example, it is determined that the A-class candidate visual model is applicable to spatial compression encoding of images containing people and has a good spatial compression effect; the B-class candidate visual base model is applicable to spatial compression encoding of images containing vehicles and has a good spatial compression effect; the C-class candidate visual base model is applicable to spatial compression encoding of images containing landscapes and has a good spatial compression effect. Determine the content type of the training image, that is, the main content contained in the image, and determine the candidate visual base model consistent with the content type in the training image as the target visual base model. For example, assuming that the content type in the training image is a landscape, then select the C-class candidate visual base model as the target visual base model to achieve an ideal spatial compression encoding effect for the training image.

[0116] Specifically, the spatial compression rate reflects the degree of reduction in the number of original image features when performing spatial compression on an image. Assume that the original image is an image with [256*256] pixels, and after compression, [16*16] encoded vectors are obtained, then the spatial compression rate is 16. When determining the target visual base model from the candidate base models according to the spatial compression rate, the spatial compression rates of the candidate visual base models can be determined, and the spatial compression rate of the pre-trained autoencoder can be determined. The candidate visual base model with a first preset ratio to the spatial compression rate of the pre-trained autoencoder is used as the target visual base model. The first preset ratio can be determined according to the actual situation, such as 1, 1 / 2, 2, etc. For example, assume that the spatial compression rates of the candidate base models include 32, 20, 16, 8, and the spatial compression rate of the pre-trained autoencoder is 16. Then, the candidate visual base model with a spatial compression rate of 16 can be selected as the target visual base model so that the number of encoded vectors of the second compressed image encoding obtained by compressing and encoding the training image through the target visual base model is the same as that of the first compressed image encoding obtained by compressing and encoding through the pre-trained autoencoder. The candidate visual base model with a spatial compression rate of 32 and / or 8 can also be selected as the target visual base model to process before the training image passes through the target visual base model and the pre-trained autoencoder, or to process after the second compressed image encoding obtained by compressing and encoding the training image through the target visual base model and the first compressed image encoding obtained by compressing and encoding through the pre-trained autoencoder to make the number of encoded vectors the same.

[0117] Specifically, the encoding dimension is the dimension of the encoded vector obtained by compressing the original data. For example, after compression, there are [16*16] encoded vectors, and each encoded vector is 32-dimensional, so the encoding dimension is 32. Determining the target visual base model from the candidate visual base models according to the encoding dimension of the autoencoder can be done by determining the encoding dimension of the candidate visual base model and the encoding dimension of the pre-trained autoencoder, and identifying the candidate visual base model whose encoding dimension is in a second preset ratio to the encoding dimension of the pre-trained autoencoder as the target visual base model. The second preset ratio can be determined according to the actual situation, such as 1, 1 / 2, 2, 3, etc. The first preset ratio and the second preset ratio can be the same or different. For example, if the encoding dimensions of the candidate visual base models include 32, 16, and 10, and the encoding dimension of the pre-trained autoencoder is 16, then the candidate visual base model with an encoding dimension of 16 can be selected as the target visual base model so that the encoding vectors of the second compressed image encoding obtained after the training image is compressed and encoded by the target visual base model are consistent with those of the first compressed image encoding obtained after being compressed and encoded by the pre-trained autoencoder. Or the candidate base model with an encoding dimension of 32 can be selected as the target visual base model, and after the encoding vectors of the second compressed image encoding obtained after the training image is compressed and encoded by the target visual base model and the first compressed image encoding obtained after being compressed and encoded by the pre-trained autoencoder, processing is performed to make the dimensions of the encoding vectors consistent for subsequent processing.

[0118] It should be noted that the above three selection conditions for the target visual base model can be randomly combined as the final selection conditions. For example, a candidate visual base model can be selected as the target visual base model if it meets any one condition or must meet a specific condition. It can also be that a candidate visual base model can be selected as the target visual base model if it meets any two conditions or must meet specific two conditions. Or it can be that it must meet all three conditions to be selected as the target visual base model. The beneficial effect of the above solution is that the selected target visual model can be more suitable for processing the training image to obtain a more ideal spatial compression encoding effect, and / or the encoding of the second compressed image obtained after the training image is spatially compressed and encoded by the target visual base model is the same as the size and dimension of the first compressed image encoding obtained after being spatially compressed and encoded by the pre-trained autoencoder, or is in a certain ratio to facilitate processing to make them consistent for subsequent further processing.

[0119] In the embodiment of the present application, if the first spatial compression rate of the pre-trained autoencoder and the second spatial compression rate of the target visual base model are inconsistent, the method further includes:

[0120] Before inputting the training image into the encoder of the pre-trained autoencoder to obtain the first compressed image encoding, or before inputting the training image into the target vision foundation model to obtain the second compressed image encoding, perform resolution expansion processing or image compression processing on the training image according to the difference ratio between the first spatial compression rate of the pre-trained autoencoder and the second spatial compression rate of the target vision foundation model.

[0121] Exemplarily, in the case where the first spatial compression rate of the pre-trained autoencoder is inconsistent with the second spatial compression rate of the target vision foundation model, before inputting the training image into the target vision foundation model, and / or before inputting the training image into the pre-trained autoencoder, the training image can be processed so that the number of encoding vectors of the second compressed image encoding obtained after the training image is compressed and encoded by the target vision foundation model is the same as that of the first compressed image encoding obtained after being compressed and encoded by the pre-trained autoencoder. Specifically, resolution expansion or image compression processing can be performed on the training image according to the difference ratio between the first spatial compression rate and the second spatial compression rate.

[0122] For example, assuming that the first spatial compression rate is less than the second spatial compression rate, the ratio of the first spatial compression rate to the second compression rate can be calculated as the difference ratio. Before inputting the training image into the target vision foundation model, resolution expansion processing is performed, and the resolution in each dimension after processing is the difference ratio times that in each dimension of the original resolution. Or, before inputting the training image into the autoencoder, image compression processing is performed, and the ratio of the resolution in each dimension after processing to that in each dimension of the original resolution is the reciprocal of the difference ratio. Or, before inputting the training image into the target vision foundation model, resolution expansion processing is performed, and the resolution in each dimension after processing is the first difference ratio times that in each dimension of the original resolution. And, before inputting the training image into the autoencoder, image compression processing is performed, and the ratio of the resolution in each dimension after processing to that in each dimension of the original resolution is the second difference ratio. The product of the first difference ratio and the reciprocal of the second difference ratio is the difference ratio. Similarly, the ratio of the second spatial compression rate to the first spatial compression rate can also be calculated as the difference ratio, and the training image is processed according to the difference ratio.

[0123] Assume that the first spatial compression rate is greater than the second spatial compression rate. Then, the ratio of the second spatial compression rate to the first spatial compression rate can be calculated as the difference ratio. Before inputting the training image into the target vision-based model, image compression processing is performed. The ratio of the dimensions of the processed resolution in each direction to the dimensions of the original resolution in each direction is the reciprocal of the difference ratio. Or, before inputting the training image into the autoencoder, resolution expansion processing is performed. The dimensions of the processed resolution in each dimension are multiples of the difference ratio of the dimensions of the original resolution in each dimension. Or, before inputting the training image into the target vision-based model, image compression processing is performed. The ratio of the dimensions of the processed resolution in each dimension to the dimensions of the original resolution in each dimension is the first difference ratio. And, before inputting the training image into the autoencoder, resolution expansion processing is performed. The dimensions of the processed resolution in each dimension are multiples of the second difference ratio of the dimensions of the original resolution in each dimension. The product of the second difference ratio and the reciprocal of the first difference ratio is the difference ratio. Similarly, the ratio of the first sample space compression rate to the second spatial compression rate can be calculated as the difference ratio, and the training image can be processed according to the difference ratio.

[0124] For example, assume that the first spatial compression rate is 4 and the second spatial compression rate is 16. Then, the ratio of the first spatial compression rate to the second spatial compression rate is used as the difference ratio, which is 1 / 4. The resolution dimension of the training image is [256*256]. Before inputting the training image into the pre-trained autoencoder, the training image can be compressed, and the dimensions of the resolution of the training image in each dimension are compressed to 1 / 4 of the original, obtaining a training image with a resolution of [64*64]. Or, before inputting the training image into the target vision-based model, the resolution of the training image can be expanded, and the dimensions of the resolution of the training image in each dimension are expanded to 4 times the original, obtaining a training image with a resolution of [1024*1024]. It is also possible to compress the training image before inputting it into the pre-trained autoencoder, compressing the dimensions of the resolution of the training image in each dimension to 1 / 2 of the original, obtaining a first training image with a resolution of [128*128]. The first training image is input into the pre-trained autoencoder to obtain the first compressed image code. And, before inputting the training image into the target vision-based model, the resolution of the training image can be expanded, and the dimensions of the resolution of the training image in each dimension are expanded to 2 times the original, obtaining a second training image with a resolution of [512*512]. The second training image is input into the target vision-based model to obtain the second compressed image code.

[0125] The beneficial effect of the above solution is that processing is performed according to the first spatial compression rate and the second spatial compression rate before compressing and encoding the training image, so that the number of encoding vectors of the first compressed image code obtained after the training image is compressed and encoded by the autoencoder and the second compressed image code obtained after being compressed and encoded by the target vision-based model is the same, facilitating subsequent further processing.

[0126] In the embodiment of the present application, if the first coding dimension of the pre-trained autoencoder is inconsistent with the second coding dimension of the target vision base model, before determining the contrast loss value according to the first compressed image coding and the second compressed image coding, the method further includes:

[0127] Constructing a first mapping space according to the first coding dimension and the second coding dimension, and converting the second compressed image coding of the second coding dimension into the second compressed image coding of the first coding dimension through the first mapping space; or,

[0128] Constructing a second mapping space according to the first coding dimension and the second coding dimension, and converting the first compressed image coding of the first coding dimension into the first compressed image coding of the second coding dimension through the second mapping space.

[0129] Exemplarily, if the first coding dimension of the pre-trained autoencoder is inconsistent with the second coding dimension of the target vision base model, before determining the contrast loss value according to the first compressed image coding and the second compressed image coding, the first compressed image coding or the second compressed image coding can be processed to make the dimensions of the coding vectors consistent.

[0130] Specifically, a first mapping space can be constructed according to the first coding dimension and the second coding dimension, and the second compressed image coding of the second coding dimension can be converted into the second compressed image coding of the first coding dimension through the first mapping space. For example, the conversion is performed based on the following formula:

[0131] F’ = WF + B;

[0132] where F’ is the second compressed image coding of the first coding dimension after conversion, F is the second compressed image coding of the second coding dimension before conversion, W is the first mapping space, d z is the first coding dimension, d f is the second coding dimension, a linear correspondence between the first coding dimension and the second coding dimension is constructed through the mapping space W, and B is the matrix weight and bias term,

[0133] Specifically, a second mapping space can also be constructed according to the first coding dimension and the second coding dimension, and the first compressed image coding of the first coding dimension can be converted into the first compressed image coding of the second coding dimension through the second mapping space. For example, the conversion is performed based on the following formula:

[0134] Z’ = W’Z + B’;

[0135] Among them, Z’ is the first compressed image encoding of the second encoding dimension after conversion, Z is the first compressed image encoding of the first encoding dimension before conversion, W’ is the second mapping space. d z is the first encoding dimension, d f is the second encoding dimension. A linear correspondence relationship between the second encoding dimension and the first encoding dimension is constructed through the mapping space W’, and B’ is the matrix weight and bias term.

[0136] The beneficial effect of the above solution is that the encoding vector dimensions of the training images may be inconsistent after being compressed and encoded by the pre-trained autoencoder and the target vision base model respectively. Through the above solution, the encoding vector dimensions of the first compressed image encoding and the second compressed image encoding can be made consistent, which is convenient for subsequent processing.

[0137] It should be noted that the solution in the above embodiment does not depend on the implementation of S220 - S250 in this application embodiment.

[0138] S220. Determine the contrast loss value according to the first compressed image encoding and the second compressed image encoding.

[0139] S230. Input the first compressed image encoding into the decoder to obtain a decoded image, and determine the first original reconstruction loss value according to the decoded image and the training image.

[0140] Exemplarily, the first compressed image encoding can be input into the decoder in the pre-trained autoencoder to obtain a decoded image, and the first original reconstruction loss value is determined according to the decoded image and the training image, which reflects the loss situation between the decoded image and the training image. Common original reconstruction loss functions include mean square error (MSE), mean absolute error (MAE), structural similarity loss (SSIM Loss), etc.

[0141] S240. Use the contrast loss value and at least the first original reconstruction loss value as the target loss value.

[0142] Exemplarily, the contrast loss value and at least the first original reconstruction loss value can be used as the target loss value, that is, the target loss value includes at least the contrast loss value and the first original reconstruction loss value, and may also include other loss values reflecting the performance of the pre-trained autoencoder.

[0143] S250. Train and optimize the pre-trained autoencoder based on the target loss value.

[0144] An embodiment of the present application provides a model training method. The first compressed image encoding is input into a decoder to obtain a decoded image, and a first original reconstruction loss value is determined according to the decoded image and the training image; the contrast loss value and at least the first original reconstruction loss value are used as the target loss value. The above solution can make the target loss value include both the contrast loss value that can reflect the performance of the encoder and the first original reconstruction loss value that can reflect the performance of the encoder and the decoder, realize the training of the encoder and the decoder in the pre-trained autoencoder, and improve the performance of the autoencoder after training optimization.

[0145] As a non-limiting implementation, the method further includes:

[0146] After completing the training optimization of the pre-trained autoencoder, an autoencoder that meets the performance requirements is obtained;

[0147] The training image is input into the encoder in the autoencoder that meets the performance requirements to obtain a third compressed image encoding;

[0148] The third compressed image encoding is input into a generation model to obtain a fourth compressed image encoding;

[0149] A first loss value is determined according to the third compressed image encoding and the fourth compressed image encoding, and the generation model is trained and optimized according to the first loss value; or,

[0150] The fourth compressed image encoding is input into the decoder in the autoencoder that meets the performance requirements to obtain a first generated image, a second original reconstruction loss value is determined according to the first generated image and the training image, and the generation model is trained and optimized according to the second original reconstruction loss value.

[0151] Exemplarily, in practical applications, an autoencoder is generally used together with a generative model to implement corresponding functions. Specifically, an image is input into the encoder of the autoencoder, and the image space is compressed into a compressed image code. The generative model generates a new compressed image code based on the compressed image code, and the decoder decodes and restores the new compressed image code generated by the generative model to obtain a decoded image. During the training process, the generative model also needs to be trained. The specific process can be as follows: After the pre-trained autoencoder is trained and optimized according to the solution of the above embodiment, an autoencoder that meets the performance requirements is obtained. The training image is input into the encoder of the autoencoder that meets the performance requirements to obtain a third compressed image code. The third compressed image code is input into the generative model to obtain a fourth compressed image code. The first loss value is determined according to the third compressed image code and the fourth compressed image code, and the generative model is trained and optimized according to the first loss value. That is, after the autoencoder is trained and optimized, the first loss value is determined according to the third compressed image code before being input into the generative model and the fourth compressed image code output by the generative model to train and optimize the generative model. In addition, the fourth compressed image code can be input into the decoder of the autoencoder that meets the performance requirements to obtain a first generated image, and the second original reconstruction loss value is determined according to the first generated image and the training image, and the generative model is trained and optimized based on the second original reconstruction loss value. The above process separates the training processes of the autoencoder and the generative model, thereby realizing orderly training. The performance of the autoencoder after training and optimization is improved, which can improve the optimization speed of the generative model.

[0152] As a non-limiting implementation manner, the method further includes:

[0153] Input the first compressed image code into the generative model to obtain a fifth compressed image code;

[0154] Determine a second loss value according to the first compressed image code and the fifth compressed image code, and train and optimize the generative model according to the second loss value; or,

[0155] Input the fifth compressed image code into the decoder of the pre-trained autoencoder to obtain a second generated image, determine a third original reconstruction loss value according to the second generated image and the training image, and train and optimize the generative model according to the third original reconstruction loss value.

[0156] In an embodiment of the present application, the pre-trained autoencoder and the generative model can also be trained simultaneously. Exemplarily, the first compressed image encoding is input into the generative model to obtain a fifth compressed image encoding. A second loss value is determined based on the first compressed image encoding and the fifth compressed image encoding. The generative model is trained and optimized according to the second loss value. The first compressed image encoding is also involved in the training and optimization of the pre-trained autoencoder in the above embodiment.

[0157] The fifth compressed image encoding can also be input into the decoder in the pre-trained autoencoder to obtain a second generated image. A third reconstruction loss value is determined based on the second generated image and the training image. The generative model is trained and optimized according to the third reconstruction loss value. That is, the loss function for training and optimizing the generative model includes the output result of the pre-trained autoencoder. At the same time, the output result of the pre-trained autoencoder is also involved in its own training and optimization, realizing the synchronous training and optimization of the pre-trained autoencoder and the generative model, and improving the training and optimization efficiency.

[0158] Figure 3 The structure diagram of a model training device provided by an embodiment of the present application. This device can execute the model training method provided by any embodiment of the present application and has the corresponding functional modules and beneficial effects for executing the method. As Figure 3 shown, the device includes:

[0159] A training image input module 310, configured to input a training image into an encoder in a pre-trained autoencoder to obtain a first compressed image encoding, and input the training image into a target visual foundation model to obtain a second compressed image encoding; wherein, the pre-trained autoencoder includes an encoder and a decoder;

[0160] A target loss value determination module 320, configured to determine a contrast loss value based on the first compressed image encoding and the second compressed image encoding, and use the contrast loss value and at least one other loss value as the target loss value; wherein, the at least one other loss value is a loss value reflecting the performance of the pre-trained autoencoder;

[0161] A training and optimization module 330, configured to train and optimize the pre-trained autoencoder based on the target loss value.

[0162] In an embodiment of the present application, determining a contrast loss value based on the first compressed image encoding and the second compressed image encoding includes:

[0163] Determining an absolute loss value based on the similarity between the first compressed image encoding and the second compressed image encoding; and / or,

[0164] Determine a relative loss value according to the first self-similarity data encoded by the first compressed image and the second self-similarity data encoded by the second compressed image;

[0165] Determine the contrast loss value according to the absolute loss value and / or the relative loss value.

[0166] In an embodiment of the present application, the target loss value determination module 320 determines an absolute loss value according to the similarity between the first compressed image encoding and the second compressed image encoding, including:

[0167] For each corresponding position element in the first compressed image encoding and the second compressed image encoding, determine the similarity between the first vector corresponding to the first compressed image encoding and the second vector corresponding to the second compressed image encoding at the same position element;

[0168] Determine whether the similarity is lower than a preset similarity threshold;

[0169] Based on the target similarity lower than the preset similarity threshold, obtain the absolute loss value.

[0170] In an embodiment of the present application, the target loss value determination module 320 obtains the absolute loss value based on the target similarity less than the preset similarity threshold, including:

[0171] Determine a first difference between the preset similarity threshold and the similarity;

[0172] Determine the rectified linear unit function value of the first difference corresponding to the same position element, and obtain the absolute loss value based on the average value of the rectified linear unit function values corresponding to each position element.

[0173] In an embodiment of the present application, the target loss value determination module 320 determines a relative loss value according to the first self-similarity data of the first compressed image encoding and the second self-similarity data of the second compressed image encoding, including:

[0174] For each corresponding position element in the first compressed image encoding and the second compressed image encoding, determine the self-similarity deviation value between the first self-similarity data and the second self-similarity data of the same position element;

[0175] Determine whether the self-similarity deviation value exceeds a preset deviation threshold;

[0176] Based on the self-similarity deviation value exceeding the preset similarity threshold, obtain the relative loss value.

[0177] In the embodiment of the present application, the target loss value determination module 320 obtains the relative loss value based on the self - similarity deviation value exceeding the preset similarity threshold, including:

[0178] Determine a second difference between the self - similarity deviation value and the preset deviation threshold;

[0179] Determine the rectified linear unit function value of the second difference corresponding to the elements at the same position, and obtain the relative loss value based on the average value of the rectified linear unit function values corresponding to the elements at each position.

[0180] In the embodiment of the present application, the target loss value determination module 320 determines the contrast loss value according to the absolute loss value and the relative loss value, including:

[0181] Determine the value of any one or more of the sum, product, power function, exponential function, logarithmic function, and trigonometric function of the absolute loss value and the relative loss value;

[0182] Take the product of the hyperparameter coefficient and / or the optimization weight and the value of any one or more of the sum, product, power function, exponential function, logarithmic function, and trigonometric function of the absolute loss value and the relative loss value as the contrast loss value.

[0183] In the embodiment of the present application, the target loss value determination module 320 determines the optimization weight including:

[0184] Determine the optimization weight according to at least one of taking a preset value, training and adjusting the values within a preset range, the bisection method, and the squeeze theorem; or,

[0185] Determine the contrast loss function gradient and the other loss function gradient, and determine the optimization weight according to the ratio of the contrast loss function gradient to the other loss function gradient, or the ratio of the other loss function gradient to the contrast loss function gradient; wherein, the other loss function is a loss function reflecting the performance of the pre - trained auto - encoder.

[0186] In the embodiment of the present application, the target loss value determination module 320 determines the contrast loss function gradient and the other loss function gradient, and determines the optimization weight according to the ratio of the contrast loss function gradient to the other loss function gradient, or the ratio of the other loss function gradient to the contrast loss function gradient, including:

[0187] Determine the first norm of the contrast loss function gradient and the second norm of the other loss function gradient;

[0188] Determine the optimization weight according to the ratio of the first norm to the second norm, or the ratio of the second norm to the first norm.

[0189] In the embodiment of the present application, the target loss value determination module 320 uses the comparison loss value and at least one other loss value as the target loss value, including:

[0190] Weight the comparison loss value with at least one other loss value as the target loss value to feedback to the pre-trained autoencoder for training optimization; wherein, at least one other loss value includes at least one of an original reconstruction loss value, a generative adversarial network loss value, and a divergence loss value.

[0191] In the embodiment of the present application, the training image input module 310 determines the target visual base model, including:

[0192] Determine the target visual base model from candidate visual base models according to at least one of the content type of the training image, the spatial compression rate of the pre-trained autoencoder, and the encoding dimension of the pre-trained autoencoder;

[0193] Determining the target visual base model from candidate visual base models according to the content type of the training image includes:

[0194] Determine the target visual base model according to a candidate visual base model whose content type of the application is consistent with the content type of the training image;

[0195] Determining the target visual base model from candidate visual base models according to the spatial compression rate of the autoencoder includes:

[0196] Determine the target visual base model according to a candidate visual base model whose spatial compression rate is in a first preset ratio to the spatial compression rate of the autoencoder;

[0197] Determining the target visual base model from candidate visual base models according to the encoding dimension of the autoencoder includes:

[0198] Determine the target visual base model according to a candidate visual base model whose encoding dimension is in a second preset ratio to the encoding dimension of the autoencoder.

[0199] In the embodiment of the present application, if the first spatial compression rate of the pre-trained autoencoder is inconsistent with the second spatial compression rate of the target visual base model, the device further includes:

[0200] A training image processing module is configured to, before inputting the training image into an encoder of a pre-trained autoencoder to obtain a first compressed image encoding, and / or before inputting the training image into a target vision base model to obtain a second compressed image encoding, perform resolution expansion processing or image compression processing on the training image according to a difference ratio between a first spatial compression rate of the pre-trained autoencoder and a second spatial compression rate of the target vision base model.

[0201] In an embodiment of the present application, if a first encoding dimension of the pre-trained autoencoder is inconsistent with a second encoding dimension of the target vision base model, before determining a contrast loss value according to the first compressed image encoding and the second compressed image encoding, the apparatus further includes:

[0202] A conversion module is configured to construct a first mapping space according to the first encoding dimension and the second encoding dimension, and convert the second compressed image encoding with the second encoding dimension into a second compressed image encoding with the first encoding dimension through the first mapping space; or,

[0203] Construct a second mapping space according to the first encoding dimension and the second encoding dimension, and convert the first compressed image encoding with the first encoding dimension into a first compressed image encoding with the second encoding dimension through the second mapping space.

[0204] In an embodiment of the present application, a target loss value determination module 320 uses the contrast loss value and at least one other loss value as the target loss value, including:

[0205] Input the first compressed image encoding into a decoder to obtain a decoded image, and determine a first original reconstruction loss value according to the decoded image and the training image;

[0206] Use the contrast loss value and at least the first original reconstruction loss value as the target loss value.

[0207] In an embodiment of the present application, the apparatus further includes: a first generation model training module, configured to:

[0208] After completing training and optimization of the pre-trained autoencoder, obtain an autoencoder that meets performance requirements;

[0209] Input the training image into an encoder of the autoencoder that meets performance requirements to obtain a third compressed image encoding;

[0210] Input the third compressed image encoding into a generation model to obtain a fourth compressed image encoding;

[0211] Determine a first loss value according to the third compressed image encoding and the fourth compressed image encoding, and train and optimize the generation model according to the first loss value; or,

[0212] Input the fourth compressed image encoding into the decoder in an autoencoder that meets the performance requirements to obtain a first generated image, determine a second original reconstruction loss value according to the first generated image and the training image, and train and optimize the generation model according to the second original reconstruction loss value.

[0213] In the embodiments of the present application, the device further includes: a first generation model training module, configured to:

[0214] Input the first compressed image encoding into the generation model to obtain a fifth compressed image encoding;

[0215] Determine a second loss value according to the first compressed image encoding and the fifth compressed image encoding, and train and optimize the generation model according to the second loss value; or,

[0216] Input the fifth compressed image encoding into the decoder in the pre-trained autoencoder to obtain a second generated image, determine a third original reconstruction loss value according to the second generated image and the training image, and train and optimize the generation model according to the third original reconstruction loss value.

[0217] A model training device provided by the embodiments of the present application can execute a model training method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method.

[0218] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described herein and / or claimed.

[0219] As Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0220] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless model training transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0221] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the model training method.

[0222] In some embodiments, the model training method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the model training method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the model training method by any other appropriate means (e.g., by means of firmware).

[0223] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0224] The computer programs for implementing the methods of this application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable model training device, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0225] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0226] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0227] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0228] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0229] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the information desired by the technical solution of this application can be achieved, and no limitation is imposed herein.

[0230] The above specific implementation manners do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A model training method, characterized in that: The method comprises: Inputting a training image into an encoder in a pre-trained autoencoder to obtain a first compressed image code, and inputting the training image into a target visual basis model to obtain a second compressed image code; wherein the pre-trained autoencoder includes an encoder and a decoder; Determine a contrast loss value according to the first compressed image encoding and the second compressed image encoding, and use the contrast loss value and at least one other loss value as target loss values; wherein at least one other loss value is a loss value reflecting the performance of the pre-trained autoencoder; The pre-trained autoencoder is trained and optimized based on the target loss value.

2. The method according to claim 1, characterized in that Determining a contrast loss value according to the first compressed image encoding and the second compressed image encoding includes: determining an absolute loss value according to a similarity between the first compressed image encoding and the second compressed image encoding; and / or, determining a relative loss value based on first self-similarity data encoded by the first compressed image and second self-similarity data encoded by the second compressed image; The comparative loss value is determined according to the absolute loss value and / or the relative loss value.

3. The method according to claim 2, characterized in that Determining an absolute loss value according to a similarity between the first compressed image encoding and the second compressed image encoding includes: For each corresponding position element in the first compressed image code and the second compressed image code, determine the similarity between a first vector in the first compressed image code and a second vector in the second compressed image code corresponding to the same position element; Determining whether the similarity is lower than a preset similarity threshold; The absolute loss value is obtained based on the target similarity that is lower than the preset similarity threshold.

4. The method according to claim 3, characterized in that Based on the target similarity less than the preset similarity threshold, obtaining the absolute loss value includes: Determine a first difference between the preset similarity threshold and the similarity; The linear rectification function value of the first difference corresponding to the elements at the same position is determined, and the absolute loss value is obtained based on the average value of the linear rectification function values ​​corresponding to the elements at each position.

5. The method according to claim 2, characterized in that: Determining a relative loss value according to first self-similarity data encoded by the first compressed image and second self-similarity data encoded by the second compressed image includes: For each corresponding position element in the first compressed image code and the second compressed image code, determine the self-similarity deviation value of the first self-similarity data and the second self-similarity data of the same position element; Determining whether the self-similarity deviation value exceeds a preset deviation threshold; The relative loss value is obtained based on the self-similarity deviation value exceeding the preset similarity threshold.

6. The method according to claim 5, characterized in that The relative loss value is obtained based on the self-similarity deviation value exceeding the preset similarity threshold, including: Determining a second difference between the self-similarity deviation value and the preset deviation threshold; The linear rectification function value of the second difference corresponding to the element at the same position is determined, and the relative loss value is obtained based on the average value of the linear rectification function values ​​corresponding to the elements at each position.

7. The method according to claim 2, characterized in that Determining the comparative loss value according to the absolute loss value and the relative loss value includes: Determine the value of any one or more of the sum, product, power function, exponential function, logarithmic function, and trigonometric function of the absolute loss value and the relative loss value; The hyperparameter coefficient and / or optimization weight is taken as the comparison loss value, and the product of the sum, product, power function, exponential function, logarithmic function, trigonometric function of the absolute loss value and the relative loss value or multiple values ​​thereof.

8. The method according to claim 7, characterized in that The process of determining the optimization weight includes: Determine the optimization weight by taking a preset value, performing training adjustment on a value within a preset range, dichotomy, or the squeeze theorem; or Determine the contrast loss function gradient and other loss function gradients, and determine the optimization weight according to the ratio of the contrast loss function gradient to the other loss function gradient, or the ratio of the other loss function gradient to the contrast loss function gradient; wherein the other loss function is a loss function that reflects the performance of the pre-trained autoencoder.

9. The method according to claim 8, characterized in that Determining the contrast loss function gradient and other loss function gradients, and determining the optimization weight according to the ratio of the contrast loss function gradient to the other loss function gradient, or the ratio of the other loss function gradient to the contrast loss function gradient, includes: Determine the first norm of the contrast loss function gradient and the second norm of the other loss function gradients; The optimization weight is determined according to a ratio of the first norm to the second norm, or a ratio of the second norm to the first norm.

10. The method according to claim 1, characterized in that Using the comparison loss value and at least one other loss value as a target loss value comprises: The contrast loss value and at least one other loss value are weighted as the target loss value to be fed back to the pre-trained autoencoder for training optimization; wherein the at least one other loss value includes at least one of the original reconstruction loss value, the generative adversarial network loss value, and the divergence loss value.

11. The method according to claim 1, characterized in that: The process of determining the target visual basic model includes: Determining a target visual base model from candidate visual base models according to at least one of a content type of the training image, a spatial compression rate of the pre-trained autoencoder, and an encoding dimension of the pre-trained autoencoder; Determining a target visual basis model from candidate visual basis models according to the content type of the training image includes: Determining the target visual base model according to a candidate visual base model whose applied content type is consistent with the content type of the training image; Determining a target visual basis model from candidate visual basis models according to the spatial compression rate of the autoencoder includes: Determining the target visual basis model according to candidate visual basis models whose spatial compression rates are in a first preset ratio to the spatial compression rates of the autoencoder; Determining a target visual basis model from candidate visual basis models according to the encoding dimension of the autoencoder includes: The target visual basis model is determined based on candidate visual basis models whose encoding dimensions are in a second preset ratio to the encoding dimension of the autoencoder.

12. The method according to claim 1, characterized in that If the first spatial compression rate of the pre-trained autoencoder is inconsistent with the second spatial compression rate of the target visual base model, the method further includes: Before the training image is input into the encoder in the pre-trained autoencoder to obtain the first compressed image code, and / or before the training image is input into the target visual basis model to obtain the second compressed image code, the training image is subjected to resolution expansion processing or image compression processing according to the difference ratio between the first spatial compression rate of the pre-trained autoencoder and the second spatial compression rate of the target visual basis model.

13. The method according to claim 1, characterized in that If the first encoding dimension of the pre-trained autoencoder is inconsistent with the second encoding dimension of the target visual basis model, before determining the contrast loss value according to the first compressed image encoding and the second compressed image encoding, the method further includes: Constructing a first mapping space according to the first coding dimension and the second coding dimension, and converting the second compressed image code of the second coding dimension into the second compressed image code of the first coding dimension through the first mapping space; or, A second mapping space is constructed according to the first coding dimension and the second coding dimension, and the first compressed image code of the first coding dimension is converted into the first compressed image code of the second coding dimension through the second mapping space.

14. The method according to claim 1, characterized in that Using the comparison loss value and at least one other loss value as a target loss value comprises: Input the first compressed image encoding into a decoder to obtain a decoded image, and determine a first original reconstruction loss value according to the decoded image and the training image; The contrast loss value and at least the first original reconstruction loss value are used as the target loss value.

15. The method according to claim 1, characterized in that The method further comprises: After completing the training optimization of the pre-trained autoencoder, an autoencoder that meets the performance requirements is obtained; Inputting the training image into an encoder in an autoencoder that meets performance requirements to obtain a third compressed image code; Inputting the third compressed image code into a generation model to obtain a fourth compressed image code; Determine a first loss value according to the third compressed image code and the fourth compressed image code, and train and optimize the generation model according to the first loss value; or, The fourth compressed image is encoded and input into a decoder in an autoencoder that meets the performance requirements to obtain a first generated image, a second original reconstruction loss value is determined based on the first generated image and the training image, and the generation model is trained and optimized based on the second original reconstruction loss value.

16. The method according to claim 1, characterized in that The method further comprises: Inputting the first compressed image code into a generation model to obtain a fifth compressed image code; Determine a second loss value according to the first compressed image code and the fifth compressed image code, and train and optimize the generation model according to the second loss value; or, The fifth compressed image is encoded and input into the decoder in the pre-trained autoencoder to obtain a second generated image, a third original reconstruction loss value is determined based on the second generated image and the training image, and the generation model is trained and optimized based on the third original reconstruction loss value.

17. A target model training device, characterized in that: The device comprises: A training image input module, used for inputting a training image into an encoder in a pre-trained autoencoder to obtain a first compressed image code, and inputting the training image into a target visual basis model to obtain a second compressed image code; wherein the pre-trained autoencoder includes an encoder and a decoder; a target loss value determination module, configured to determine a contrast loss value according to the first compressed image code and the second compressed image code, and use the contrast loss value and at least one other loss value as a target loss value; wherein the at least one other loss value is a loss value reflecting the performance of the pre-trained autoencoder; A training optimization module is used to perform training optimization on the pre-trained autoencoder based on the target loss value.

Citation Information

Cited By

  • Training method of variational auto-encoder, video generation method and corresponding device

    CN120471107A

  • Variational autoencoder training method, video generation method and corresponding device

    CN120471107B

  • Visual generation model training method and device based on linear network layer, equipment, medium and product

    CN121413694A

  • Linear network layer-based visual generation model training method and device, equipment, medium and product

    CN121413694B