Image evaluation model processing method, image processing method, and related apparatus

CN118861705BActive Publication Date: 2026-09-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410917057.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2026-09-04
Estimated Expiration
2044-07-09

AI Technical Summary

Technical Problem

[0004]然而,采用人工或相似度评估的方式,误差较大,影响图文匹配度的评估准确度

Benefits of technology

[0026] The image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product described above all involve training an image evaluation model using the aforementioned image evaluation model processing method. In the fixed-length training phase, multiple first sample image sequences of the same length are used to train the image evaluation model. The sample images are ordered according to their matching degree. For each sample image in each first sample image sequence, a first predicted matching degree with the first sample text is predicted. Based on the first predicted matching degree and the order of the sample images, a fixed-length sequence loss value is determined. This fixed-length sequence loss value is then used to train the image evaluation model. Through training, the image evaluation model learns to predict, i.e., evaluate, the matching degree between the sample image and the first sample text. The ability to match the text between images is enhanced through several methods. During the variable-length training phase, multiple second-sample image sequences of varying lengths are used to train the image evaluation model. For each sample image in each second-sample image sequence, a second predicted matching degree with the second-sample text is predicted. Based on the second predicted matching degree and the sorting order of the sample images in the multiple second-sample image sequences, a variable-length sequence loss value is determined. This loss value is then used to train the image evaluation model, allowing for further adjustments. This enables the model to learn the ability to predict, and evaluate, the second predicted matching degree between the sample image and the second-sample text using second-sample image sequences of different lengths, thereby further improving the accuracy of the image-text matching degree evaluation. Therefore, the image evaluation model trained using the above image evaluation model processing method can accurately determine the target matching degree between the generated image and the text, improving the accuracy of the image-text matching degree evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861705B_ABST
    Figure CN118861705B_ABST
Patent Text Reader

Abstract

The application relates to an image evaluation model processing method, an image processing method and related devices. The method comprises the following steps: in a fixed-length training stage, a plurality of first sample image sequences of a fixed length are acquired, a first prediction matching degree with first sample text is predicted for each sample image in each first sample image sequence through an image evaluation model, and the image evaluation model is trained based on the first prediction matching degree and the sorting order of the sample images in the plurality of first sample image sequences; in an indefinite-length training stage, a plurality of second sample image sequences of an indefinite length are acquired, a second prediction matching degree with second sample text is predicted for each sample image in each second sample image sequence through the image evaluation model, and the image evaluation model is trained based on the second prediction matching degree and the sorting order of the sample images in the plurality of second sample image sequences. The method can improve the evaluation accuracy of the image-text matching degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image evaluation model processing method, an image processing method, and related apparatus. Background Technology

[0002] With the development of artificial intelligence and image processing technologies, text-generated images are increasingly used. Text-generated images refer to images generated from text; for example, multiple images can be generated from the same text. These images are called generated images. To evaluate the matching degree between the generated image and the text, the generated image needs to be evaluated. The matching degree between the generated image and the text, or simply image-text matching degree, reflects the degree of similarity between the content presented by the generated image and the content expressed by the text.

[0003] In traditional techniques, the generated images are typically evaluated manually or based on the similarity between the text and the generated image.

[0004] However, using manual or similarity assessment methods results in significant errors, affecting the accuracy of image-text matching assessment. Summary of the Invention

[0005] Therefore, it is necessary to provide an image evaluation model processing method, image processing method, and related apparatus that can improve the evaluation accuracy of image-text matching degree in response to the above-mentioned technical problems.

[0006] On one hand, this application provides an image evaluation model processing method, comprising: in a fixed-length training phase, acquiring a plurality of fixed-length first sample image sequences, each first sample image sequence including sample images generated using first sample text, and the included sample images being sorted according to their matching degree with the first sample text, wherein each of the plurality of first sample image sequences has the same sequence length, the sequence length referring to the number of included sample images; using an image evaluation model, for each sample image in each of the first sample image sequences, predicting a first predicted matching degree with the first sample text, determining a fixed-length sequence loss value based on the first predicted matching degree and the sorting order of the sample images in the plurality of first sample image sequences, and utilizing the fixed-length sequence loss value to further refine the image processing method. The image evaluation model is trained using a long sequence loss value. During the variable-length training phase, multiple second sample image sequences of variable length are acquired. Each second sample image sequence includes sample images generated using second sample text, and the included sample images are ordered according to their matching degree with the second sample text. Variable length means that at least two second sample image sequences have different sequence lengths. Using the image evaluation model, for each sample image in each second sample image sequence, a second predicted matching degree with the second sample text is predicted. Based on the second predicted matching degree and the sorting order of the sample images in the multiple second sample image sequences, a variable-length sequence loss value is determined, and the image evaluation model is trained using the variable-length sequence loss value.

[0007] On the other hand, this application also provides an image evaluation model processing apparatus, comprising: a first sequence acquisition module, configured to acquire a plurality of fixed-length first sample image sequences during a fixed-length training phase, each first sample image sequence comprising sample images generated using first sample text, and the included sample images being sorted according to their matching degree with the first sample text, wherein each of the plurality of first sample image sequences has the same sequence length, the sequence length referring to the number of included sample images; and a first training module, configured to, through an image evaluation model, predict a first predicted matching degree with the first sample text for each sample image in each of the first sample image sequences, determine a fixed-length sequence loss value based on the first predicted matching degree and the sorting order of the sample images in the plurality of first sample image sequences, and utilize the... The image evaluation model is trained using a fixed-length sequence loss value. A second sequence acquisition module is used to acquire multiple second sample image sequences of variable length during the variable-length training phase. Each second sample image sequence includes sample images generated using second sample text, and the included sample images are ordered according to their matching degree with the second sample text. Variable length means that at least two second sample image sequences have different sequence lengths. A second training module is used to predict a second predicted matching degree with the second sample text for each sample image in each second sample image sequence using the image evaluation model. Based on the second predicted matching degree and the sorting order of the sample images in the multiple second sample image sequences, a variable-length sequence loss value is determined, and the image evaluation model is trained using the variable-length sequence loss value.

[0008] In some embodiments, the image evaluation model processing apparatus further includes a set acquisition module, which is configured to acquire a set of multiple sample image sequences determined based on multiple sample images corresponding to the first sample text. The sample images are images generated using the first sample text. The sequence lengths of the sample image sequences in different sample image sequence sets are different, while the sequence lengths of the sample image sequences in the same sample image sequence set are the same. The sample image sequence includes at least a portion of the sample images in the multiple sample images, and the sample images in the sample image sequence are sorted according to their degree of matching with the first sample text. The first sequence acquisition module is further configured to acquire a portion or all of the sample image sequences from one of the multiple sample image sequence sets to obtain a set of multiple first sample image sequences of a fixed length.

[0009] In some embodiments, the first sample text and the second sample text are the same text, and the second sequence acquisition module is further configured to determine at least two sample image sequence sets from the plurality of sample image sequence sets; and to acquire part or all of the sample image sequences from each of the at least two sample image sequence sets to obtain a plurality of second sample image sequences of variable length.

[0010] In some embodiments, the set acquisition module is further configured to acquire multiple different global sequences obtained by sorting the multiple sample images according to the image-text matching degree determined by each of the multiple objects, wherein the image-text matching degree is the matching degree between each sample image and the first sample text; and determine a set of sample image sequences corresponding to multiple specified sequence lengths based on the multiple different global sequences, wherein the sequence length of the sample image sequences in the set of sample image sequences is the corresponding specified sequence length, and the relative sorting order of the sample images in the sample image sequences is consistent with the relative sorting order in each of the global sequences.

[0011] In some embodiments, the set acquisition module is further configured to determine a common sequence that satisfies a preset condition based on the plurality of different global sequences. The preset condition is that the relative sorting order of each sample image contained therein is consistent in each of the global sequences, and that the relative sorting order is no longer consistent in each of the global sequences after adding any sample image. For each specified sequence length among the plurality of specified sequence lengths, at least one subsequence with a sequence length of the specified sequence length is obtained from each common sequence with a sequence length greater than or equal to the specified sequence length, and a sample image sequence set corresponding to the specified sequence length is formed.

[0012] In some embodiments, the first training module is further configured to, for each first sample image sequence, determine the loss value generated by the first sample image sequence based on the sorting order of the sample images in the first sample image sequence and the first prediction matching degree of the sample images; and perform comprehensive calculation on the loss values ​​generated by each first sample image sequence to obtain a fixed-length sequence loss value.

[0013] In some embodiments, the first training module is further configured to: determine at least one pair of adjacent sample images in the first sample image sequence according to the sorting order of the sample images in the first sample image sequence, wherein the pair of adjacent sample images includes two adjacent sample images in the first sample image sequence; for each pair of adjacent sample images, determine the difference between the first predicted matching degrees of the sample images in the pair of adjacent sample images to obtain the matching degree difference corresponding to the pair of adjacent sample images; and obtain the loss value generated by the first sample image sequence based on the matching degree difference corresponding to each pair of adjacent sample images.

[0014] In some embodiments, the loss value generated by the first sample image sequence includes a first loss value. The first training module is further configured to normalize the matching degree difference corresponding to each of the adjacent sample image pairs to obtain the normalized matching degree difference corresponding to the adjacent sample image pairs; and to obtain the first loss value generated by the first sample image sequence based on the normalized matching degree difference corresponding to each of the adjacent sample image pairs.

[0015] In some embodiments, the loss value generated by the first sample image sequence includes a second loss value. The first training module is further configured to perform a logarithmic transformation on the matching degree difference corresponding to each of the adjacent sample image pairs to obtain the transformed matching degree difference corresponding to the adjacent sample image pairs; use the matching degree difference corresponding to each of the adjacent sample image pairs as the weight of the transformed matching degree difference corresponding to the corresponding transformed matching degree difference, perform a weighted calculation on the transformed matching degree difference corresponding to each of the adjacent sample image pairs, and obtain the second loss value generated by the first sample image sequence based on the weighted result.

[0016] In some embodiments, the loss value generated by the first sample image sequence includes a third loss value. The first training module is further configured to determine the matching degree difference between the first predicted matching degree and the minimum matching degree of the sample image with the smallest matching degree in the first sample image sequence; sum the matching degree differences and the minimum matching degree to obtain a summation result; and obtain the third loss value generated by the first sample image sequence based on the summation result, wherein the third loss value is negatively correlated with the summation result.

[0017] On the other hand, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described image evaluation model processing method.

[0018] On the other hand, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described image evaluation model processing method.

[0019] On the other hand, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described image evaluation model processing method.

[0020] The aforementioned image evaluation model processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product, in the fixed-length training phase, use multiple first sample image sequences of the same sequence length to train the image evaluation model. The sample images are sorted according to their matching degree. For each sample image in each first sample image sequence, a first predicted matching degree with the first sample text is predicted. Based on the first predicted matching degree and the sorting order of the sample images, a fixed-length sequence loss value is determined. The fixed-length sequence loss value is used to train the image evaluation model. Thus, through training, the image evaluation model can learn the ability to predict, i.e., evaluate, the matching degree between the sample image and the first sample text. In the variable-length training phase, multiple second sample image sequences of variable length are used to train the image evaluation model. For each sample image in each second sample image sequence, a second predicted matching degree with the second sample text is predicted. Based on the second predicted matching degree and the sorting order of the sample images in the multiple second sample image sequences, a variable-length sequence loss value is determined. The image evaluation model is trained using the variable-length sequence loss value, which can further adjust the image evaluation model so that it can learn to predict the second predicted matching degree between the sample image and the second sample text for second sample image sequences of different lengths, thereby further improving the evaluation accuracy of image-text matching degree.

[0021] On the other hand, this application also provides an image processing method, the method comprising: acquiring text features obtained by encoding text; acquiring image features obtained by encoding a generated image generated using the text; fusing the image features with the text features to obtain image-text fusion features corresponding to the generated image; and inputting the image-text fusion features into a trained image evaluation model to obtain the target matching degree between the generated image and the text, wherein the image evaluation model is trained by the above-described image evaluation model processing method.

[0022] On the other hand, this application also provides an image processing apparatus, comprising: a text feature acquisition module for acquiring text features obtained by encoding text; an image feature acquisition module for acquiring image features obtained by encoding a generated image generated using the text; a feature fusion module for fusing the image features with the text features to obtain image-text fusion features corresponding to the generated image; and a matching degree acquisition module for inputting the image-text fusion features into a trained image evaluation model to obtain the target matching degree between the generated image and the text, wherein the image evaluation model is trained using the above-described image evaluation model processing method.

[0023] On the other hand, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described image processing method.

[0024] On the other hand, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described image processing method.

[0025] On the other hand, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described image processing method.

[0026] The image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product described above all involve training an image evaluation model using the aforementioned image evaluation model processing method. In the fixed-length training phase, multiple first sample image sequences of the same length are used to train the image evaluation model. The sample images are ordered according to their matching degree. For each sample image in each first sample image sequence, a first predicted matching degree with the first sample text is predicted. Based on the first predicted matching degree and the order of the sample images, a fixed-length sequence loss value is determined. This fixed-length sequence loss value is then used to train the image evaluation model. Through training, the image evaluation model learns to predict, i.e., evaluate, the matching degree between the sample image and the first sample text. The ability to match the text between images is enhanced through several methods. During the variable-length training phase, multiple second-sample image sequences of varying lengths are used to train the image evaluation model. For each sample image in each second-sample image sequence, a second predicted matching degree with the second-sample text is predicted. Based on the second predicted matching degree and the sorting order of the sample images in the multiple second-sample image sequences, a variable-length sequence loss value is determined. This loss value is then used to train the image evaluation model, allowing for further adjustments. This enables the model to learn the ability to predict, and evaluate, the second predicted matching degree between the sample image and the second-sample text using second-sample image sequences of different lengths, thereby further improving the accuracy of the image-text matching degree evaluation. Therefore, the image evaluation model trained using the above image evaluation model processing method can accurately determine the target matching degree between the generated image and the text, improving the accuracy of the image-text matching degree evaluation. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a diagram illustrating the application environment of the image processing methods in some embodiments;

[0029] Figure 2 This is a flowchart illustrating the image evaluation model processing method in some embodiments;

[0030] Figure 3 This is a schematic diagram of sample images in some embodiments;

[0031] Figure 4 This is a schematic diagram of the image evaluation model processing method in some embodiments;

[0032] Figure 5 This is a schematic diagram illustrating the principle of generating image-text fusion features in some embodiments;

[0033] Figure 6 This is a schematic diagram illustrating the principle of image-text fusion in some embodiments;

[0034] Figure 7 This is a schematic diagram of the image evaluation model processing method in some embodiments;

[0035] Figure 8 This is a flowchart illustrating the common sequences obtained in some embodiments;

[0036] Figure 9 Here are function curves in some embodiments;

[0037] Figure 10 This is a flowchart illustrating the image processing method in some embodiments;

[0038] Figure 11 This is a schematic diagram illustrating the principles of some embodiment image processing methods;

[0039] Figure 12 These are application scenario diagrams of some embodiment image processing methods;

[0040] Figure 13 This is a flowchart illustrating the image processing method in some other embodiments;

[0041] Figure 14 This is a structural block diagram of the image evaluation model processing device in some embodiments;

[0042] Figure 15 This is a structural block diagram of the image processing apparatus in some embodiments;

[0043] Figure 16 These are internal structural diagrams of the computer device in some embodiments;

[0044] Figure 17 This is a diagram showing the internal structure of a computer device in some embodiments. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0046] The image evaluation model processing method and image processing method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on other network servers.

[0047] Specifically, server 104 acquires text features obtained by encoding the text, and acquires image features obtained by encoding the generated image using the text. Server 104 fuses the image features and text features to obtain the image-text fusion features corresponding to the generated image. Server 104 inputs the image-text fusion features into a trained image evaluation model to obtain the target matching degree between the generated image and the text. The text can be sent from terminal 102 to server 104, or the text can be generated by server 104.

[0048] The image evaluation model can be trained using an image evaluation model processing method, which includes: In a fixed-length training phase, server 104 acquires multiple fixed-length first sample image sequences. Each first sample image sequence includes sample images generated using first sample text, and the included sample images are ordered according to their matching degree with the first sample text. The multiple first sample image sequences have the same sequence length. Sequence length refers to the number of included sample images. Server 104, using the image evaluation model, predicts a first predicted matching degree with the first sample text for each sample image in each first sample image sequence. Based on the first predicted matching degree and the order of the sample images in the multiple first sample image sequences, a fixed-length sequence loss value is determined, and the image evaluation model is trained using the fixed-length sequence loss value. In a variable-length training phase, server 104 acquires multiple variable-length second sample image sequences. Each second sample image sequence includes sample images generated using second sample text, and the included sample images are ordered according to their matching degree with the second sample text. Variable-length means that at least two second sample image sequences have different sequence lengths. Server 104 uses an image evaluation model to predict a second predicted matching degree with the second sample text for each sample image in each second sample image sequence. Based on the second predicted matching degree and the sorting order of the sample images in multiple second sample image sequences, it determines a variable-length sequence loss value and uses the variable-length sequence loss value to train the image evaluation model.

[0049] The terminal 102 can be, but is not limited to, a desktop computer, laptop computer, smartphone, tablet computer, IoT device, and portable wearable device. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The cloud server provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal 102 and the server 104 can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0050] The image evaluation model processing method in this application can be implemented based on artificial intelligence (AI) and online media. The solutions provided in the embodiments of this application involve technologies such as machine learning in artificial intelligence, which are specifically illustrated through the following embodiments:

[0051] In some embodiments, such as Figure 2 As shown, an image evaluation model processing method is provided. This method can be executed by a terminal or a server, or it can be executed by both a terminal and a server. This method can be applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 202 to 208. Wherein:

[0052] Step 202: In the fixed-length training phase, multiple first sample image sequences of fixed length are obtained. Each first sample image sequence includes sample images generated using first sample text, and the included sample images are sorted according to their matching degree with the first sample text. The multiple first sample image sequences have the same sequence length, and the sequence length refers to the number of included sample images.

[0053] In this system, the first sample text corresponds to multiple sample images, which are generated using the first sample text. Therefore, sample images are also a type of generated image, meaning images generated based on text. The sequence of first sample images is obtained from the multiple sample images corresponding to the first sample text. The sample images corresponding to the first sample text are used to present the content described by the first sample text. The multiple sample images corresponding to the first sample text can be generated using a text-generated image generation model. The text-generated image generation model is a trained neural network model used to generate images that match the text, so as to present the content described by the text using images. The text-generated image generation model can be any model that can generate images that match the text, such as the stable-diffusion open-source model. The txt2img method in the stable-diffusion open-source model can generate text-matched images. For example, it can generate corresponding plot images based on text describing the plot. For example, if the text describing the plot is "Spring water falls, water droplets hit the mountain rocks, the clear sound breaks the silence among the pines, and also touches the heart of the sleepless person in the pine cabin," and if 9 images need to be generated, then 9 images are generated through 9 random generation processes, such as... Figure 3 As shown, nine images were generated based on the text describing the plot. The sample text can be sent from the terminal to the server or generated by the server.

[0054] The order of sample images in the first sample image sequence is determined by the degree of matching between the sample images and the first sample text. For example, sample images with a higher degree of matching are placed before sample images with a lower degree of matching, that is, the matching degree of the sample images placed earlier is greater than that of the sample images placed later; or, sample images with a higher degree of matching are placed after sample images with a lower degree of matching, that is, the matching degree of the sample images placed earlier is less than that of the sample images placed later.

[0055] The first sample text can be any text. The first sample text can correspond to multiple sample images, and the sample images corresponding to the first sample text are images generated using the first sample text. The first sample text can correspond to a set of multiple sample image sequences, and the set of sample image sequences includes at least one sample image sequence. The sample image sequence includes at least a portion of the sample images corresponding to the first sample text, and the sample images in the sample image sequence are sorted according to the degree of matching between the sample images and the first sample text. This sorting can be from highest to lowest matching degree or from lowest to highest matching degree. For example, if the sample text corresponds to 10 sample images, and the sample image sequence of length 2 includes two sample images from these 10 sample images, namely sample image 1 and sample image 2, if the degree of matching between sample image 1 and the sample text is greater than the degree of matching between sample image 2 and the sample text, and they are sorted from highest to lowest matching degree, the sample image sequence is "sample image 1, sample image 2".

[0056] In this set of multiple sample image sequences, the sequence lengths of the sample image sequences differ across different sets, while the sequence lengths of the sample image sequences within the same set are the same. Sequence length refers to the length of a sample image sequence, i.e., the number of sample images it contains. For example, the first sample text corresponds to M sets of sample image sequences, where sample image sequence set j is the j-th set of these M sets, 1 ≤ j ≤ M. Each sample image sequence in sample image sequence set j has the same sequence length; for example, each sample image sequence in sample image sequence set j has a sequence length of 2, meaning that each sample image sequence in sample image sequence set j contains 2 sample images. The sequence length of the sample image sequence in sample image sequence set j is different from the sequence length of the sample image sequence in sample image sequence set d. Sample image sequence set d is the d-th sample image sequence set among these M sample image sequence sets, and sample image sequence set j and sample image sequence set d are two different sample image sequence sets, 1≤d≤M. For example, the sequence length of each sample image sequence in sample image sequence set j is 2, and the sequence length of each sample image sequence in sample image sequence set d is 3.

[0057] Specifically, the training process of the image evaluation model includes one or more fixed-length training phases, where "multiple" means at least two. The sequence lengths of the sample image sequences used in training differ in different fixed-length training phases, while the sequence lengths of the sample image sequences used in training within the same fixed-length training phase are the same. In each fixed-length training phase, the server acquires multiple first sample image sequences of a fixed length. For example, the server can acquire at least some or all of the sample image sequences from one set of multiple sample image sequences corresponding to the first sample text, thus obtaining multiple first sample image sequences of a fixed length. In this fixed-length training phase, the image evaluation model is trained using the first sample text and these multiple first sample image sequences. Furthermore, the sequence lengths of the first sample image sequences acquired in different fixed-length training phases are different. For example, the training process includes two fixed-length training phases. In the first fixed-length training phase, multiple first sample image sequences with a sequence length of 2 are obtained, and the image evaluation model is trained using the first sample text and the multiple first sample image sequences with a sequence length of 2. In the second fixed-length training phase, multiple first sample image sequences with a sequence length of 3 are obtained, and the image evaluation model is trained using the first sample text and the multiple first sample image sequences with a sequence length of 3.

[0058] In some embodiments, when the training process includes multiple fixed-length training stages, each fixed-length training stage is executed sequentially. The sequence length of the first sample image sequence used in the training of an earlier fixed-length training stage is less than the sequence length of the first sample image sequence used in the training of a later fixed-length training stage. For example, the training process includes two fixed-length training stages. In the first fixed-length training stage, multiple first sample image sequences, each with a sequence length of 2, are obtained, and the image evaluation model is trained using the first sample text and the multiple first sample image sequences with a sequence length of 2. In the second fixed-length training stage, multiple first sample image sequences, each with a sequence length of 3, are obtained, and the image evaluation model is trained using the first sample text and the multiple first sample image sequences with a sequence length of 3. Of course, when the training process includes multiple fixed-length training stages, the sequence length of the first sample image sequence used in the training of an earlier fixed-length training stage can be greater than the sequence length of the first sample image sequence used in the training of a later fixed-length training stage. Because the sequence length of the first sample image sequence used in the initial fixed-length training phase is shorter than that used in the subsequent fixed-length training phase, it is easier to obtain a shorter sequence image sequence compared to a longer sequence image sequence. Therefore, the number of first sample image sequences used in the initial fixed-length training phase is likely to be greater than the number of first sample image sequences used in the subsequent fixed-length training phase. Thus, training the image evaluation model with a shorter sequence image sequence first can expand the learning data space, allowing the image evaluation model to learn the basic sample space distribution. Then, using a shorter sequence image sequence to train the image evaluation model can make the image evaluation model converge stably, thereby improving the training effect of the image evaluation model.

[0059] In some embodiments, each sample text in the sample text set participates in training the image evaluation model, i.e., in the fixed-length training phase. The first sample text is any sample text in the sample text set; multiple sample texts refer to at least two. Each sample text can correspond to multiple sample images, where the sample images are generated using the sample text. Each sample text can correspond to multiple sets of sample image sequences. For example, if N sample texts participate in training the image evaluation model, sample text i is the i-th sample text among these N sample texts, where 1 ≤ i ≤ N. Sample text i corresponds to M sets of sample image sequences. The sequence lengths of the sample image sequences in different sets are different, while the sequence lengths of the sample image sequences in the same set are the same. Each set of sample image sequences includes at least one sample image sequence. The sample image sequence includes at least a portion of the sample images corresponding to the sample text, and the sample images in the sample image sequence are sorted according to the degree of matching between the sample image and the sample text, either from largest to smallest matching degree or from smallest to largest matching degree.

[0060] In some embodiments, in each fixed-length training phase, the sequence length of the sample image sequences used in the training of that fixed-length training phase is determined to obtain the target sequence length corresponding to the fixed-length training phase. For each sample text in the sample text set, a set of sample image sequences corresponding to the target sequence length is determined from multiple sets of sample image sequences corresponding to that sample text. This results in a set of sample image sequences corresponding to the target sequence length for each sample text in the sample text set. The sequence length of the sample image sequences included in the set of sample image sequences corresponding to the target sequence length is the target sequence length. For example, if the target sequence length is 2 and there are 10 sample texts, for each of these 10 sample texts, a set of sample image sequences with a sequence length of 2 is determined from multiple sets of sample image sequences corresponding to that sample text. This results in 10 sets of sample image sequences corresponding to the target sequence length of 2. In the fixed-length training phase, the server uses each sample text in the sample text set and the set of sample image sequences corresponding to the target sequence length for each sample text to train the image evaluation model. For example, for each sample text, the server can obtain some or all of the sample image sequences from the set of sample image sequences corresponding to the target sequence length of the sample text, resulting in multiple fixed-length sample image sequences for that sample text. Based on the sample text and the multiple fixed-length sample image sequences for that sample text, an image evaluation model is trained. During the fixed-length training phase, the sample text set can be divided into multiple subsets, each including a portion of the sample text in the sample text set. For each subset, the parameters of the image evaluation model are adjusted using the sample text in the subset and the multiple fixed-length sample image sequences for that sample text. For example, if the sample text set contains N sample texts, and each subset includes bs sample texts, then N / bs subsets are obtained. After adjusting the parameters of the image evaluation model using these N / bs subsets, one epoch is completed. The first sample text is any sample text in the sample text set. When the sample file is the first sample file, the multiple fixed-length sample image sequences for that sample text are the multiple fixed-length first sample image sequences.

[0061] In some embodiments, the server inputs the first sample text into the text-to-image generation model to generate at least one image matching the first sample text. The image generation by the text-to-image generation model is random, and the images generated each time the first sample text is input into the model may not be exactly the same. The first sample image may be generated by the server, or the first sample text may be sent to the server by the terminal. The terminal displays a text-to-image interface, obtains the text input into the interface, and sends a text-to-image request carrying that text to the server. In response to the text-to-image request, the server uses the text in the request as the first sample text.

[0062] Step 204: Using the image evaluation model, for each sample image in each first sample image sequence, predict the first predicted matching degree with the first sample text. Based on the first predicted matching degree and the sorting order of the sample images in multiple first sample image sequences, determine the fixed-length sequence loss value, and use the fixed-length sequence loss value to train the image evaluation model.

[0063] The image evaluation model is a model that needs to be trained; it may be untrained or trained but require further training. The prediction matching degree is the degree of matching between the sample image predicted by the image evaluation model and the first sample text.

[0064] Specifically, for each first sample image sequence, the server can construct an image-text pair by combining each sample image in the sequence with the first sample text. Each image-text pair includes a sample image and the first sample text. For each image-text pair, the server can encode the first sample text to obtain its sample text features, encode the sample image to obtain its sample image features, and fuse the sample text features and sample image features to obtain fused image-text features. These fused features are then input into the image evaluation model to obtain the first predicted matching degree corresponding to the image-text pair. The output of the image evaluation model can be used as the first predicted matching degree, or the output can be normalized, and the normalized result can be used as the first predicted matching degree. Normalization can be implemented using any normalization function, such as the sigmoid function. The first predicted matching degree corresponding to the image-text pair refers to the matching degree between the sample image and the first sample text in the image-text pair. The server can determine the loss value generated by the first sample image sequence based on the first predicted matching degree corresponding to each image-text pair and the sorting order of the sample images in the first sample image sequence. Having obtained the loss values ​​generated by each of the multiple first sample image sequences, the server performs a comprehensive calculation on these loss values ​​to obtain a fixed-length sequence loss value for the first sample text. This comprehensive calculation includes, but is not limited to, mean calculation or summation calculation. For example, the server can use the mean of the loss values ​​generated by the multiple first sample image sequences as the fixed-length sequence loss value for the first sample text, and can use this fixed-length sequence loss value to adjust the parameters of the image evaluation model. The dimensions of the sample text features and sample image features can be the same, and the fusion can be multiplication or addition at corresponding positions. The dimensions of the sample text features and sample image features can also be different.

[0065] In some embodiments, the first sample text is a sample text from a sample text set. Each sample text in the sample text set participates in the fixed-length training phase. The sample text set is split into multiple subsets. The server can sequentially adjust the parameters of the image evaluation model based on each subset. For example, the server determines a subset from the multiple subsets, determines the fixed-length sequence loss value generated by each sample text in that subset, performs a comprehensive calculation (e.g., mean calculation) on the fixed-length sequence loss values ​​generated by each sample text in that subset to obtain a comprehensive loss value generated by that subset, and uses this comprehensive loss value to adjust the parameters of the image evaluation model. Then, the server returns to the step of determining a subset from the multiple subsets until all the subsets have been traversed, thus completing one round of training. During the fixed-length training phase, the server can perform multiple rounds of training, with each round using the same method.

[0066] In some embodiments, the server can input sample text, such as a first sample text, into a text encoder for encoding to obtain sample text features. The text encoder is a trained neural network. The text encoder can be any neural network capable of encoding text, such as the BLIP model (Bootstrapping Language-Image Pre-training, a unified visual language understanding and generation pre-trained model). The server can also input sample images into an image encoder for encoding to obtain sample image features. The image encoder is a trained neural network. The image encoder can be any neural network capable of encoding images, such as VIT (Vision Transformer). The text encoder can also be a BERT (Bidirectional Encoder Representation from Transformers) model.

[0067] In some embodiments, such as Figure 4 As shown, the image evaluation model consists of a Multi-Layer Perceptron (MLP) and a prediction layer. The MLP comprises K stacked linear activation layers, where K can be set or preset as needed; K is greater than or equal to 1 or 2, for example, K equals 3 or 4, etc. The server normalizes the output of the prediction layer to obtain the prediction matching degree. Each linear activation layer contains one linear layer and one activation layer. The activation layer can be implemented using any activation function; for example, it can be implemented using the dropout function, thus becoming a dropout layer (randomly discarded layer). The number in the dropout layer indicates the proportion of features randomly discarded; for example, Dropout Layer (0.2) indicates a random discard proportion of 0.2. Dropout is a commonly used activation function, its main function being to randomly discard a portion of neurons during neural network training to enhance the network's generalization ability. Of course, other activation functions can also be used to implement the activation layer. Table 1 shows the parameter table of the image evaluation model composed of an MLP structure and a prediction layer, where the MLP structure contains 3 linear activation layers. In Table 1, "channel" represents the output dimension of the linear layer.

[0068] Table 1

[0069]

[0070] For example, the first sample image sequence is "sample image 1, sample image 2, sample image 3", such as... Figure 4 As shown, for the image-text pair generated by the first sample text and sample image 1, the first sample text is input into the text encoder for encoding to obtain sample text features, and sample image 1 is input into the image encoder for encoding to obtain sample image features. Figure 4 The crossed-out pattern within the circle indicates fusion. Then, the sample text features and sample image features are fused to obtain the image-text fused features corresponding to the sample image. The server inputs the image-text fused features into the image evaluation model to obtain the first predicted matching degree between sample image 1 and the first sample text. Similarly, for the image-text pair generated from the first sample text and sample image 2, the first sample text is input into the text encoder for encoding to obtain sample text features, and sample image 2 is input into the image encoder for encoding to obtain sample image features. The sample text features and sample image features are fused to obtain the image-text fused features corresponding to the sample image, and the image-text fused features are input into the image evaluation model to obtain the first predicted matching degree between sample image 2 and the first sample text. The same process applies to the image-text pair generated from the first sample text and sample image 3, and will not be elaborated further here. Figure 4 The double-headed arrow indicates that the model parameters are shared, meaning that there is actually only one image evaluation model.

[0071] In some embodiments, the server can generate image-text fusion features using an image-text encoder. For example... Figure 5 The image-to-text encoder shown includes a text encoder and an image encoder. The server can input a sample image into the image encoder to obtain sample image features, input the first sample text into the text encoder to obtain sample text features, and then fuse the sample image features and sample text features through the text encoder to obtain image-text fused features. The image encoder includes, but is not limited to, the ViT (Vision Transformer) model. The text encoder includes, but is not limited to, the BLIP model or the BERT model. Figure 6 It shows Figure 4 The input part is input part 1. If Figure 4 China adopts Figure 5 The image-text encoder generates image-text fusion features, then Figure 4 The input section 1 can be changed to Figure 6 Input part 2 takes the image-text fusion features output from input part 2 and inputs them into the image evaluation model to predict the matching degree between the first sample text and the sample image.

[0072] In some embodiments, the image evaluation model can be a multi-layer MLP model, where multi-layer means at least two layers, such as four or five layers.

[0073] In some embodiments, the sample text features may include multiple sub-text features, each of which is a vector. The multiple sub-text features are at least two, for example, Nt sub-text features. Nt can be approximately the same length as the sample text. For example, the sample image features are Nt×768 vectors (Nt vectors of length 768), and the sample image features are 1x768 vectors. The sample image features are a single vector. The length of the sample image features is consistent with the length of the sub-text features, for example, both are vectors of length 768. For each sample image in the first sample image sequence, the server can fuse the multiple sub-text features in the sample text features with the sample image features of the sample image, respectively, to obtain multiple sub-image-text fusion features corresponding to the sample image. The server can then perform statistical analysis on these multiple sub-image-text fusion features to obtain the image-text fusion features corresponding to the sample image. This statistical analysis can be mean calculation, maximum value calculation, or minimum value calculation, for example, max pooling, mean pooling, or minimum pooling. The server can also multiply the sub-text features with the sample image features at corresponding positions to obtain the sub-image-text fusion features corresponding to the sample image. Alternatively, the server can add or calculate the mean of the sub-text features and sample image features at corresponding positions to obtain the sub-text-image fusion features corresponding to the sample image.

[0074] In some embodiments, the server can perform max pooling on the sub-image-text fusion features corresponding to the sample image to obtain the image-text fusion features corresponding to the sample image. For example, using a 1×768 vector of sample image features, Nt 1×768 vectors of sample text features are multiplied at the same position (image constraint), resulting in Nt 1x768 vectors, which are the sub-image-text fusion features. Then, max pooling is performed on the resulting Nt 1x768 vectors to obtain a 1x768 vector output, which is the image-text fusion feature. Figure 7As shown, inputting the first sample text into the text encoder generates sample text features containing multiple sub-text features, and inputting the sample image into the image encoder generates sample image features. The crossed-out patterns in the circles represent fusion modules used for fusion, such as multiplying corresponding positions, to obtain multiple sub-image-text fusion features. These multiple sub-image-text fusion features are then input into a pooling layer for pooling, resulting in image-text fusion features. Finally, these image-text fusion features are input into the image evaluation model. The text encoder, image encoder, and fusion module are networks in the CLIP (Contrastive Language-Image Pre-Training) model. By fusing multiple sub-text features from the sample text features with the sample image features of the sample image, multiple sub-image-text fusion features corresponding to the sample image are obtained. Statistical analysis of these multiple sub-image-text fusion features yields the image-text fusion features corresponding to the sample image. This allows for a thorough fusion of image and text features, improving the training performance of the image evaluation model.

[0075] Step 206: In the variable-length training phase, multiple second sample image sequences of variable length are obtained. Each second sample image sequence includes sample images generated using second sample text, and the included sample images are sorted according to their matching degree with the second sample text. Variable length means that at least two second sample image sequences have different sequence lengths.

[0076] The training process of the image evaluation model includes at least one variable-length training phase. This variable-length training phase can occur after any fixed-length training phase. The second sample text can be any text, and the first and second sample texts can be the same text or different texts. The second sample text can be a sample text from the sample text set, and the first and second sample texts can be the same sample text from the sample text set or different sample texts.

[0077] The order of sample images in the second sample image sequence is determined by the degree of matching between the sample images and the second sample text. For example, sample images with a higher degree of matching are placed before sample images with a lower degree of matching, that is, the matching degree of the sample images placed earlier is greater than that of the sample images placed later; or, sample images with a higher degree of matching are placed after sample images with a lower degree of matching, that is, the matching degree of the sample images placed earlier is less than that of the sample images placed later.

[0078] At least two of the plurality of second sample image sequences have different sequence lengths. For example, the plurality of second sample image sequences may include a sample image sequence with a sequence length of 2 and a sample image sequence with a sequence length of 3. The sequence length of the first sample image sequence may be less than or equal to at least one of the sequence lengths of the plurality of second sample image sequences. Of course, the sequence length of the first sample image sequence may also be greater than the sequence lengths of the plurality of second sample image sequences.

[0079] The second sample text can correspond to multiple sample images, which are images generated using the second sample text. The second sample text can also correspond to a set of multiple sample image sequences, each containing at least one sample image sequence. These sequences include at least a subset of the multiple sample images corresponding to the second sample text, and the images are ordered according to their degree of matching with the second sample text. This order can be from highest to lowest matching degree or vice versa. Within the multiple set of sample image sequences corresponding to the second sample text, the sequence lengths of the sample image sequences differ between different sets, while the sequence lengths are the same within the same set.

[0080] Specifically, when the first sample text and the second sample text are the same text, the server can obtain multiple second sample image sequences of variable length based on the multiple sample image sequence sets corresponding to the first sample text. When the first sample text and the second sample text are different texts, the server can obtain multiple second sample image sequences of variable length based on the multiple sample image sequence sets corresponding to the second sample text. For example, the server can determine at least two sample image sequence sets from the multiple sample image sequence sets corresponding to the second sample text, and obtain at least a portion of the sample image sequences from each of these at least two sample image sequence sets to obtain multiple second sample image sequences of variable length. For example, the second sample text corresponds to three sample image sequence sets, namely sample image sequence set 1, sample image sequence set 2, and sample image sequence set 3, where the sequence length of the sample image sequences in sample image sequence set 1 is 2, the sequence length of the sample image sequences in sample image sequence set 2 is 3, and the sequence length of the sample image sequences in sample image sequence set 3 is 4. The server can determine at least two sample image sequence sets from these three sets of sample image sequences. From each determined set, it obtains at least a portion of the sample image sequences, resulting in multiple second sample image sequences of variable length. The at least a portion of the sample image sequences can be one or more sequences, with "multiple" meaning at least two. For example, if sample image sequence set 1 and sample image sequence set 2 are determined, at least a portion of the sample image sequences are obtained from sample image sequence set 1, and at least a portion of the sample image sequences are also obtained from sample image sequence set 2. The resulting sample image sequences are then multiple second sample image sequences of variable length. Since the sequence lengths of the sample image sequences in different sets are different, it is obvious that at least two of the obtained sample image sequences will have different sequence lengths.

[0081] In some embodiments, the training process may include multiple variable-length training phases, and for each sample text, multiple variable-length second sample image sequences may be used for training in different variable-length training phases, which may be the same or different.

[0082] In some embodiments, each sample text in the sample text set participates in the training of the variable-length training phase. The second sample text is a sample text in the sample text set. In the variable-length training phase, at least two specified sequence lengths can be determined. For each sample text, from the multiple sample image sequence sets corresponding to the sample text, the sample image sequence sets corresponding to the at least two specified sequence lengths are determined respectively. From each sample image sequence set corresponding to the at least two specified sequence lengths, at least a portion of the sample image sequence is obtained to obtain multiple sample image sequences of variable length for that sample text. When the sample text is the second sample text, multiple second sample image sequences of variable length are obtained. The sequence length of the sample image sequence in the sample image sequence set corresponding to the specified sequence length is the specified sequence length. In the variable-length training phase, each sample text and multiple sample image sequences of variable length for each sample text can be used to train the image evaluation model, for example, training can be performed in batches, and multiple rounds of training can be performed.

[0083] In some embodiments, the image evaluation model includes a multi-layer perceptron (MLP), and training is performed in batches within each stage. The training process can be as follows: 1) Parameter initialization: The multi-layer perceptron in the image evaluation model is initialized using a Gaussian normal distribution with parameters (0, 1). 2) Setting learning parameters: All modules in the image evaluation model are trained using a learning rate of 0.005 and the SGD (stochastic gradient descent) gradient update method. The learning rate is reduced to 0.1 times its original value every 10 epochs. 3) Training method: Assume that the total training data consists of N data points, which correspond to N sample texts. In the fixed-length training phase, one data point refers to a sequence of multiple sample images of a fixed length corresponding to one sample text. In the variable-length training phase, one data point refers to a sequence of multiple sample images of a variable length corresponding to one sample text. The N data points are trained in batches, with each batch consisting of 1s data points. There are a total of N / 1s batches. Each N / 1s batch is completed, which represents the completion of one epoch. Each epoch processes the entire set of N data points once, until the average epoch loss no longer decreases at a certain epoch. Specifically, for each sample text in the `bs` data sets (e.g., multiple sample image sequences of variable length corresponding to the sample text), a loss value is calculated for each sequence. The average of these individual loss values ​​is then used as the loss value for the variable-length sequence generated by the sample text. Finally, the average of these loss values ​​is calculated to obtain the total loss value for the `bs` data sets, which represents the total loss for one batch. The parameters of the image evaluation model are then adjusted using this total loss value.

[0084] Training process under a certain batch of data: (1) Forward model (forward calculation): Each sample image in the sample image sequence is paired with the sample text to form an image-text pair. The sample images and sample text in the image-text pair are encoded and fused to obtain the image-text fusion feature. The image-text fusion feature is input into the image evaluation model. The output of the image evaluation model is normalized to obtain the predicted matching degree between the sample text and the sample image in the image-text pair. Normalization can be achieved by using a normalization function, such as the sigmoid function, to normalize the output of the image evaluation model to 0~1 to obtain the predicted matching degree of the sample image. (2) Calculate the loss based on the predicted matching degree. (3) Backward model (backward calculation): The image evaluation model backpropagates the loss to the network in the image evaluation model to calculate the gradient of each parameter of the network. (4) Update the model parameters: Update the parameters of the image evaluation model based on the gradient of each parameter in the network.

[0085] Step 208: Using the image evaluation model, for each sample image in each second sample image sequence, predict the second predicted matching degree with the second sample text. Based on the second predicted matching degree and the sorting order of the sample images in multiple second sample image sequences, determine the variable-length sequence loss value, and train the image evaluation model using the variable-length sequence loss value.

[0086] Specifically, for each second sample image sequence, the server can construct an image-text pair with each sample image in the sequence and the second sample text. Each image-text pair includes one sample image and the second sample text. For each image-text pair, the server can encode the second sample text to obtain its sample text features, encode the sample image to obtain its sample image features, and fuse the sample text features and sample image features to obtain image-text fusion features. These fusion features are then input into the image evaluation model to obtain the second predicted matching degree corresponding to the image-text pair. This second predicted matching degree refers to the matching degree between the sample image and the second sample text in the image-text pair. The server can determine the loss value generated by the second sample image sequence based on the second predicted matching degree corresponding to each image-text pair and the sorting order of the sample images in the second sample image sequence. Having obtained the loss values ​​generated by each of the multiple second sample image sequences, the server performs a comprehensive calculation on these loss values ​​to obtain a variable-length sequence loss value for the second sample text. This comprehensive calculation includes, but is not limited to, mean calculation or summation calculation. For example, the server can use the average of the loss values ​​generated by the multiple second sample image sequences as the loss value of the variable-length sequence generated by the second sample text.

[0087] In some embodiments, the first sample text is a sample text from a sample text set. Each sample text in the sample text set participates in the training of the variable-length training phase. The sample text set is split into multiple subsets. The server can sequentially adjust the parameters of the image evaluation model according to each subset. For example, the server determines a subset from the multiple subsets, determines the variable-length sequence loss value generated by each sample text in the subset, performs a comprehensive calculation (e.g., mean calculation) on the variable-length sequence loss values ​​generated by each sample text in the subset to obtain the comprehensive loss value generated by the subset, and adjusts the parameters of the image evaluation model using the comprehensive loss value generated by the subset. Then, the server returns to the step of determining a subset from the multiple subsets until all the multiple subsets are traversed, thereby completing one round of training. In the fixed-length training phase, the server can perform multiple rounds of training, with each round using the same method.

[0088] In the above image evaluation model processing method, during the fixed-length training phase, multiple first sample image sequences of the same length are used to train the image evaluation model. The sample images are sorted according to their matching degree. For each sample image in each first sample image sequence, a first predicted matching degree with the first sample text is predicted. Based on the first predicted matching degree and the sorting order of the sample images, a fixed-length sequence loss value is determined. The image evaluation model is trained using the fixed-length sequence loss value. Thus, through training, the image evaluation model can learn the ability to predict, i.e., evaluate, the matching degree between the sample image and the first sample text. During the variable-length training phase, multiple second sample image sequences of variable length are used to train the image evaluation model. For each sample image in each second sample image sequence, a second predicted matching degree with the second sample text is predicted. Based on the second predicted matching degree and the sorting order of the sample images in the multiple second sample image sequences, a variable-length sequence loss value is determined. The image evaluation model is trained using the variable-length sequence loss value. This allows the image evaluation model to be further adjusted so that it can learn the ability to predict, i.e. evaluate, the second predicted matching degree between the sample image and the second sample text for second sample image sequences of different lengths. This can further improve the evaluation accuracy of the image-text matching degree.

[0089] In some embodiments, a set of multiple sample image sequences determined based on multiple sample images corresponding to a first sample text is obtained. The sample images are images generated using the first sample text. The sequence lengths of the sample image sequences in different sample image sequence sets are different, while the sequence lengths of the sample image sequences in the same sample image sequence set are the same. The sample image sequence includes at least a portion of the sample images from the multiple sample images, and the sample images in the sample image sequence are sorted according to their degree of matching with the first sample text. Obtaining multiple first sample image sequences of a fixed length includes: obtaining a portion or all of the sample image sequences from one of the multiple sample image sequence sets to obtain multiple first sample image sequences of a fixed length.

[0090] Each set of sample image sequences corresponds one-to-one with a specified sequence length. The sequence length of the sample image sequences in the set of sample image sequences corresponding to the specified sequence length is that specified sequence length. The specified sequence length can be set as needed. The set of multiple sample image sequences determined based on the multiple sample images corresponding to the first sample text is the set of multiple sample image sequences corresponding to the first sample text.

[0091] The degree of matching between each sample image and the first sample text can be determined based on the image-text matching degrees determined by multiple objects. Each object's determination of the image-text matching degree is the degree of matching between each sample image and the first sample text as determined by that object. For each sample image, the average of the matching degrees determined by the multiple objects can be used as the matching degree between that sample image and the first sample text. These multiple objects can be multiple individuals or multiple trained image-text matching evaluation models. These trained image-text matching evaluation models are used to predict the matching degree between images and text. Any two trained image-text matching evaluation models may have different model structures, different training data, or different training methods. The accuracy of the evaluation results of each trained image-text matching evaluation model is greater than or equal to an accuracy threshold, which can be set as needed, such as 95% or 96%. For example, consider three objects and ten sample images (images 1 through 10). For image 1, if the matching degree between image 1 and the first sample text is determined to be a, b, and c respectively, then the average of a, b, and c is taken as the matching degree between image 1 and the first sample text. The matching degree can be represented numerically, for example, using a value between 0 and 1; a larger value indicates a greater matching degree. The matching degree reflects the similarity between the content presented in the sample image and the content described in the sample text. A higher matching degree indicates a more similar content between the sample image and the sample text. A higher matching degree can be understood as a better image quality.

[0092] Specifically, the server can determine multiple specified sequence lengths. For each specified sequence length, it can obtain a specified number of sample images from multiple sample images corresponding to the first sample text, and sort the specified number of sample images according to the degree of matching to obtain a sequence of sample images corresponding to the specified sequence length. This method can be used to obtain multiple sequence of sample images corresponding to the specified sequence length, and these multiple sequence of sample images corresponding to the specified sequence length can be combined to form a set of sequence of sample images corresponding to the specified sequence length.

[0093] In some embodiments, the server determines the target sequence length corresponding to the fixed-length training phase, which can be set as needed. The target sequence length is one of the multiple specified sequence lengths. The server determines the set of sample image sequences corresponding to the target sequence length from the multiple sample image sequences, and obtains some or all of the sample image sequences from the set of sample image sequences corresponding to the target sequence length to obtain multiple first sample image sequences of fixed length.

[0094] In this embodiment, since the sequence lengths of sample image sequences in different sample image sequence sets are different, and the sequence lengths of sample image sequences in the same sample image sequence set are the same, the sequence lengths of each sample image sequence obtained from the same sample image sequence set are all the same. Therefore, multiple first sample image sequences of a fixed length can be quickly obtained from the same sample image sequence set.

[0095] In some embodiments, the first sample text and the second sample text are the same text. Obtaining multiple second sample image sequences of variable length includes: determining at least two sample image sequence sets from multiple sample image sequence sets; and obtaining a portion or all of the sample image sequences from each of the at least two sample image sequence sets to obtain multiple second sample image sequences of variable length.

[0096] Specifically, in the variable-length training phase, the server can determine at least two specified sequence lengths from multiple specified sequence lengths, determine the sample image sequence sets corresponding to the at least two specified sequence lengths from the multiple sample image sequence sets, and obtain some or all sample image sequences from the sample image sequence sets corresponding to the at least two specified sequence lengths respectively, thus obtaining multiple second sample image sequences of variable length. Each obtained sample image sequence is a multiple second sample image sequence of variable length.

[0097] In this embodiment, since the sequence lengths of sample image sequences in different sample image sequence sets are different, and the sequence lengths of sample image sequences in the same sample image sequence set are the same, multiple second sample image sequences of variable length can be quickly obtained from multiple sample image sequence sets.

[0098] In some embodiments, obtaining a set of multiple sample image sequences determined based on multiple sample images corresponding to the first sample text includes: obtaining multiple different global sequences obtained by sorting the multiple sample images according to the image-text matching degree determined by each of the multiple objects, wherein the image-text matching degree is the matching degree between each sample image and the first sample text; determining a set of sample image sequences corresponding to multiple specified sequence lengths based on the multiple different global sequences, wherein the sequence length of the sample image sequence in the set of sample image sequences is the corresponding specified sequence length, and the relative sorting order of the sample images in the sample image sequence is consistent with the relative sorting order in each global sequence.

[0099] While each sample image corresponding to the first sample text is a generated image used to present the content described by the first sample text, there are differences between different sample images. Therefore, the degree of matching between the sample images and the first sample text is also inconsistent. The degree of matching is used to reflect the similarity between the content presented by the sample image and the content described by the first sample text. The higher the degree of matching, the more similar the content presented by the sample image is to the content described by the first sample text.

[0100] These multiple objects correspond one-to-one with the global sequence, and "multiple objects" refers to at least two people. If there are multiple people, these multiple people can be any number of individuals, such as those with a good eye for images or experts who are highly perceptive and can accurately judge the matching degree between images and text. Each person's image-text matching degree corresponds one-to-one with their assessment of the matching degree between each sample image and the first sample text. In other words, the matching degree between the sample images and the first sample text in the image-text matching degree is the degree of matching determined by that person. The matching degree can be represented numerically, for example, using a value between 0 and 1, with a larger value indicating a greater matching degree. Due to differences in aesthetic judgment, the determined matching degree may vary for the same sample text. The global sequence includes multiple sample images corresponding to the first sample text.

[0101] For example, consider three people and ten sample images (Image 1-Image 10). Since these three people have different aesthetic preferences, the global sequences obtained by sorting the images according to their respective assessments of image-text matching are as follows: Global Sequence 1: "Image 1, Image 3, Image 5, Image 8, Image 9, Image 10, Image 7, Image 6, Image 4, Image 2"; Global Sequence 2: "Image 1, Image 5, Image 8, Image 7, Image 3, Image 9, Image 10, Image 2, Image 4, Image 6"; Global Sequence 3: "Image 2, Image 1, Image 5, Image 9, Image 8, Image 10, Image 4, Image 7, Image 6, Image 2". A higher matching degree can be interpreted as a better image quality. Therefore, the global sequences can be sorted from best to worst image quality, or from worst to best image quality.

[0102] Each sample image sequence includes at least a subset of the sample images from the plurality of sample images, and the relative order of the sample images in the sample image sequence is consistent with the relative order in each global sequence. The relative order refers to considering only the order of a few elements within the whole, regardless of their overall order within the whole. For example, for the five letters a, b, c, d, e, if the relative order of a and b is required to be a before b, then it is sufficient for a to be placed before b; the order of a and b within the five letters is irrelevant. For instance, in global sequences 1 through 3, images 1, 5, and 9 have the same relative order in all three sequences; therefore, the sequence "image 1, image 5, image 9" can be considered a sample image sequence of length 3. The sample images in a sample image sequence can be continuous or discontinuous within the global sequence. Different sample image sequences can have the same or different lengths. Each sample image sequence contains at least two sample images.

[0103] Since the relative sorting order of the sample images in the sample image sequence is consistent with the relative sorting order in each global sequence, and the sample images in each global sequence are sorted according to the degree of matching, the sample images in the sample image sequence are also sorted according to the degree of matching. When the sample images in each global sequence are sorted from largest to smallest in terms of degree of matching, the sample images in the sample image sequence are also sorted from largest to smallest in terms of degree of matching. When the sample images in each global sequence are sorted from smallest to largest in terms of degree of matching, the sample images in the sample image sequence are also sorted from smallest to largest in terms of degree of matching.

[0104] Since the relative order of the sample images in the sample image sequence is consistent with their relative order in each global sequence, for any two sample images in the sample image sequence, such as sample image 1 and sample image 2, if sample image 1 is listed before sample image 2 in each global sequence, it means that all the users determined that the matching degree of sample image 1 is greater than that of sample image 2, or that the matching degree of sample image 1 is less than that of sample image 2. Because different people may have different matching degrees for the same sample image, the sorting of sample images in the sample image sequence according to their matching degree can be understood as sorting them according to the magnitude of their matching degree; sample images with a higher matching degree are listed before sample images with a lower matching degree, or vice versa.

[0105] Specifically, the server can obtain the image-text matching degree determined by each of the multiple individuals. This matching degree includes the degree of matching between each sample image and the first sample text. For each matching degree, the server can sort the multiple sample images according to their matching degree with the first sample text, generating a global sequence. For example, the multiple sample images can be sorted in ascending order of matching degree to generate a global sequence; or, they can be sorted in descending order of matching degree to generate a global sequence.

[0106] In some embodiments, for each specified sequence length among multiple specified sequence lengths, the server can determine at least two sample images whose relative sorting order is consistent in each global sequence, and sort the sample images of the specified sequence length according to their relative sorting order in the global sequence to obtain a sample image sequence corresponding to the specified sequence length. The at least one sample image sequence corresponding to the specified sequence length obtained in this way constitutes a set of sample image sequences corresponding to the specified sequence length. For example, if the specified sequence length is 3, and images 1, 5, and 9 in global sequences 1 to 3 have the same relative sorting order, then the sequence "image 1, image 5, image 9" can be used as the sample image sequence corresponding to the specified sequence length 3.

[0107] In this embodiment, multiple different global sequences are obtained by sorting multiple sample images according to the image-text matching degree determined by multiple individuals. The image-text matching degree is the degree of matching between each sample image and the sample text. Based on the multiple different global sequences, a set of sample image sequences corresponding to multiple specified sequence lengths is determined. The relative sorting order of the sample images in the sample image sequence is consistent with the relative sorting order in each global sequence. Therefore, the relative sorting order of each sample image in the sample image sequence is the order jointly determined by multiple individuals. Thus, the relative sorting order of each sample image in the sample image sequence conforms to human cognition. Therefore, by training the image evaluation model using the sample image sequences in the sample image sequence set, the image evaluation model can learn knowledge that conforms to human cognition. This ensures that when the image evaluation model is used to evaluate the matching degree between the image and the text, i.e., the image-text matching degree, the evaluation result obtained conforms to human cognition and human aesthetic standards, thereby improving the evaluation accuracy of the image-text matching degree.

[0108] In some embodiments, determining a set of sample image sequences corresponding to multiple specified sequence lengths based on multiple different global sequences includes: determining a common sequence that satisfies a preset condition based on multiple different global sequences, wherein the preset condition is that the relative sorting order of each sample image contained therein is consistent in each global sequence, and the relative sorting order is no longer consistent in each global sequence after adding any sample image; for each specified sequence length among multiple specified sequence lengths, obtaining at least one subsequence with a specified sequence length from each common sequence whose sequence length is greater than or equal to the specified sequence length, and forming a set of sample image sequences corresponding to the specified sequence length.

[0109] The common sequence is a sequence that satisfies a preset condition. This condition requires that the relative order of the included sample images in each global sequence be consistent, and that adding any sample image disrupts this consistency. Therefore, multiple different common sequences may exist. "Adding any sample image disrupts the consistency of the relative order in each global sequence" means that after adding any sample image from the multiple sample images corresponding to the first sample text to the common sequence, the common sequence no longer satisfies the requirement that the relative order of the included sample images in each global sequence be consistent. It should be noted that the added sample image is one that was not originally present in the common sequence.

[0110] A subsequence in a common sequence is a sequence formed by at least two consecutive sample images in the common sequence, and the order of the sample images in the subsequence is the same as their order in the common sequence.

[0111] Specifically, the length of the common sequence can be greater than or equal to 2. For each specified sequence length, a common sequence with a sequence length greater than or equal to that specified sequence length is determined. The server can obtain each subsequence with a sequence length of the specified sequence length from each common sequence with a sequence length greater than or equal to that specified sequence length. The server can use the subsequence as the sample image sequence corresponding to the specified sequence length. Alternatively, since it is possible to obtain the same subsequence from different common sequences, the obtained subsequences can be deduplicated, and each remaining subsequence after deduplication can be used as the sample image sequence corresponding to the specified sequence length. The server then uses all the sample image sequences corresponding to the specified sequence length to form a set of sample image sequences corresponding to the specified sequence length. For example, if the specified sequence length is 3, based on global sequence 1 to global sequence 3, the common sequences that meet the preset conditions and have a sequence length greater than or equal to 3 are: "Image 1, Image 5, Image 9, Image 10, Image 2", "Image 1, Image 5, Image 9, Image 10, Image 4", "Image 1, Image 5, Image 9, Image 10, Image 6", "Image 1, Image 5, Image 8, Image 10, Image 2", "Image 1, Image 5, Image 8, Image 10, Image 4", "Image 1, Image 5, Image 8, Image 10, Image 6", "Image 3, Image 9, Image 10, Image 2", "Image 3, Image 9, Image 10, Image 4", and "Image 3, Image 9, Image 10, Image 6". For each common sequence with a sequence length greater than or equal to 3, obtain the subsequences with a sequence length of 3. After removing duplicates from the subsequences with a sequence length of 3, the remaining subsequences are used as sample image sequences corresponding to sequence length 3. Each subsequence is a sample image sequence. Then, the sample image sequences corresponding to sequence length 3 are combined to form a set of sample image sequences corresponding to sequence length 3.

[0112] In some embodiments, the first sample text may be a prompt word, such as Figure 8 As shown, the server can use a text-based image generation model to generate multiple images for prompt words, then obtain the results of multiple people sorting these multiple images, and then generate multiple common sequences based on the sorting results, and generate multiple sample image sequences based on at least some of the common sequences.

[0113] In some embodiments, to obtain a subsequence with a specified sequence length from a common sequence whose sequence length is greater than or equal to the specified sequence length, the server can determine a starting sample image from the common sequence whose sequence length is greater than or equal to the specified sequence length in a forward-to-back order, obtain a subsequence whose length starts from the starting sample image and is consistent with the specified sequence length, update the starting sample image using sample images from the obtained subsequence, for example, using any sample image after the first sample image in the obtained subsequence as the new starting sample image, such as using the last sample image in the subsequence as the new starting sample image to update the starting sample image, and return to the previous execution to obtain a subsequence whose length starts from the starting sample image and is consistent with the specified sequence length. If the sequence length of the subsequence from the starting sample image to the end of the sample image in the common sequence is less than the specified sequence length, then a subsequence whose length is consistent with the specified sequence length is obtained from the common sequence, including the starting sample image and all sample images after the starting sample image, that is, the subsequence whose length is consistent with the specified sequence length is obtained by using sample images before the starting sample image to complete the subsequence.

[0114] Taking a specified sequence length of 3, where the last sample image in a subsequence is the new starting sample image, as an example, for the common sequence "image 1, image 5, image 9, image 10, image 2", the resulting subsequences are "image 1, image 5, image 9" and "image 9, image 10, image 2"; for the common sequence "image 3, image 9, image 10, image 2", the resulting subsequences are "image 3, image 9, image 10" and "image 9, image 10, image 2". It can be seen that "image 9, image 10, image 2" appears twice, therefore, it is necessary to remove duplicates from "image 9, image 10, image 2" to obtain the remaining "image 1, image 5, image 9", "image 9, image 10, image 2", and "image 3, image 9, image 10".

[0115] In this embodiment, since the preset condition is that the relative sorting order of each sample image in each global sequence is consistent, and the relative sorting order in each global sequence is no longer consistent after adding any sample image, it is possible to make the relative sorting order of each sample image in the common sequence consistent in both the common sequence and the global sequence, and to make the common sequence as long as possible. This allows for a wider range of values ​​for the specified sequence length, enabling flexible selection of the specified sequence length and allowing for the selection of as many specified sequence lengths as possible, thereby improving the flexibility and diversity of the obtained sample image sequence.

[0116] In some embodiments, determining a fixed-length sequence loss value based on a first predicted matching degree and the sorting order of sample images in a plurality of first sample image sequences includes: for each first sample image sequence, determining the loss value generated by the first sample image sequence according to the sorting order of sample images in the first sample image sequence and the first predicted matching degree of the sample images; and comprehensively calculating the loss values ​​generated by each first sample image sequence to obtain a fixed-length sequence loss value.

[0117] The first predicted matching degree of the sample image refers to the first predicted matching degree between the sample image and the first sample image. The comprehensive calculation can be a mean calculation or a summation calculation, etc. When each sample text in the sample text set participates in training, and the first sample text is a sample text in the sample text set, a fixed-length sequence loss value is obtained for each sample text. The fixed-length sequence loss value represents the loss value generated by multiple fixed-length sample image sequences.

[0118] Specifically, the server calculates the mean of the loss values ​​generated by each first sample image sequence and uses the calculated mean as the fixed-length sequence loss value. Alternatively, the server can sum the difference in loss values ​​generated by each first sample image sequence and use the sum as the fixed-length sequence loss value.

[0119] In some embodiments, the process of calculating the loss value of a variable-length sequence is the same as the process of calculating the loss value of a fixed-length sequence. Specifically, for each second sample image sequence of the second sample text, the loss value generated by the second sample image sequence is determined based on the sorting order of the sample images in the second sample image sequence and the first predicted matching degree of the sample images in the second sample image sequence. The loss values ​​generated by each second sample image sequence are then comprehensively calculated, for example, by averaging or summing, to obtain the loss value of the variable-length sequence.

[0120] In this embodiment, the loss values ​​generated by each first sample image sequence are comprehensively calculated to obtain a fixed-length sequence loss value. This allows the fixed-length sequence loss value to represent the loss value generated by multiple fixed-length first sample image sequences. By training the image evaluation model using the fixed-length sequence loss value, the image evaluation model can learn to correctly evaluate the matching degree between images and text.

[0121] In some embodiments, determining the loss value generated by the first sample image sequence based on the sorting order of the sample images in the first sample image sequence and the first predicted matching degree of the sample images includes: determining at least one pair of adjacent sample images in the first sample image sequence based on the sorting order of the sample images in the first sample image sequence, wherein the pair of adjacent sample images includes two adjacent sample images in the first sample image sequence; for each pair of adjacent sample images, determining the difference between the first predicted matching degrees of the sample images in the pair of adjacent sample images to obtain the matching degree difference corresponding to the pair of adjacent sample images; and obtaining the loss value generated by the first sample image sequence based on the matching degree difference corresponding to each pair of adjacent sample images.

[0122] The adjacent sample image pair includes two adjacent sample images in the first sample image sequence. For example, if the first sample image sequence is "image 3, image 9, image 10, image 2", then each adjacent sample image pair in the first sample image sequence is "image 3, image 9", "image 9, image 10", and "image 10, image 2".

[0123] Specifically, for each pair of adjacent sample images, the server can subtract the first predicted matching degree of the less matched sample image from the first predicted matching degree of the more matched sample image in the pair to obtain the matching degree difference for that pair. The loss value generated by the first sample image sequence is negatively correlated with the matching degree difference for each pair of adjacent sample images. The loss value of the fixed-length sequence is positively correlated with the loss value generated by the first sample image sequence, thus the fixed-length sequence loss value is negatively correlated with the matching degree difference. Therefore, the larger the matching degree difference, the smaller the fixed-length sequence loss value. Since the larger the matching degree difference, the more likely the first predicted matching degree of the more matched sample image is to be greater than the first predicted matching degree of the less matched sample image, the parameters of the image evaluation model can be updated in a way that reduces the fixed-length sequence loss value. For example, stochastic gradient descent can be used to update the parameters of the image evaluation model to increase the matching degree difference and make the first predicted matching degree of the more matched sample image... The matching degree is greater than the first predicted matching degree of the sample image with a lower matching degree. This makes the first predicted matching degree obtained by the image evaluation model conform to the image-text matching degree determined by the multiple people. For example, if the multiple people believe that the matching degree between sample image 1 and the first sample text is greater than that between sample image 2 and the first sample text, by adjusting the parameters of the image evaluation model, the first predicted matching degree between sample image 1 and the sample text can be made greater than that between sample image 2 and the first sample text. This makes the result obtained by the image evaluation model conform to human aesthetic standards, that is, conform to human preferences, and make it more in line with actual needs.

[0124] In some embodiments, the server can obtain a minimum matching degree, which is a lower limit of the matching degree. The minimum matching degree can be set as needed, for example, to 0. The server can subtract the minimum matching degree from the first predicted matching degree of the sample image with the lowest matching degree in the first sample image sequence to obtain the matching degree difference between the sample image with the lowest matching degree and the minimum matching degree. The loss value generated by the first sample image sequence is negatively correlated with the matching degree difference corresponding to each pair of adjacent sample images, and the loss value generated by the first sample image sequence is negatively correlated with the matching degree difference between the sample image with the lowest matching degree and the minimum matching degree.

[0125] In some embodiments, the server can perform a weighted summation of the differences in matching degree to obtain a weighted result. The weight used for weighting is greater than 0. The loss value generated by the first sample image sequence is determined based on the weighted result. The loss value generated by the first sample image sequence is negatively correlated with the weighted result.

[0126] The matching degree differences include: the matching degree difference corresponding to each pair of adjacent sample images, and the matching degree difference between the sample image with the smallest matching degree and the minimum matching degree. For example, the first sample image sequence is "Image 3, Image 9, Image 10, Image 2", and the sample images in the first sample image sequence are sorted in descending order of matching degree. The first predicted matching degree of Image 3 is p1, the first predicted matching degree of Image 9 is p2, the first predicted matching degree of Image 10 is p3, the first predicted matching degree of Image 2 is p4, and the minimum matching degree is Pmin. p4-Pmin=a1, p1-p2=a2, p2-p3=a3, p3-p4=a4, then the matching degree differences are a1~a4.

[0127] In some embodiments, the process of determining the loss value generated by the second sample image sequence is the same as the process of determining the loss value generated by the first sample image sequence, and will not be described again here.

[0128] In this embodiment, the loss value generated by the first sample image sequence is obtained based on the matching degree difference corresponding to each adjacent sample image pair. Thus, by training the image evaluation model, the image evaluation model can learn to predict the correct matching degree for different sample images.

[0129] In some embodiments, the loss value generated by the first sample image sequence includes a first loss value. The loss value generated by the first sample image sequence is obtained based on the matching degree difference corresponding to each adjacent sample image pair, including: for each adjacent sample image pair, normalizing the matching degree difference corresponding to the adjacent sample image pair to obtain the normalized matching degree difference corresponding to the adjacent sample image pair; and obtaining the first loss value generated by the first sample image sequence based on the normalized matching degree difference corresponding to each adjacent sample image pair.

[0130] Normalization can be implemented using any normalization function, including but not limited to the sigmoid function. The normalized difference in matching degree is negatively correlated with the first loss value. The normalized difference in matching degree is positively correlated with the difference in matching degree.

[0131] Specifically, for each normalized match difference, the server can map the normalized match difference to a value within a preset numerical range to obtain a mapped value. The mapped value is positively correlated with the normalized match difference. Based on each mapped value, a first loss value is obtained, which is negatively correlated with each mapped value. For example, the mapped values ​​can be summed, and the negative of the sum can be used as the first loss value. All values ​​within the preset numerical range are less than 0. The preset numerical range is, for example, (-inf, 0), where -inf represents negative infinity.

[0132] For example, the formula for calculating the first loss value is: (1). Among them, Represents the first loss value. Represents the sequence length. This represents the i-th matching degree difference among all matching degree differences. For the logsigmoid function, , This represents the sigmoid function. The function maps an input numeric value to a value in the range (-inf, 0). -inf represents negative infinity. The sigmoid function maps an input numeric value to the range 0~1. For example... Figure 9 The graph shown is a curve of the logsigmoid function. In formula (1), The larger, the better The smaller the value, the better it satisfies the optimization direction of loss minimization. This can be achieved through training. The probability of obtaining a larger value increases, thus making The higher the probability of obtaining an integer, the better it satisfies the requirement of constant positive difference.

[0133] Wherein, constant positive difference means that the matching degree difference is positive. For example, the first sample image sequence is "Image 3, Image 9, Image 10, Image 2", and the sample images in the first sample image sequence are sorted in descending order of matching degree. The first predicted matching degree of Image 3 is p1, the first predicted matching degree of Image 9 is p2, the first predicted matching degree of Image 10 is p3, the first predicted matching degree of Image 2 is p4, the minimum matching degree is Pmin, p4-Pmin=a1, p1- p2=a2, p2- p3=a3, p3-p4=a4, then the matching degree differences are a1 to a4. For the difference to be constant positive, it is required that a1>0, a2>0, a3>0, a4>0, that is, p4<p3<p2<p1 is satisfied, wherein min score <p4<p3<p2<p1< max score, min score refers to Pmin, max score refers to the maximum matching degree, that is, the upper limit of matching degree, for example, the maximum matching degree is 1.

[0134] In this embodiment, normalizing the matching degree differences corresponding to adjacent sample image pairs and obtaining the first loss value generated by the first sample image sequence based on the normalized matching degree differences respectively corresponding to each adjacent sample image pair can simplify calculation.

[0135] In some embodiments, the loss value generated by the first sample image sequence includes a second loss value. Obtaining the loss value generated by the first sample image sequence based on the matching degree differences respectively corresponding to each adjacent sample image pair includes: for each adjacent sample image pair, performing logarithmic transformation on the matching degree difference corresponding to the adjacent sample image pair to obtain the transformed matching degree difference corresponding to the adjacent sample image pair; taking the matching degree difference respectively corresponding to each adjacent sample image pair as the weight of the corresponding transformed matching degree difference, performing weighted calculation on the transformed matching degree differences respectively corresponding to each adjacent sample image pair, and obtaining the second loss value generated by the first sample image sequence based on the weighted result.

[0136] Wherein, the matching degree difference is positively correlated with the transformed matching degree difference. The second loss value is negatively correlated with the weighted result.

[0137] Specifically, the server may perform logarithmic transformation on the matching degree difference corresponding to the adjacent sample image pair to obtain a logarithmic value, which is the transformed matching degree difference corresponding to the adjacent sample image pair.

[0138] In some embodiments, the server may adjust the matching degree difference corresponding to the adjacent sample image pair by using a preset adjustment parameter to obtain an adjusted matching degree difference, and perform non-linear transformation on the adjusted matching degree difference to obtain the transformed matching degree difference corresponding to the adjacent sample image pair.

[0139] For example, the calculation formula of the second loss value is: (2). Among them, Represents the second loss value. To preset adjustment parameters, To avoid numerical errors The value can be set as needed, for example, to 1e-5. If the value is negative, then the second loss value generated by the first sample image sequence is determined to be 0. If the value is not negative, the second loss value is calculated according to formula (2). Formula (2) is a variant of the cross-entropy loss function. In formula (1), when any For example When the value is maximized, it will result in a maximum loss (the dominant loss), making it difficult to optimize the losses from other differences, which will increase the loss. However, in formula (2), the following formula is used: As a weight, - The smaller The larger the value, the smaller the second loss value, which allows for different... The difference between the two loss values ​​is not too large. The loss value generated by the first sample image sequence is positively correlated with the second loss value. During training, the image evaluation model is adjusted to reduce the second loss value, thereby reducing the loss value generated by the first sample image sequence. This allows the training process to... Increase The larger, the different The greater the likelihood that the difference between them will not be too large.

[0140] In this embodiment, a logarithmic transformation is performed on the matching degree difference corresponding to adjacent sample image pairs to obtain the transformed matching degree difference corresponding to adjacent sample image pairs. The matching degree difference corresponding to each adjacent sample image pair is used as the weight of the transformed matching degree difference corresponding to the corresponding adjacent sample image pair. The transformed matching degree difference corresponding to each adjacent sample image pair is weighted and calculated. Based on the weighted result, the second loss value generated by the first sample image sequence is obtained. Thus, the cross-entropy loss function or a variant of the cross-entropy loss function can be used to calculate the second loss value, so that the difference between different matching degree differences is not too large, thereby adjusting the relative distribution of the matching degree difference and improving the rationality of the matching degree difference.

[0141] In some embodiments, the loss value generated by the first sample image sequence includes a third loss value. The loss value generated by the first sample image sequence is obtained based on the matching degree difference corresponding to each adjacent sample image pair. This includes: determining the matching degree difference between the first predicted matching degree of the sample image with the smallest matching degree in the first sample image sequence and the minimum matching degree; summing each matching degree difference and the minimum matching degree to obtain a summation result; and obtaining the third loss value generated by the first sample image sequence based on the summation result. The third loss value is negatively correlated with the summation result.

[0142] Specifically, the server can determine the sample image with the lowest matching degree based on the sorting order of the sample images in the first sample image sequence. For example, if the sample images in the first sample image sequence are arranged in descending order of matching degree, then the last sample image in the first sample image sequence is the sample image with the lowest matching degree.

[0143] In some embodiments, the server can sum the difference between the first predicted matching degree of the sample image with the lowest matching degree and the minimum matching degree, and the matching degree difference corresponding to each pair of adjacent sample images, to obtain a summation result.

[0144] In some embodiments, the server can normalize the summation result to obtain a normalized summation result, perform logarithmic calculation on the normalized summation result to obtain the corresponding logarithmic value, and obtain the third loss value generated by the first sample image sequence based on the corresponding logarithmic value of the normalized summation result. The third loss value is negatively correlated with the corresponding logarithmic value of the normalized summation result. For example, the formula for calculating the third loss value can be: (3). For the logsigmoid function. In formula (3), In fact with For example, if the sequence length of the first sample image sequence is 4, then p4-Pmin=a1, p1-p2=a2, p2-p3=a3, p3-p4=a4. . The first predicted match score is the highest-matching sample image in the first sample image sequence, which is reduced during training. In fact, it increases This can increase , making Taking a relatively large value makes the value more reasonable.

[0145] In some embodiments, the loss value generated by the first sample image sequence may include at least one of a first loss value, a second loss value, or a third loss value. For example, the server may use any one of the first, second, or third loss values ​​as the loss value generated by the first sample image sequence. Alternatively, the loss value generated by the first sample image sequence may be the result of a weighted calculation of any two of the first, second, or third loss values. Or, the loss value generated by the first sample image sequence may be the result of a weighted calculation of the first, second, and third loss values. The loss value of the first sample image sequence is positively correlated with the first, second, and third loss values. For example, the formula for calculating the loss value of the first sample image sequence is: (4), of which, The weight of the first loss value, The weights for the second loss value, The weight of the third loss value. , , All are greater than 0. , , It can be configured as needed, for example... 0.5 0.05 It is 0.1. The value range is -1 to 1. For example, if the sequence length of the first sample image sequence is 4, then... .

[0146] It should be noted that the formula for calculating the loss value of the first sample image sequence can be used to calculate the loss value of any sample image sequence, such as calculating the loss value of the second sample image sequence.

[0147] In this embodiment, the difference between the first predicted matching degree of the sample image with the smallest matching degree in the first sample image sequence and the minimum matching degree is determined. The matching degree differences and the minimum matching degree are summed to obtain the summation result. Based on the summation result, the third loss value generated by the first sample image sequence is obtained. The third loss value is negatively correlated with the summation result. Thus, by training the image evaluation model, the first predicted matching degree of the sample image with the largest matching degree in the first sample image sequence can be maximized, thereby improving the accuracy of the image evaluation model.

[0148] In some embodiments, such as Figure 10 As shown, an image processing method is provided. This method can be executed by a terminal or a server, or by both a terminal and a server. Taking the application of this method to a server as an example, it includes the following steps 1002 to 1008. Wherein:

[0149] Step 1002: Obtain the text features obtained by encoding the text.

[0150] The text can be sent from the terminal to the server, obtained by the server from other devices, or generated by the server. Text features can be encoded by the server, pre-stored on the server, or obtained by the server from other devices.

[0151] Specifically, the server obtains the text, inputs it into a text encoder for encoding, and obtains the text features.

[0152] Step 1004: Obtain the image features obtained by encoding the generated image generated using text.

[0153] The generated image is based on the text. The generated image can be generated by the server or obtained by the server from other devices. Image features can be obtained by the server encoding the generated image or obtained by the server from other devices.

[0154] Specifically, the generated image can be generated using a text-based image generation model. The server inputs text into the text-based image generation model to obtain a generated image using the text.

[0155] In some embodiments, the terminal inputs the generated image into an image encoder for encoding to obtain image features.

[0156] Step 1006: Fuse the image features and text features to obtain the image-text fusion features corresponding to the generated image.

[0157] Specifically, when the image features and text features have the same dimension, such as being vectors of the same length, fusion can be achieved by multiplying, adding, or averaging the corresponding positions of the image features and text features to obtain the image-text fused features.

[0158] In some embodiments, the text features may include multiple sub-text features, each of which is a vector. The image features are also vectors. The length of the image features is the same as the length of the sub-text features. The server fuses the multiple sub-text features in the text features with the image features respectively to obtain multiple sub-image-text fusion features corresponding to the generated image. The server then performs statistical analysis, such as pooling, on these multiple sub-image-text fusion features corresponding to the generated image to obtain the image-text fusion features corresponding to the generated image.

[0159] Step 1008: Input the image-text fusion features into the trained image evaluation model to obtain the target matching degree between the generated image and the text. The image evaluation model is trained using the image evaluation model processing method.

[0160] Specifically, image evaluation models are used to predict the matching degree between text and images. For example... Figure 11 As shown, the server inputs text into a text encoder to obtain text features, then inputs the generated image from the text into an image encoder to obtain image features. A fusion module then fuses the image features with the text features to obtain the image-text fused features corresponding to the generated image. These fused features are then input into a trained image evaluation model to obtain the target matching degree. For example, the server normalizes the output of the image evaluation model to obtain the target matching degree. When multiple generated images are generated from text, the target matching degree between each of these generated images and the text can be predicted separately.

[0161] In some embodiments, before obtaining the text features obtained by encoding the text, the method further includes: generating multiple generated images based on the text in response to a text-based image generation request instructing the generation of an image based on the text; the method further includes: returning at least one target generated image among the multiple generated images based on the target matching degree of each of the multiple generated images with the text. Specifically, the terminal displays a text-based image interface, obtains the text input into the text-based image interface, and sends a text-based image request carrying the text to the terminal. The text is the text that the user wants to generate an image. In response to the text-based image request, the server inputs the text into a text-based image generation model and generates multiple generated images. It should be noted that there can be different text-based image generation models, and the application scenarios of different text-based image generation models can be different. Different text-based image generation models are used to generate images in different scenarios. The server can use a text-based image generation model consistent with the scenario of the text to obtain the generated image. For example, in a pet photo show application, the terminal displays a text-based image interface through the application and obtains text describing the pet through the text-based image interface. If the server determines that the text describes the pet, it can input the text into a text-based image generation model used to generate pet images to obtain the generated image. The server can input text into the text-to-image generation model multiple times to generate multiple generated images.

[0162] In some embodiments, for each generated image, the server uses an image evaluation model to obtain the target matching degree between the generated image and the text. At least one target generated image is selected from the plurality of generated images in descending order of target matching degree. The target generated image refers to the selected generated image, and the selected at least one target generated image is returned to the terminal. The terminal can display each received target generated image. For example, the server can sort the plurality of generated images in descending order of target matching degree to obtain a sequence of generated images, and return the first k generated images in the sequence to the terminal. k is greater than or equal to 1.

[0163] In this embodiment, based on the target matching degree of each of the multiple generated images with the text, at least one target generated image is returned from the multiple generated images. Thus, by using the target matching degree obtained by the image evaluation model, at least one target generated image can be returned from the multiple generated images. This allows the generated image with better image effect and in line with human aesthetics to be returned, thereby accurately providing the user with the generated image they need and improving the interaction efficiency in text-to-image scenarios.

[0164] In some embodiments, obtaining text features obtained by encoding text includes: encoding multiple texts generated using at least one keyword to obtain text features for each text; the method further includes: for each text, statistically analyzing the target matching degree between at least two generated images and the text, obtaining a matching degree statistical value corresponding to the text; the matching degree statistical value is used to reflect the ability of the text to generate images; and selecting at least one target text from the multiple texts based on the matching degree statistical values ​​corresponding to each text. The matching degree statistical value corresponding to the text reflects the ability of the text to generate images; the larger the matching degree statistical value, the stronger the ability of the text to generate images, i.e., the better the generated image and the more it meets the requirements. Statistical analysis of the target matching degree includes, but is not limited to, calculating the average value of the target matching degree.

[0165] In some embodiments, the terminal sends a text-based image request carrying at least one keyword to the server. In response to the text-based image request, the server inputs the at least one keyword into a large language model to generate multiple texts. Each of these multiple texts is a feasible sentence organized around the at least one keyword. For example, if a user wants to generate an image with the keywords "dog," "beach," "sun umbrella," "yacht," and "swimming," then "dog," "beach," "sun umbrella," "yacht," and "swimming" are all keywords. A large language model can generate multiple possible sentences, such as "A dog is sitting under a sun umbrella on the beach, many people are swimming on the nearby shore, and a yacht is slowly sailing in the distance," or "A dog is swimming at the beach, there are many sun umbrellas on the beach, and a yacht is smoking in the distance." Large language models include, but are not limited to, question-and-answer type large language models. After the server generates multiple texts, it inputs each text into a text-to-image generation model to generate multiple generated images. For each generated image, the target matching degree between the generated image and the text is determined. The generated images are then sorted in descending order of target matching degree to obtain a sequence of generated images. The target matching degree of the first k1 generated images in this sequence is statistically analyzed to obtain the corresponding matching degree statistics for each text. Alternatively, the server can statistically analyze the target matching degree of each generated image against the text to obtain the corresponding matching degree statistics for each text.

[0166] In some embodiments, the target matching degree is a numerical value obtained after normalizing the output of the image evaluation model. The output is a numerical value, which can be understood as a score; the higher the score, the greater the target matching degree obtained after normalization. The server can sort the multiple generated images in descending order of score to obtain a sequence of generated images. The scores of the first k1 generated images in the sequence before target matching degree normalization with the text are statistically analyzed to obtain the score statistics for the text. The score statistics are used as an indicator to measure the text generation capability.

[0167] In some embodiments, the server can sort the multiple texts in descending order of their matching scores to obtain a text sequence, and select at least one target text from the text sequence. The target text is the selected text. For example, the first k2 texts can be selected from the text sequence to obtain k2 target texts. The server can return the at least one target text to the terminal, and for each target text, return at least one generated image generated using that target text to the terminal. For example, the server can sort the multiple generated images of the target text in descending order of their target matching scores to obtain a sequence of generated images corresponding to the target text, and return the first k3 generated images from the sequence to the terminal. k1~k3 are greater than or equal to 1. The terminal can display the received generated images and text. Figure 12 As shown, at least one keyword is input into the large language model. The large language model generates multiple sentences based on the keyword. A text-to-image generation model is used to generate multiple images for each sentence. An image evaluation model is used to score and sort each image to obtain the image sequence corresponding to each sentence. For example, the multiple images are sorted from largest to smallest based on the score. Then, the average score of the top k1 images in the image sequence is determined. The sentences are arranged according to the average score to obtain the sentence sequence. The top k2 sentences and the top k3 images in the image sequence corresponding to each of the top k2 sentences are returned.

[0168] In this embodiment, multiple texts are generated based on at least one keyword, and the ability of each text to generate an image is evaluated. Texts with strong generation capabilities and images with good image effects are returned, thereby automatically providing high-quality text and images and improving the interaction efficiency in text-to-image scenarios.

[0169] In the above image processing method, the image evaluation model is trained using the aforementioned image evaluation model processing method. In the fixed-length training phase, multiple first sample image sequences of the same length are used to train the image evaluation model. The sample images are sorted according to their matching degree. For each sample image in each first sample image sequence, a first predicted matching degree with the first sample text is predicted. Based on the first predicted matching degree and the sorting order of the sample images, a fixed-length sequence loss value is determined. The image evaluation model is trained using this fixed-length sequence loss value. Thus, through training, the image evaluation model learns the ability to predict, i.e., evaluate, the matching degree between the sample image and the first sample text. In the variable-length training phase... In the long training phase, multiple second sample image sequences of variable length are used to train the image evaluation model. For each sample image in each second sample image sequence, a second predicted matching degree with the second sample text is predicted. Based on the second predicted matching degree and the sorting order of the sample images in the multiple second sample image sequences, a variable-length sequence loss value is determined. This variable-length sequence loss value is used to train the image evaluation model, allowing for further adjustments. This enables the image evaluation model to learn the ability to predict, and evaluate, the second predicted matching degree between sample images and second sample text using second sample image sequences of different lengths, thereby further improving the accuracy of image-text matching degree evaluation. Therefore, the image evaluation model trained using the above image evaluation model processing method can accurately determine the target matching degree between the generated image and the text, improving the accuracy of image-text matching degree evaluation.

[0170] In some embodiments, such as Figure 13 As shown, an image processing method is provided. This method can be executed by a terminal or a server, or by both a terminal and a server. Taking the application of this method to a server as an example, it includes the following steps 1302 to 1320. Wherein:

[0171] Step 1302: Obtain multiple different global sequences by sorting multiple sample images according to the image-text matching degree determined by each person.

[0172] Among them, the sample images are images generated using the same sample text, and the image-text matching degree is the degree of matching between each sample image and the sample text.

[0173] Step 1304: Based on multiple different global sequences, determine the common sequence that meets the preset conditions.

[0174] The preset condition is that the relative sorting order of each sample image in the global sequence is consistent, and adding any sample image will not satisfy the requirement of maintaining the relative sorting order in the global sequence. There can be one or more common sequences that satisfy the preset condition.

[0175] Step 1306: Determine multiple specified sequence lengths. For each specified sequence length, obtain at least one subsequence with a specified sequence length from each common sequence with a sequence length greater than or equal to the specified sequence length, and form a set of sample image sequences corresponding to the specified sequence length.

[0176] Among them, the sequence lengths of sample image sequences in different sample image sequence sets are different, while the sequence lengths of sample image sequences in the same sample image sequence set are the same.

[0177] Step 1308: In the fixed-length training phase, determine the specified sequence length corresponding to the fixed-length training phase, and obtain some or all of the sample image sequences from the sample image sequence set corresponding to the specified sequence length of the fixed-length training phase to obtain multiple sample image sequences of fixed length. The specified sequence length corresponding to the fixed-length training phase is less than or equal to the length threshold.

[0178] The length threshold is set as needed, and its value should be relatively small, such as 2 or 3. If the length threshold is the smallest specified sequence length among the multiple specified sequence lengths, then the specified sequence length corresponding to the fixed-length training phase is the smallest specified sequence length among the multiple specified sequence lengths. If there is at least one specified sequence length among the multiple specified sequence lengths that is less than or equal to the length threshold, the specified sequence length corresponding to the fixed-length training phase can be any of the at least one specified sequence length that is less than or equal to the length threshold.

[0179] Step 1310: For each sample image sequence in a fixed-length sequence of multiple sample images, predict the first predicted matching degree between each sample image in the sample image sequence and the sample text using the image evaluation model. Based on the first predicted matching degree and the sorting order of the sample images in the fixed-length sequence of multiple sample images, determine the fixed-length sequence loss value for the sample text, and train the image evaluation model using the fixed-length sequence loss value for the sample text.

[0180] Step 1312: In the variable-length training phase, at least two sample image sequence sets are determined from the sample image sequence sets corresponding to multiple specified sequence lengths respectively. From the determined at least two sample image sequence sets, some or all sample image sequences are obtained respectively to obtain multiple sample image sequences of variable length.

[0181] Step 1314: For each sample image sequence in multiple sample image sequences of variable length, predict the second predicted matching degree between each sample image in the sample image sequence and the sample text through the image evaluation model. Based on the second predicted matching degree and the sorting order of the sample images in multiple sample image sequences of fixed length, determine the variable length sequence loss value for the sample text. Use the variable length sequence loss value for the sample text to train the image evaluation model to obtain the trained image evaluation model.

[0182] The training sample texts can be multiple, and can go through one or more fixed-length training stages or one or more variable-length training stages.

[0183] Step 1316: Obtain the text features obtained by encoding the text, and obtain the image features obtained by encoding the generated image generated from the text.

[0184] Step 1318: Fuse the image features and text features to obtain the image-text fusion features corresponding to the generated image.

[0185] Step 1320: Input the image-text fusion features into the trained image evaluation model to obtain the target matching degree between the generated image and the text.

[0186] In this embodiment, since it is easier to obtain sample image sequences with shorter sequences than those with longer sequences, and the specified sequence length for the fixed-length training phase is less than or equal to the length threshold, the number of sample image sequences participating in the training phase is relatively large. Therefore, the image evaluation model can learn the basic sample space distribution through the fixed-length training phase. In contrast, the variable-length training phase allows for the adjustment of the image evaluation model using multiple sample image sequences of varying lengths. This enables the image evaluation model to be applicable to sample image sequences of different lengths, correctly evaluating the matching degree for each sample image in the sequence and improving the accuracy of the matching degree evaluation. Furthermore, since the relative order of the sample images in the sample image sequence is the order jointly determined by multiple people, the relative order of the sample images in the sample image sequence conforms to human cognition. Therefore, training the image evaluation model with the sample image sequence allows the image evaluation model to learn knowledge that conforms to human cognition. This ensures that when the image evaluation model is used to evaluate the generated image, the evaluation result conforms to human cognition and aesthetic standards. Thus, inputting the image-text fusion features corresponding to the generated image into the trained image evaluation model to obtain the target matching degree between the generated image and the text can make the evaluated target matching degree conform to human cognition and aesthetic standards, thereby improving the evaluation accuracy of the generated image.

[0187] The image processing method provided in this application can be applied to any application scenario that requires evaluating the matching degree between text and generated images. For example, the image processing method provided in this application can be used for advertising design. The server can obtain advertising copy, generate multiple generated images using the advertising copy, encode text features of the advertising copy, encode image features of the generated images, fuse image features with text features to obtain image-text fusion features corresponding to the generated images, input the image-text fusion features into a trained image evaluation model to obtain the target matching degree between the generated images and the advertising copy, and return at least one target generated image with a higher target matching degree among the multiple generated images to the terminal based on the target matching degree of each of the multiple generated images with the advertising copy. Thus, the terminal can use the target generated image as material for advertising design.

[0188] For example, the image processing method provided in this application can be used for game design. For instance, a server can obtain descriptive text describing a game scene or a game character, generate multiple images based on the descriptive text, encode text features from the descriptive text, encode image features from the generated images, fuse image features with text features to obtain image-text fusion features corresponding to the generated image, input these image-text fusion features into a trained image evaluation model to obtain the target matching degree between the generated image and the descriptive text, and return at least one target generated image with a higher target matching degree among the multiple generated images to the terminal. When the descriptive text describes a game scene, the target generated image is an image of that game scene, allowing the terminal to use the target generated image as material for designing the game scene. When the descriptive text describes a game character, the target generated image is an image of that game character, allowing the terminal to use the target generated image as material for designing the game character.

[0189] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0190] Based on the same inventive concept, this application also provides an image evaluation model processing apparatus for implementing the image evaluation model processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more image evaluation model processing apparatus embodiments provided below can be found in the limitations of the image evaluation model processing method described above, and will not be repeated here.

[0191] In some embodiments, such as Figure 14 As shown, an image evaluation model processing device is provided, including: a first sequence acquisition module 1402, a first training module 1404, a second sequence acquisition module 1406, and a second training module 1408, wherein:

[0192] The first sequence acquisition module 1402 is used to acquire multiple first sample image sequences of fixed length during the fixed-length training phase. Each first sample image sequence includes sample images generated using first sample text, and the included sample images are sorted according to their matching degree with the first sample text. The multiple first sample image sequences have the same sequence length, and the sequence length refers to the number of included sample images.

[0193] The first training module 1404 is used to evaluate the image by predicting the first predicted matching degree with the first sample text for each sample image in each first sample image sequence, determining the fixed-length sequence loss value based on the first predicted matching degree and the sorting order of the sample images in multiple first sample image sequences, and training the image evaluation model using the fixed-length sequence loss value.

[0194] The second sequence acquisition module 1406 is used to acquire multiple second sample image sequences of variable length during the variable length training phase. Each second sample image sequence includes sample images generated using second sample text, and the included sample images are sorted according to their degree of matching with the second sample text. Variable length means that at least two second sample image sequences have different sequence lengths.

[0195] The second training module 1408 is used to evaluate the image by predicting the second predicted matching degree with the second sample text for each sample image in each second sample image sequence, determining the variable-length sequence loss value based on the second predicted matching degree and the sorting order of the sample images in multiple second sample image sequences, and training the image evaluation model using the variable-length sequence loss value.

[0196] In some embodiments, the image evaluation model processing apparatus further includes a set acquisition module, which is used to acquire a set of multiple sample image sequences determined based on multiple sample images corresponding to the first sample text. The sample images are images generated using the first sample text. The sequence lengths of the sample image sequences in different sample image sequence sets are different, while the sequence lengths of the sample image sequences in the same sample image sequence set are the same. The sample image sequence includes at least a portion of the sample images among the multiple sample images, and the sample images in the sample image sequence are sorted according to their degree of matching with the first sample text. The first sequence acquisition module 1402 is further used to acquire a portion or all of the sample image sequences from one of the multiple sample image sequence sets to obtain a set of multiple first sample image sequences of a fixed length.

[0197] In some embodiments, the first sample text and the second sample text are the same text. The second sequence acquisition module 1406 is further configured to determine at least two sample image sequence sets from a plurality of sample image sequence sets; and to acquire part or all of the sample image sequences from each of the at least two sample image sequence sets to obtain a plurality of second sample image sequences of variable length.

[0198] In some embodiments, the set acquisition module is further configured to acquire multiple different global sequences obtained by sorting multiple sample images according to the image-text matching degree determined by each of the multiple objects, wherein the image-text matching degree is the degree of matching between each sample image and the first sample text; and determine a set of sample image sequences corresponding to multiple specified sequence lengths based on the multiple different global sequences, wherein the sequence length of the sample image sequence in the set of sample image sequences is the corresponding specified sequence length, and the relative sorting order of the sample images in the sample image sequence is consistent with the relative sorting order in each global sequence.

[0199] In some embodiments, the set acquisition module is further configured to determine a common sequence that meets a preset condition based on multiple different global sequences. The preset condition is that the relative sorting order of each sample image contained therein is consistent in each global sequence, and that the relative sorting order in each global sequence is no longer consistent after adding any sample image. For each specified sequence length among multiple specified sequence lengths, at least one subsequence with a specified sequence length is obtained from each common sequence with a sequence length greater than or equal to the specified sequence length, and a sample image sequence set corresponding to the specified sequence length is formed.

[0200] In some embodiments, the first training module 1404 is further configured to, for each first sample image sequence, determine the loss value generated by the first sample image sequence based on the sorting order of the sample images in the first sample image sequence and the first prediction matching degree of the sample images; and perform comprehensive calculation on the loss values ​​generated by each first sample image sequence to obtain a fixed-length sequence loss value.

[0201] In some embodiments, the first training module 1404 is further configured to determine at least one pair of adjacent sample images in the first sample image sequence according to the sorting order of the sample images in the first sample image sequence, wherein the pair of adjacent sample images includes two adjacent sample images in the first sample image sequence; for each pair of adjacent sample images, determine the difference between the first predicted matching degrees of the sample images in the pair of adjacent sample images to obtain the matching degree difference corresponding to the pair of adjacent sample images; and obtain the loss value generated by the first sample image sequence based on the matching degree difference corresponding to each pair of adjacent sample images.

[0202] In some embodiments, the loss value generated by the first sample image sequence includes a first loss value. The first training module 1404 is further configured to normalize the matching degree difference corresponding to each adjacent sample image pair to obtain the normalized matching degree difference corresponding to the adjacent sample image pair; and obtain the first loss value generated by the first sample image sequence based on the normalized matching degree difference corresponding to each adjacent sample image pair.

[0203] In some embodiments, the loss value generated by the first sample image sequence includes a second loss value. The first training module 1404 is further configured to perform a logarithmic transformation on the matching degree difference corresponding to each adjacent sample image pair to obtain the transformed matching degree difference corresponding to the adjacent sample image pair; use the matching degree difference corresponding to each adjacent sample image pair as the weight of the transformed matching degree difference corresponding to the corresponding transformed matching degree difference, perform weighted calculation on the transformed matching degree difference corresponding to each adjacent sample image pair, and obtain the second loss value generated by the first sample image sequence based on the weighted result.

[0204] In some embodiments, the loss value generated by the first sample image sequence includes a third loss value. The first training module 1404 is further configured to determine the matching degree difference between the first predicted matching degree of the sample image with the smallest matching degree in the first sample image sequence and the minimum matching degree; sum the matching degree differences and the minimum matching degree to obtain a summation result; and obtain the third loss value generated by the first sample image sequence based on the summation result. The third loss value is negatively correlated with the summation result.

[0205] Each module in the aforementioned image evaluation model processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0206] Based on the same inventive concept, this application also provides an image processing apparatus for implementing the image processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more image processing apparatus embodiments provided below can be found in the limitations of the image processing method described above, and will not be repeated here.

[0207] In some embodiments, such as Figure 15 As shown, an image processing apparatus is provided, including: a text feature acquisition module 1502, an image feature acquisition module 1504, a feature fusion module 1506, and a matching degree acquisition module 1508, wherein:

[0208] The text feature acquisition module 1502 is used to acquire text features obtained by encoding the text.

[0209] The image feature acquisition module 1504 is used to acquire image features obtained by encoding the generated image generated from text.

[0210] The feature fusion module 1506 is used to fuse image features and text features to obtain the image-text fusion features corresponding to the generated image.

[0211] The matching degree module 1508 is used to input the image-text fusion features into the trained image evaluation model to obtain the target matching degree between the generated image and the text. The image evaluation model is trained by the above image evaluation model processing method.

[0212] Each module in the aforementioned image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0213] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 16As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data involved in the image evaluation model processing method or image processing method. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an image evaluation model processing method or image processing method.

[0214] In some embodiments, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 17 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an image evaluation model processing method or an image processing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0215] Those skilled in the art will understand that Figure 16 and Figure 17The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0216] In some embodiments, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the image evaluation model processing method described above.

[0217] In some embodiments, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the image evaluation model processing method described above.

[0218] In some embodiments, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the image evaluation model processing method described above.

[0219] In some embodiments, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the image processing method described above.

[0220] In some embodiments, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the image processing method described above.

[0221] In some embodiments, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the image processing method described above.

[0222] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0223] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0224] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0225] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An image evaluation model processing method, characterized in that, The method includes: During the fixed-length training phase, multiple first sample image sequences of fixed length are obtained. Each first sample image sequence includes sample images generated using first sample text, and the included sample images are sorted according to their matching degree with the first sample text. The multiple first sample image sequences have the same sequence length, and the sequence length refers to the number of included sample images. The image evaluation model predicts a first predicted matching degree with the first sample text for each sample image in each first sample image sequence. Based on the first predicted matching degree and the sorting order of the sample images in the plurality of first sample image sequences, a fixed-length sequence loss value is determined, and the image evaluation model is trained using the fixed-length sequence loss value. During the variable-length training phase, multiple second sample image sequences of variable length are obtained. Each second sample image sequence includes sample images generated using second sample text, and the included sample images are sorted according to their degree of matching with the second sample text. Variable length means that at least two second sample image sequences have different sequence lengths. The image evaluation model predicts a second predicted matching degree with the second sample text for each sample image in each second sample image sequence. Based on the second predicted matching degree and the sorting order of the sample images in the plurality of second sample image sequences, a variable-length sequence loss value is determined, and the image evaluation model is trained using the variable-length sequence loss value.

2. The method according to claim 1, characterized in that, The method further includes: Obtain a set of multiple sample image sequences determined based on multiple sample images corresponding to the first sample text. The sample images are images generated using the first sample text. The sequence lengths of the sample image sequences in different sample image sequence sets are different, while the sequence lengths of the sample image sequences in the same sample image sequence set are the same. The sample image sequence includes at least a portion of the sample images from the multiple sample images, and the sample images in the sample image sequence are sorted according to their degree of matching with the first sample text. The step of obtaining a fixed-length sequence of multiple first sample images includes: From one of the multiple sample image sequence sets, obtain some or all of the sample image sequences to obtain multiple first sample image sequences of fixed length.

3. The method according to claim 2, characterized in that, The first sample text and the second sample text are the same text. The step of obtaining multiple second sample image sequences of variable length includes: Determine at least two sample image sequence sets from the plurality of sample image sequence sets; From each of the at least two sample image sequence sets, a portion or all of the sample image sequences are obtained to obtain multiple second sample image sequences of variable length.

4. The method according to claim 2, characterized in that, The step of obtaining a set of multiple sample image sequences determined based on multiple sample images corresponding to the first sample text includes: Multiple different global sequences are obtained by sorting the multiple sample images according to the image-text matching degree determined by each of the multiple objects, wherein the image-text matching degree is the degree of matching between each sample image and the first sample text. Based on the multiple different global sequences, a set of sample image sequences corresponding to multiple specified sequence lengths is determined. The sequence length of the sample image sequences in the set of sample image sequences is the corresponding specified sequence length. The relative sorting order of the sample images in the sample image sequences is consistent with the relative sorting order in each of the global sequences.

5. The method according to claim 4, characterized in that, The step of determining a set of sample image sequences corresponding to multiple specified sequence lengths based on the multiple different global sequences includes: Based on the multiple different global sequences, a common sequence that meets a preset condition is determined. The preset condition is that the relative sorting order of each sample image in each global sequence is consistent, and the relative sorting order in each global sequence is no longer consistent after adding any sample image. For each specified sequence length among the plurality of specified sequence lengths, at least one subsequence with a sequence length of the specified sequence length is obtained from each common sequence with a sequence length greater than or equal to the specified sequence length, forming a sample image sequence set corresponding to the specified sequence length.

6. The method according to any one of claims 1 to 5, characterized in that, The step of determining the fixed-length sequence loss value based on the first predicted matching degree and the sorting order of the sample images in the plurality of first sample image sequences includes: For each of the first sample image sequences, the loss value generated by the first sample image sequence is determined based on the sorting order of the sample images in the first sample image sequence and the first prediction matching degree of the sample images. The loss values ​​generated by each of the first sample image sequences are calculated together to obtain the fixed-length sequence loss value.

7. The method according to claim 6, characterized in that, The step of determining the loss value generated by the first sample image sequence based on the sorting order of the sample images in the first sample image sequence and the first predicted matching degree of the sample images includes: Based on the sorting order of the sample images in the first sample image sequence, at least one pair of adjacent sample images in the first sample image sequence is determined, and the pair of adjacent sample images includes two adjacent sample images in the first sample image sequence. For each pair of adjacent sample images, the difference between the first predicted matching degrees of the sample images in the pair of adjacent sample images is determined to obtain the matching degree difference corresponding to the pair of adjacent sample images. The loss value generated by the first sample image sequence is obtained based on the matching degree difference corresponding to each of the adjacent sample image pairs.

8. The method according to claim 7, characterized in that, The loss value generated by the first sample image sequence includes a first loss value. The step of obtaining the loss value generated by the first sample image sequence based on the matching degree difference corresponding to each of the adjacent sample image pairs includes: For each pair of adjacent sample images, the matching degree difference corresponding to the pair of adjacent sample images is normalized to obtain the normalized matching degree difference corresponding to the pair of adjacent sample images. Based on the normalized matching degree difference corresponding to each of the adjacent sample image pairs, the first loss value generated by the first sample image sequence is obtained.

9. The method according to claim 7, characterized in that, The loss value generated by the first sample image sequence includes a second loss value. The step of obtaining the loss value generated by the first sample image sequence based on the matching degree difference corresponding to each of the adjacent sample image pairs includes: For each pair of adjacent sample images, a logarithmic transformation is performed on the matching degree difference corresponding to the pair of adjacent sample images to obtain the transformed matching degree difference corresponding to the pair of adjacent sample images. The matching degree difference corresponding to each of the adjacent sample image pairs is used as the weight of the corresponding transformed matching degree difference. The transformed matching degree difference corresponding to each of the adjacent sample image pairs is weighted and calculated. The second loss value generated by the first sample image sequence is obtained based on the weighted result.

10. The method according to claim 7, characterized in that, The loss value generated by the first sample image sequence includes a third loss value. The step of obtaining the loss value generated by the first sample image sequence based on the matching degree difference corresponding to each of the adjacent sample image pairs includes: Determine the difference in matching degree between the first predicted matching degree and the minimum matching degree of the sample image with the lowest matching degree in the first sample image sequence; The summation result is obtained by summing the differences in matching degree and the minimum matching degree. A third loss value is obtained based on the summation result, and the third loss value is negatively correlated with the summation result.

11. An image processing method, characterized in that, The method includes: Obtain the text features obtained by encoding the text; Obtain the image features encoded from the generated image generated using the text; The image features and the text features are fused to obtain the image-text fusion features corresponding to the generated image; The image-text fusion features are input into a trained image evaluation model to obtain the target matching degree between the generated image and the text. The image evaluation model is trained by any one of the methods of claims 1 to 10.

12. An image evaluation model processing device, characterized in that, The device includes: The first sequence acquisition module is used to acquire multiple first sample image sequences of fixed length during the fixed-length training phase. Each first sample image sequence includes a sample image generated using first sample text, and the included sample images are sorted according to their matching degree with the first sample text. The multiple first sample image sequences have the same sequence length, and the sequence length refers to the number of included sample images. The first training module is used to predict a first predicted matching degree with the first sample text for each sample image in each first sample image sequence through an image evaluation model, determine a fixed-length sequence loss value based on the first predicted matching degree and the sorting order of the sample images in the plurality of first sample image sequences, and train the image evaluation model using the fixed-length sequence loss value. The second sequence acquisition module is used to acquire multiple second sample image sequences of variable length during the variable length training phase. Each second sample image sequence includes a sample image generated using the second sample text, and the included sample images are sorted according to their degree of matching with the second sample text. Variable length means that at least two second sample image sequences have different sequence lengths. The second training module is used to predict a second predicted matching degree with the second sample text for each sample image in each second sample image sequence using the image evaluation model, determine a variable-length sequence loss value based on the second predicted matching degree and the sorting order of the sample images in the plurality of second sample image sequences, and train the image evaluation model using the variable-length sequence loss value.

13. An image processing apparatus, characterized in that, The device includes: The text feature acquisition module is used to acquire text features obtained by encoding the text. The image feature acquisition module is used to acquire image features encoded from the generated image generated using the text; The feature fusion module is used to fuse the image features and the text features to obtain the image-text fusion features corresponding to the generated image; The matching degree acquisition module is used to input the image-text fusion features into a trained image evaluation model to obtain the target matching degree between the generated image and the text. The image evaluation model is trained by any one of the methods of claims 1 to 10.

14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.

16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Target identification method, model training method and electronic equipment

    CN116168380A

  • Figure graph model training method and text graph method

    CN116935169A