Training of large-scale cultural graph models and cultural graph methods, devices, equipment, and media
Through multi-stage training and screening and fine-tuning of image-text data pairs, the problems of low training efficiency and poor results of large Chinese text-to-image models were solved, the model's prediction accuracy and image quality were improved, and the user experience was enhanced.
Patent Information
- Application Number
- CN202411505398.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-10-25
AI Technical Summary
The existing technology has problems of low training efficiency and poor results when training large Chinese text-to-graph models, especially when directly training based on English models or omitting the alignment process.
Using multiple image-text data pairs, the text-image large model is trained in multiple stages in sequence. The model parameters are screened through evaluation indicators and fine-tuned in the last training stage. The final adjustment is made using high-quality image-text data pairs with high aesthetic scores to improve the accuracy and aesthetics of the model parameters.
It improves the training effect and output image quality of the large-scale cultural image model, improves the user experience in cultural image scenarios, and enhances the prediction accuracy and aesthetics of the model.
Smart Images

Figure CN119516016B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of AI (Artificial Intelligence), specifically to technical fields such as deep learning, large models, and computer vision. It can be used in application fields such as generative search, intelligent document editing, intelligent assistants, and intelligent e-commerce, and in particular to the training of large-scale text-graph models and text-graph methods, devices, equipment, and media. Background Art
[0002] AI painting, one of the key applications of artificial intelligence, has achieved significant breakthroughs in recent years. It can automatically generate images of various styles based on user input or prompts, providing powerful auxiliary tools for artists, designers, and creators, and bringing new possibilities to the field of digital creativity. Therefore, when users search for images, they are no longer limited to existing image libraries. Instead, they can create new image content and styles by generating images based on their needs and creativity. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device and medium for training a large-scale culture graph model.
[0004] According to a first aspect of the present disclosure, a method for training a large-scale cultural graph model is provided, comprising:
[0005] Using multiple image-text data pairs, sequentially perform multiple training phases on the Wensheng graph large model; wherein the model parameters of the Wensheng graph large model to be trained in the i-th training phase are obtained by screening based on the evaluation indicators of multiple sets of model parameters of the Wensheng graph large model trained in the i-1-th training phase, where i is a positive integer greater than 1;
[0006] Determine the model parameters to be fine-tuned from the multiple sets of model parameters of the large Wensheng graph model obtained from the last training phase;
[0007] determining a first image-text data pair from the plurality of image-text data pairs based on first quality scores and aesthetic scores of sample images in the plurality of image-text data pairs;
[0008] The first image-text data pair is used to fine-tune the parameters of the model to be fine-tuned to obtain target model parameters corresponding to the large text-image model.
[0009] According to a second aspect of the present disclosure, a method for generating a cultural image is provided, comprising:
[0010] Get input text;
[0011] Calling a large text-image model with built-in target model parameters to process the input text to obtain a target image;
[0012] The large model of the cultural graph is obtained by training using the method proposed in the embodiment of the first aspect.
[0013] According to a third aspect of the present disclosure, a training device for a large-scale Wensheng graph model is provided, comprising:
[0014] an execution module, configured to sequentially execute multiple training phases on the Wensheng graph large model using multiple image-text data pairs; wherein the model parameters of the Wensheng graph large model to be trained in the i-th training phase are obtained by screening based on the evaluation indicators of multiple sets of model parameters of the Wensheng graph large model trained in the i-1-th training phase, where i is a positive integer greater than 1;
[0015] A screening module is used to determine the model parameters to be fine-tuned from the multiple sets of model parameters of the Wenshengtu large model trained in the last training phase;
[0016] a determining module, configured to determine a first image-text data pair from the plurality of image-text data pairs based on a first quality score and an aesthetic score of a sample image in the plurality of image-text data pairs;
[0017] The fine-tuning module is used to fine-tune the parameters of the model to be fine-tuned using the first image-text data pair to obtain target model parameters corresponding to the large text-image model.
[0018] According to a fourth aspect of the present disclosure, there is provided a cultural image device, comprising:
[0019] Acquisition module, used to obtain input text;
[0020] A calling module is used to call a large text-image model with built-in target model parameters to process the input text to obtain a target image;
[0021] Among them, the large model of the cultural graph is obtained by training using the device proposed in the embodiment of the third aspect.
[0022] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:
[0023] at least one processor; and
[0024] a memory communicatively connected to the at least one processor; wherein,
[0025] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method of the large model of the Vincent graph proposed in the first aspect of the present disclosure, or execute the Vincent graph method proposed in the second aspect of the present disclosure.
[0026] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium of computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method for training the Vincent graph large model proposed in the first aspect of the present disclosure, or to execute the Vincent graph method proposed in the second aspect of the present disclosure.
[0027] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the method for training the large Vincent graph model proposed in the first aspect of the present disclosure, or, when executed, implements the Vincent graph method proposed in the second aspect of the present disclosure.
[0028] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0030] Figure 1 A flowchart of a method for training a large-scale cultural graph model provided in the first embodiment of the present disclosure;
[0031] Figure 2 A flowchart of a method for training a large-scale cultural graph model provided in the second embodiment of the present disclosure;
[0032] Figure 3 A flowchart of a method for training a large-scale Wensheng graph model provided in the third embodiment of the present disclosure;
[0033] Figure 4 A flowchart of a method for training a large-scale Wensheng graph model provided in the fourth embodiment of the present disclosure;
[0034] Figure 5 A flowchart of a method for training a large-scale Wensheng graph model provided in the fifth embodiment of the present disclosure;
[0035] Figure 6 A flowchart of a method for training a large-scale Wensheng graph model provided in the sixth embodiment of the present disclosure;
[0036] Figure 7 A flowchart of a method for training a large-scale Wensheng graph model provided in the seventh embodiment of the present disclosure;
[0037] Figure 8 A schematic diagram of the model training principle provided by an embodiment of the present disclosure;
[0038] Figure 9 A schematic flow chart of the Wensheng diagram method provided in the eighth embodiment of the present disclosure;
[0039] Figure 10 A schematic diagram of the structure of a training device for a large-scale Wensheng graph model provided in Example 9 of the present disclosure;
[0040] Figure 11 This is a schematic diagram of the structure of the Wenshengtu device provided in the tenth embodiment of the present disclosure;
[0041] Figure 12 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0042] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0043] Currently, companies can combine the development of AIGC (Artificial Intelligence Generated Content) technology and users' image search needs to develop native AI painting products. Through these AI painting products, creative images of different scenes and styles can be intelligently generated, providing users with new image "retrieval" methods and realizing a generational change in the field of image search.
[0044] Among them, the underlying technology that AI painting products rely on is the Chinese cultural image model. Therefore, how to train a good Chinese cultural image model is a very important core task.
[0045] Among the related technologies, it is based on the open source Chinese / English text-based graph architecture, directly training the company's own business data to obtain a large Chinese text-based graph model.
[0046] For example, after starting training on the English-based text-graph model, the Chinese-based text-graph model is trained. However, the alignment process of this training method is relatively slow.
[0047] Alternatively, training can be started directly based on the Chinese text-based graph model. Although this training method omits the alignment process, it is directly trained on the company's own training data, which has low training efficiency and poor training results.
[0048] Therefore, in response to at least one of the above-mentioned problems, the present disclosure proposes a training method, device, equipment and medium for a large-scale cultural graph model.
[0049] The following describes the training of the large-scale model of the culture graph and the culture graph method, device, equipment and medium of the embodiments of the present disclosure with reference to the accompanying drawings.
[0050] Large models are machine learning models with large parameters and complex computational structures. They are typically built from deep neural networks and contain billions or even hundreds of billions of parameters. Large models are designed to improve their expressiveness and predictive performance, enabling them to handle more complex tasks and data. Large models are widely used in various fields, including natural language processing, computer vision, speech recognition, and recommendation systems.
[0051] Figure 1 Schematic diagram of the flow of the training method of the large-scale cultural graph model provided in the first embodiment of the present disclosure.
[0052] The embodiment of the present disclosure uses the example of configuring the training method of the large-scale cultural graph model in a training device for the large-scale cultural graph model. The training device can be applied to any electronic device so that the electronic device can perform the training function of the large-scale cultural graph model.
[0053] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a mobile phone, tablet computer, personal digital assistant, wearable device, etc., which are hardware devices with various operating systems, touch screens and / or display screens.
[0054] like Figure 1 As shown, the training method of the large text-image model may include the following steps S101 to S104:
[0055] In step S101, a plurality of image-text data pairs are used to sequentially perform a plurality of training stages on the large model of the culture graph; wherein the model parameters of the large model of the culture graph to be trained in the i-th training stage are obtained by screening the evaluation indicators of the multiple sets of model parameters of the large model of the culture graph trained in the i-1-th training stage.
[0056] Wherein, i is a positive integer greater than 1.
[0057] Among them, each image-text data pair includes a sample text and a corresponding sample image. It should be noted that the present disclosure does not limit the method of obtaining the image-text data pair. For example, the image-text data pair can be obtained from an existing training set, or the image-text data pair can be obtained from a self-owned training set, or the image-text data pair can be collected online, such as using web crawler technology to collect image-text data pairs online, or the image-text data pairs can be manually generated, and so on.
[0058] Among them, multiple training stages are pre-configured training stages, including but not limited to: semantic tone alignment stage, resolution improvement stage, multi-size training stage, etc.
[0059] The evaluation metrics are used to indicate the accuracy or quality of the model parameters of the large Wensheng graph model. These metrics can include both positive and negative metrics, which are not limited in this disclosure. For example, the evaluation metrics can include accuracy, precision, loss (or loss value), and the like.
[0060] In the embodiment of the present disclosure, multiple image-text data pairs can be used to sequentially perform multiple training stages on the text-image model.
[0061] As an example, for the first of multiple training phases, multiple image-text data pairs can be used to perform this first training phase on the large-scale Wensheng-Graph model, thereby obtaining multiple sets of model parameters corresponding to the large-scale Wensheng-Graph model trained in the first training phase. That is, during the first training phase, the model parameters of the large-scale Wensheng-Graph model can be adjusted multiple times to obtain multiple sets of adjusted model parameters.
[0062] For example, in the first training stage, after each x steps (where a step refers to a basic iterative unit in the training process, usually corresponding to a complete process of forward propagation and backward propagation) are trained, the model parameters at this time can be used as a set of model parameters obtained by training in the first training stage.
[0063] In the present disclosure, quality evaluation can also be performed on multiple sets of model parameters of the Vincent graph large model trained in the first training stage to obtain evaluation indicators of the multiple sets of model parameters of the Vincent graph large model trained in the first training stage. Based on the evaluation indicators of the multiple sets of model parameters of the Vincent graph large model trained in the first training stage, the model parameters of the Vincent graph large model to be trained in the second training stage are determined from the multiple sets of model parameters of the Vincent graph large model trained in the first training stage.
[0064] For example, the model parameters with the best evaluation index can be used as the model parameters of the large Wensheng graph model to be trained in the second training stage.
[0065] For the i-th (i is a positive integer greater than or equal to 2) training stage among multiple training stages, multiple image-text data pairs can be used to execute the i-th training stage on the large text-image model to be trained in the i-th training stage to obtain multiple sets of model parameters corresponding to the large text-image model trained in the i-th training stage.
[0066] In the present disclosure, quality evaluation can also be performed on multiple sets of model parameters of the Vincent graph large model trained in the i-th training stage to obtain evaluation indicators of the multiple sets of model parameters of the multiple Vincent graph large models trained in the i-th training stage. Based on the evaluation indicators of the multiple sets of model parameters of the Vincent graph large model trained in the i-th training stage, the model parameters of the Vincent graph large model to be trained in the i+1-th training stage are determined from the multiple sets of model parameters of the Vincent graph large model trained in the i-th training stage.
[0067] Step S102 : determining the model parameters to be fine-tuned from the multiple sets of model parameters of the large Wensheng graph model obtained through training in the last training phase.
[0068] In an embodiment of the present disclosure, a set of model parameters to be fine-tuned (referred to as model parameters to be fine-tuned in the present disclosure) may be determined from multiple sets of model parameters of the large Wensheng graph model trained in the last training phase.
[0069] As a possible implementation method, quality evaluation can be performed on multiple sets of model parameters of the Wensheng graph large model trained in the last training stage to obtain evaluation indicators of the multiple sets of model parameters of the Wensheng graph large model trained in the last training stage. Based on the evaluation indicators, the model parameters to be fine-tuned are determined from the multiple sets of model parameters of the Wensheng graph large model trained in the last training stage.
[0070] As an example, when the evaluation indicator is a positive indicator (such as accuracy), the set of model parameters with relatively higher evaluation indicators can be used as the model parameters to be fine-tuned.
[0071] As another example, when the evaluation indicator is a negative indicator (such as loss), the set of model parameters with relatively low evaluation indicators can be used as the model parameters to be fine-tuned.
[0072] Step S103 : determining a first image-text data pair from the plurality of image-text data pairs based on the first quality scores and aesthetic scores of the sample images in the plurality of image-text data pairs.
[0073] Among them, the first quality score is used to indicate the image quality of the sample image. Exemplarily, the first quality score of the sample image can be calculated based on the image clarity, image resolution, semantic relevance between the sample image and the sample text (in the same image-text data pair as the sample image), etc.
[0074] Among them, the aesthetic score is used to indicate the aesthetic quality of the sample image, that is, whether the sample image conforms to aesthetics. For example, the aesthetic score of the sample image can be calculated based on the visual aesthetic elements in the sample image (aesthetic elements or visual elements such as color, light and shadow, contrast, composition, theme expression and emotional communication).
[0075] As an example, the image-text data pair containing the sample image having a relatively high first quality score and a relatively high aesthetic score may be used as the first image-text data pair.
[0076] Step S104 : using the first image-text data pair, fine-tuning the parameters of the model to be fine-tuned, so as to obtain target model parameters corresponding to the large text-image model.
[0077] In the embodiment of the present disclosure, the first image-text data pair can be used to fine-tune the model parameters to be fine-tuned to obtain the final set of model parameters obtained by training corresponding to the large text-image model, which are recorded as target model parameters in the present disclosure.
[0078] The training method of the Vincent graph large model of the embodiment of the present disclosure adopts multiple training stages to systematically train the Vincent graph large model, which can improve the training effect of the Vincent graph large model. In addition, in each training stage, the quality of the multiple sets of model parameters of the Vincent graph large model obtained by training in the previous training stage is evaluated to obtain evaluation indicators. Based on the evaluation indicators, high-precision model parameters are screened from the multiple sets of model parameters of the Vincent graph large model obtained by training in the previous training stage, and used as the model parameters of the Vincent graph large model to be trained in this training stage. This can improve the prediction quality of the Vincent graph large model finally trained and improve the user experience in the Vincent graph scene. In addition, using high-quality and aesthetically pleasing image and text data pairs to fine-tune the model parameters of the Vincent graph large model trained in the last training stage can improve the aesthetics and quality of the image output by the Vincent graph large model finally trained, further improving the user experience in the Vincent graph scene.
[0079] It should be noted that in the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, and are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0080] In order to clearly illustrate how any embodiment of the present disclosure determines a first image-text data pair from multiple image-text data pairs based on the first quality score and aesthetic score of the sample images in the multiple image-text data pairs, the present disclosure also proposes a training method for a large text-image model.
[0081] Figure 2 This is a flow chart of the training method of the large-scale Wensheng graph model provided in the second embodiment of the present disclosure.
[0082] like Figure 2 As shown, the training method of the large text-image model may include the following steps S201 to S207:
[0083] Step S201 : using multiple pairs of text-image data, sequentially executing multiple training phases on the large text-image model.
[0084] The model parameters of the Vincent graph large model to be trained in the i-th training stage are obtained by screening based on the evaluation indicators of multiple sets of model parameters of the Vincent graph large model trained in the i-1-th training stage, and i is a positive integer greater than 1.
[0085] Step S202 : determining the model parameters to be fine-tuned from the multiple sets of model parameters of the large Wensheng graph model obtained from the last training phase.
[0086] For explanations of steps S201 to S202 , reference may be made to the relevant descriptions in any embodiment of the present disclosure, and no further details will be given here.
[0087] Step S203 : determining a first quality score of the sample image in any image-text data pair according to the attribute information of the sample image in any image-text data pair and the semantic relevance between the sample image and the sample text in any image-text data pair.
[0088] The attribute information includes but is not limited to: image clarity, image resolution, etc.
[0089] In an embodiment of the present disclosure, for any one of a plurality of image-text data pairs, the semantic relevance between the sample image and the sample text in the image data pair can be calculated, and based on the semantic relevance and the attribute information of the sample image in the image data pair, the quality score of the sample image in the image-text data pair (referred to as the first quality score in the present disclosure) is calculated.
[0090] Among them, the first quality score is positively correlated with image clarity, the first quality score is positively correlated with image resolution, and the first quality score is positively correlated with semantic relevance.
[0091] Step S204 : determining candidate image-text data pairs from the plurality of image-text data pairs according to the first quality scores of the sample images in the plurality of image-text data pairs.
[0092] In an embodiment of the present disclosure, multiple image-text data pairs can be screened based on the first quality scores of sample images in the multiple image-text data pairs to obtain candidate image-text data pairs, wherein the first quality scores of sample images in the candidate image-text data pairs are relatively high. For example, the first quality scores of sample images in the candidate image-text data pairs can be higher than a set quality threshold.
[0093] Step S205 : determining an aesthetic score of the sample image in the candidate image-text data pair based on the visual aesthetic elements in the sample image in the candidate image-text data pair.
[0094] Among them, visual aesthetic elements include aesthetic elements or visual elements such as color, light and shadow, contrast, composition, theme expression and emotional communication.
[0095] In the embodiment of the present disclosure, the aesthetic score of the sample image in each candidate image-text data pair may be calculated based on the visual aesthetic elements in the sample image in the candidate image-text data pair.
[0096] For example, based on set rules, the visual aesthetic elements in the sample image in the candidate image-text data pair may be scored to obtain the aesthetic score of the sample image in the candidate image-text data pair.
[0097] For example, visual aesthetic elements in the sample images in the candidate image-text data pairs may be scored manually to obtain aesthetic scores for the sample images in the candidate image-text data pairs.
[0098] For example, a sample image in a candidate image-text data pair may be provided to an outsourced person, who then scores the visual aesthetic elements in the sample image to obtain an aesthetic score.
[0099] Step S206 : determining a first image-text data pair from the candidate image-text data pairs based on the aesthetic scores of the sample images in the candidate image-text data pairs.
[0100] In an embodiment of the present disclosure, candidate image-text data pairs can be screened based on the aesthetic scores of sample images in the candidate image-text data pairs to obtain a first image-text data pair, wherein the aesthetic score of the sample image in the first image-text data pair is relatively high. Exemplarily, the aesthetic score of the sample image in the first image-text data pair can be higher than a set score threshold.
[0101] Step S207 : using the first image-text data pair, fine-tune the parameters of the model to be fine-tuned to obtain target model parameters corresponding to the large text-image model.
[0102] For explanation of step S207, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0103] The training method of the large-scale cultural image model of the disclosed embodiment uses a high-quality and aesthetically pleasing first image-text data pair to fine-tune the model parameters of the large-scale cultural image model, thereby improving the aesthetics and quality of the image output by the final trained large-scale cultural image model and the user experience in the cultural image scenario.
[0104] In order to clearly illustrate how to obtain multiple image-text data pairs in any embodiment of the present disclosure, the present disclosure also proposes a training method for a large text-image model.
[0105] Figure 3This is a flow chart of the training method of the large-scale cultural graph model provided in the third embodiment of the present disclosure.
[0106] like Figure 3 As shown, the training method of the large text-image model may include the following steps S301 to S308:
[0107] Step S301 : Acquire a plurality of initial data pairs; wherein the initial data pairs include sample texts in a specified language and sample images corresponding to the sample texts.
[0108] The designated language refers to a language that is adapted to the target business scenario (e.g., Chinese), and the target business scenario refers to the business scenario in which the final trained text-based graph model is applied.
[0109] There is no restriction on the method of obtaining the initial data pairs. For example, the initial data pairs can be obtained from an existing training set, or the initial data pairs can be obtained from a self-owned training set, or the initial data pairs can be collected online, such as using web crawler technology to collect the initial data pairs online, or the initial data pairs can be generated manually, etc.
[0110] In any embodiment of the present disclosure, after obtaining multiple initial data pairs, text cleaning can be performed on the sample texts in the multiple image-text data pairs. For example, the sample texts can be filtered for HTML (Hypertext Markup Language) tags and special characters (such as emoticons, extra spaces, etc.), spelling errors in the sample texts can be corrected, and blank lines can be removed.
[0111] Therefore, by performing text cleaning on the sample text, the cleaned sample text can be made more suitable for subsequent processing and analysis of the large text-graph model, thereby improving the prediction accuracy of the model.
[0112] Step S302: Obtain the semantic relevance between the sample text and the sample image in the same initial data pair.
[0113] As an example, for any initial text data pair, the sample text in the initial data pair can be text-encoded to obtain text features, and the sample image in the initial data pair can be image-encoded to obtain image features. The text features and image features are mapped to a common feature space, and the semantic correlation between the text features and the image features is calculated in the common feature space.
[0114] As another example, for any initial text data pair, the sample text in the initial data pair can be text-encoded to obtain text features, and the sample image in the initial data pair can be image-encoded to obtain image features. The interaction between the text features and the image features is calculated through the cross-attention mechanism to generate weighted text features and weighted image features. The weighted text features and weighted image features are mapped to a common feature space to calculate the semantic correlation between the weighted text features and the weighted image features in the common feature space.
[0115] It should be understood that, in actual applications, other algorithms may also be used to calculate the semantic relevance between the sample text and the sample image, and the embodiments of the present disclosure do not limit this.
[0116] Step S303: Obtain a second quality score of the sample image in the same initial data pair.
[0117] The second quality score is used to indicate the image quality of the sample image.
[0118] In an embodiment of the present disclosure, a quality score (referred to as a second quality score in the present disclosure) of the sample image in the initial data pair may be calculated.
[0119] In any embodiment of the present disclosure, the second quality score can be calculated by following steps a to c:
[0120] Step a: Based on an image processing algorithm, perform image recognition on the sample images in the same initial data pair to obtain a recognition result; wherein the recognition result is used to indicate whether the sample image has at least one of black edges, mosaics, and visual illusions.
[0121] Step b: Obtain attribute information of sample images in the same initial data pair.
[0122] The attribute information includes but is not limited to: image clarity, image resolution, etc.
[0123] Step c: determining a second quality score of the sample image in the same initial data pair based on the attribute information and the recognition result of the sample image in the same initial data pair.
[0124] Among them, the second quality score is positively correlated with image clarity; the second quality score is also positively correlated with image resolution; when the recognition result indicates that the sample image has at least one of black edges, mosaics and visual illusions, the second quality score is relatively low, and when the recognition result indicates that the sample image does not have black edges, mosaics and visual illusions at the same time, the second quality score is relatively high.
[0125] In this way, the quality score of the sample image can be calculated by integrating the attribute information of the sample image (such as image clarity, image resolution, etc.) and the recognition results indicating whether the sample image has black edges, mosaics, or visual illusions, which can improve the rationality and accuracy of the quality score calculation.
[0126] Step S304 : determining a plurality of image-text data pairs from the plurality of initial data pairs based on the semantic relevance and the second quality scores of the plurality of initial data pairs.
[0127] In the embodiment of the present disclosure, multiple initial data pairs may be screened based on their semantic relevance and second quality scores to obtain multiple image-text data pairs.
[0128] As an example, an initial data pair whose semantic relevance is higher than a set relevance threshold and whose second quality score is higher than a set quality score threshold may be used as the finally obtained image-text data pair.
[0129] Step S305 , using multiple image-text data pairs, sequentially executing multiple training phases on the large text-image model.
[0130] The model parameters of the Vincent graph large model to be trained in the i-th training stage are obtained by screening based on the evaluation indicators of multiple sets of model parameters of the Vincent graph large model trained in the i-1-th training stage, and i is a positive integer greater than 1.
[0131] Step S306 , determining the model parameters to be fine-tuned from the multiple sets of model parameters of the large Wensheng graph model obtained through training in the last training phase.
[0132] Step S307 : determining a first image-text data pair from the plurality of image-text data pairs based on the first quality scores and aesthetic scores of the sample images in the plurality of image-text data pairs.
[0133] Step S308 : fine-tune the parameters of the model to be fine-tuned using the first image-text data pair to obtain target model parameters corresponding to the large text-image model.
[0134] For explanations of steps S305 to S308 , please refer to the relevant descriptions in any embodiment of the present disclosure and will not be repeated here.
[0135] The training method of the large-scale text-graph model of the disclosed embodiment can filter low-quality text-graph data pairs and use the high-quality text-graph data pairs obtained by screening to train the large-scale text-graph model, which can improve the training effect of the large-scale text-graph model, that is, improve the prediction accuracy of the large-scale text-graph model and improve the user experience in the text-graph scenario.
[0136] In order to clearly illustrate how to obtain a large model of a cultural graph in any embodiment of the present disclosure, the present disclosure also proposes a training method for a large model of a cultural graph.
[0137] Figure 4 This is a flow chart of the training method of the large-scale Wensheng graph model provided in the fourth embodiment of the present disclosure.
[0138] like Figure 4 As shown, the training method of the large text-image model may include the following steps S401 to S407:
[0139] Step S401: Determine an initial large model that is adapted to a target business scenario from a plurality of base large models.
[0140] The target business scenario refers to the business scenario in which the final trained large-scale cultural graph model is applied.
[0141] In an embodiment of the present disclosure, multiple base large models can be screened to obtain an initial large model that is adapted to the target business scenario.
[0142] For example, an initial large model that is adapted to the target business scenario may be manually selected from a plurality of base large models.
[0143] For example, there are a large number of open source large-scale cultural graph models on the Internet, such as the English large-scale cultural graph models include the SD (Stable Diffusion) series and the Flux series; the Chinese large-scale cultural graph models include Ketu and Hunyuan. You can choose the appropriate large-scale cultural graph model according to application requirements to take into account both model size and model effect.
[0144] Step S402: When the initial large model is not compatible with the specified language, a text encoder compatible with the specified language is used to replace the text encoder in the initial large model.
[0145] The designated language refers to the language that is suitable for the target business scenario (such as Chinese).
[0146] In the embodiment of the present disclosure, it is possible to determine whether the initial large model is compatible with the specified language. If so, the initial large model is used as the text-graph large model to be trained. If not, a text encoder compatible with the specified language is used to replace the text encoder in the initial large model to obtain an updated initial large model.
[0147] For example, taking Chinese as the designated language, when the initial large model is a Chinese text-graph large model, it can be determined that the initial large model is adapted to the designated language. At this time, the initial large model can be directly used as the text-graph large model to be trained; and when the initial large model is an English text-graph large model, the text encoder in the initial large model can be replaced with the Chinese version of the text encoder to obtain an updated initial large model.
[0148] Step S403: perform model alignment on the updated initial large model to obtain a large model of the Wensheng graph.
[0149] In the disclosed embodiment, the updated initial large model may also be aligned to obtain a large model of the cultural graph to be trained.
[0150] In any embodiment of the present disclosure, the following steps 1 to 4 may be used to perform model alignment on the updated initial large model:
[0151] Step 1: Determine a third image-text data pair for model alignment from multiple image-text data pairs.
[0152] As a possible implementation method, the third image-text data pair can be obtained by screening using the following steps A to B:
[0153] Step A: Query the resolution threshold for model alignment.
[0154] The resolution threshold is a threshold pre-configured for the model alignment task.
[0155] Step B: screening multiple image-text data pairs according to a resolution threshold to obtain a third image-text data pair; wherein the resolution of the sample image in the third image-text data pair is lower than the resolution threshold.
[0156] As another possible implementation, the third image-text data pair can be obtained by screening using the following steps A' to B':
[0157] Step A': sort the plurality of image-text data pairs from small to large according to the resolution of the corresponding sample images to obtain a sorted sequence.
[0158] Step B': taking a set number of image-text data pairs that are sorted first in the sorting sequence as the third image-text data pairs.
[0159] Therefore, using relatively low-resolution image-text data pairs as training data for model alignment can enable the large-scale text-image model to quickly align the semantics of the text side and the image side, thereby improving the alignment efficiency of the large-scale text-image model.
[0160] Step 2: Use the text encoder in the updated initial large model to encode the sample text in the third image-text data pair to obtain text features.
[0161] Step 3: Use the updated initial large model to encode the sample image in the third image-text data pair to obtain image features.
[0162] Step 4: Based on the semantic similarity between text features and image features, perform model alignment on the updated initial large model to obtain the text-image large model to be trained.
[0163] As an example, the value of the loss function (referred to as the loss value in this disclosure) can be determined based on the semantic similarity between text features and image features, so that the updated initial large model can be adjusted or trained based on the loss value.
[0164] Among them, the loss value is negatively correlated with the above-mentioned semantic similarity, that is, the loss value is used to measure the degree of mismatch between the sample image and the sample text.
[0165] In summary, by aligning the updated initial large model based on the semantic similarity between text features and image features, the text-to-image large model can iteratively optimize its parameters to better convert text in a specified language into corresponding image features and generate images that are highly semantically similar to the text in the specified language, thereby improving the prediction accuracy of the text-to-image large model.
[0166] Step S404: Acquire multiple image-text data pairs, and use the multiple image-text data pairs to sequentially perform multiple training phases on the large text-image model.
[0167] The model parameters of the Vincent graph large model to be trained in the i-th training stage are obtained by screening based on the evaluation indicators of multiple sets of model parameters of the Vincent graph large model trained in the i-1-th training stage, and i is a positive integer greater than 1.
[0168] Step S405 , determining the model parameters to be fine-tuned from the multiple sets of model parameters of the large Wensheng graph model obtained through training in the last training phase.
[0169] Step S406 : determining a first image-text data pair from the plurality of image-text data pairs based on the first quality scores and aesthetic scores of the sample images in the plurality of image-text data pairs.
[0170] Step S407 : using the first image-text data pair, fine-tune the parameters of the model to be fine-tuned to obtain target model parameters corresponding to the large text-image model.
[0171] The explanation of steps S404 to S407 can be found in the relevant description of any embodiment of the present disclosure, and will not be repeated here.
[0172] The training method of the large Wensheng graph model in the embodiment of the present disclosure can take into account both the model size and model effect of the final large Wensheng graph model through model selection and model alignment operations.
[0173] As a possible implementation method, the first training stage is taken as an example of the semantic tonal alignment stage. In order to clearly illustrate how the semantic tonal alignment stage is performed on the large model of the text graph in any embodiment of the present disclosure, the present disclosure also proposes a training method for the large model of the text graph.
[0174] Figure 5 This is a flowchart of the training method of the large-scale Wensheng graph model provided in the fifth embodiment of the present disclosure.
[0175] like Figure 5 As shown, the training method of the large text-image model may include the following steps S501 to S508:
[0176] Step S501 : determining a second image-text data pair adapted to the semantic tonality alignment stage from a plurality of image-text data pairs.
[0177] The number of the second graphic-text data pair may be at least one.
[0178] In an embodiment of the present disclosure, multiple image-text data pairs may be screened to obtain a second image-text data pair adapted to the semantic tonality alignment stage.
[0179] As an example, the second image-text data pair may be randomly selected from a plurality of image-text data pairs, or a high-quality second image-text data pair may be selected from a plurality of image-text data pairs.
[0180] Step S502 : Process the sample text in the second image-text data pair using the text-image macro model to obtain a first output image.
[0181] In an embodiment of the present disclosure, the text-image macro model may be used to process the sample text in the second text-image data pair to obtain an image output by the text-image macro model, which is referred to as the first output image in the present disclosure.
[0182] Step S503 : training the large Wensheng-image model based on the style differences and semantic differences between the first output image and the sample image in the second image-text data pair to obtain multiple sets of model parameters corresponding to the large Wensheng-image model trained in the semantic tone alignment stage.
[0183] In an embodiment of the present disclosure, the style difference (or style tonality difference) between the first output image and the sample image in the second image-text data pair can be calculated based on the image style of the first output image and the image style of the sample image in the second image-text data pair.
[0184] In the embodiment of the present disclosure, the semantic difference between the first output image and the sample image in the second image-text data pair may be calculated based on the image semantics of the first output image and the image semantics of the sample image in the second image-text data pair.
[0185] Therefore, in the present disclosure, the Wensheng graph large model can be trained according to the above-mentioned style differences and semantic differences to obtain multiple sets of model parameters corresponding to the Wensheng graph large model trained in the semantic tone alignment stage.
[0186] As an example, the value of the loss function corresponding to the semantic tonal alignment stage (hereinafter referred to as the loss value) can be determined based on the above-mentioned style differences and semantic differences. Therefore, the model parameters of the large Wenshengtu model can be adjusted based on the loss value to obtain multiple sets of model parameters corresponding to the large Wenshengtu model trained in the semantic tonal alignment stage.
[0187] Among them, the loss value is positively correlated with the style difference, and the loss value is positively correlated with the semantic difference.
[0188] In any embodiment of the present disclosure, during the training of the large-scale Wensheng graph model in the semantic tonality alignment stage, hyperparameters of the large-scale Wensheng graph model (such as the learning rate (which determines the step size of the model parameter update in each iteration), batch size (batch size, i.e., the number of samples passed to the model for training in one iteration), image size, etc.) can be dynamically adjusted, thereby improving the training speed and efficiency of the large-scale Wensheng graph model.
[0189] Step S504 , for the i-th training stage, based on the evaluation indicators of the multiple sets of model parameters of the Vincent graph large model trained in the i-1-th training stage, determine the model parameters of the Vincent graph large model to be trained in the i-th training stage from the multiple sets of model parameters of the Vincent graph large model trained in the i-1-th training stage.
[0190] Wherein, i is a positive integer greater than 1.
[0191] Step S505 , using multiple image-text data pairs, executing the i-th training phase on the large text-image model to be trained in the i-th training phase.
[0192] Step S506 , determining the model parameters to be fine-tuned from the multiple sets of model parameters of the large Wensheng graph model obtained through training in the last training phase.
[0193] Step S507 : determining a first image-text data pair from the plurality of image-text data pairs based on the first quality scores and aesthetic scores of the sample images in the plurality of image-text data pairs.
[0194] Step S508 : using the first image-text data pair, fine-tune the parameters of the model to be fine-tuned to obtain target model parameters corresponding to the large text-image model.
[0195] For explanations of steps S504 to S508 , reference may be made to the relevant descriptions in any embodiment of the present disclosure, and will not be repeated here.
[0196] The training method of the culture image large model in the embodiment of the present disclosure can enable the culture image large model to fully learn the image semantics and style tonality by performing a semantic tonality alignment stage on the culture image large model, thereby improving the learning effect and prediction accuracy of the culture image large model.
[0197] In any embodiment of the present disclosure, an example is given in which a non-first training stage includes a resolution enhancement stage. In order to clearly illustrate how the resolution enhancement stage is performed on the large model of the culture graph in any embodiment of the present disclosure, the present disclosure also proposes a training method for the large model of the culture graph.
[0198] Figure 6 This is a flow chart of the training method of the large-scale Wensheng graph model provided in the sixth embodiment of the present disclosure.
[0199] like Figure 6 As shown, the training method of the large text-image model may include the following steps S601 to S609:
[0200] Step S601 , for a first training phase among multiple training phases, multiple image-text data pairs are used to perform the first training phase on the large model of text-graph, and multiple sets of model parameters of the large model of text-graph trained in the first training phase are obtained.
[0201] Exemplarily, the first training stage may be a semantic tonality alignment stage. It should be noted that the first training stage may also be other training stages, such as a multi-scale training stage, etc., and the embodiments of the present disclosure do not impose any limitation on this.
[0202] For explanation of step S601, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0203] Step S602 : For a resolution improvement stage other than the first training stage in the multiple training stages, first evaluation indicators of multiple sets of model parameters of the Vincent graph large model trained in the previous training stage are obtained.
[0204] In an embodiment of the present disclosure, for a non-first training stage in multiple training stages, such as a resolution improvement stage, quality evaluation can be performed on multiple sets of model parameters of the Vincent graph large model trained in the previous training stage to obtain a first evaluation index.
[0205] In any embodiment of the present disclosure, the first evaluation index can be obtained by following steps X to Z:
[0206] Step X: Obtain a test data pair; wherein the test data pair includes a test text and a test image.
[0207] The method for obtaining the test data pair is similar to that for obtaining the initial data pair, and will not be described in detail here.
[0208] Step Y: For any set of model parameters of the Wenshengtu large model obtained by training in the previous training stage, the test text is input into the Wenshengtu large model with any set of model parameters built in, and the image output by the Wenshengtu large model is obtained, which is recorded as the second output image in this disclosure.
[0209] Step Z: Determine a first evaluation metric for any of the above sets of model parameters based on the degree of difference between the second output image and the test image.
[0210] Among them, when the first evaluation indicator is a positive indicator, the first evaluation indicator is negatively correlated with the above-mentioned difference degree, that is, the smaller the difference degree, the larger the first evaluation indicator, and conversely, the larger the difference degree, the smaller the first evaluation indicator.
[0211] Among them, when the first evaluation indicator is a negative indicator, the first evaluation indicator is positively correlated with the above-mentioned difference degree, that is, the greater the difference degree, the greater the first evaluation indicator, and conversely, the smaller the difference degree, the smaller the first evaluation indicator.
[0212] This allows us to calculate the evaluation metrics for a set of model parameters of the Wensheng graph model based on the loss of the Wensheng graph model on the test data pairs (calculated by the degree of difference, where the loss is positively correlated with the degree of difference). This can improve the effectiveness and accuracy of the evaluation metric calculation.
[0213] In any embodiment of the present disclosure, the first evaluation index can be obtained by following steps W' to Z':
[0214] Step W': Get the evaluation text.
[0215] There is no restriction on the method of obtaining the evaluation text. For example, the evaluation text can be provided manually, or the evaluation text can be collected online, or the evaluation text can be obtained from an existing test set, etc.
[0216] Step X': For any set of model parameters of the large Wensheng graph model trained in the previous training stage, input the evaluation text into the large Wensheng graph model with any set of model parameters built in, and obtain the image output by the Wensheng graph model, which is recorded as the third output image in this disclosure.
[0217] Step Y': Obtaining the image quality of the third output image; wherein the image quality is used to indicate image aesthetics and / or image clarity.
[0218] As an example, a machine review operator can be constructed to calculate the picture beauty and image clarity of the third output image, and the image quality of the third output image can be calculated based on the picture beauty and image clarity.
[0219] Step Z': determining a first evaluation index of any of the above-mentioned sets of model parameters according to the image quality of the third output image and the semantic relevance between the third output image and the evaluation text.
[0220] In the embodiment of the present disclosure, the semantic relevance between the third output image and the evaluation text can be calculated, and the image quality and semantic relevance of the third output image can be comprehensively considered to calculate the first evaluation index of any of the above-mentioned sets of model parameters.
[0221] In the case where the first evaluation index is a positive index, the first evaluation index is positively correlated with the image quality of the third output image, and the first evaluation index is also positively correlated with the semantic relevance.
[0222] In the case where the first evaluation index is a negative index, the first evaluation index is negatively correlated with the image quality of the third output image, and the first evaluation index is also negatively correlated with the semantic relevance.
[0223] In this way, the image quality of the output image of the comprehensive text-image model and the semantic relevance between the output image and the evaluation text can be used to calculate the evaluation index of a set of model parameters of the text-image model, which can improve the effectiveness, rationality and accuracy of the evaluation index calculation.
[0224] It should be understood that, in actual applications, other quality assessment algorithms may also be used to obtain the first assessment index of any set of model parameters of the Wensheng graph large model, and the embodiments of the present disclosure do not limit this.
[0225] Step S603 : Based on the first evaluation indicator, the model parameters of the Vincent graph large model to be trained in the resolution enhancement phase are determined from the multiple sets of model parameters of the Vincent graph large model trained in the previous training phase.
[0226] For example, the set of model parameters with the highest first evaluation index may be used as the model parameters of the large Wensheng graph model to be trained in the resolution improvement stage.
[0227] Step S604 : dividing the plurality of image-text data pairs to obtain groups of a plurality of resolutions; wherein the sample images in the image-text data pairs in the same group have the same resolution.
[0228] In the embodiment of the present disclosure, multiple graphic data pairs can be divided based on the resolution of sample images in the multiple graphic data pairs to obtain groups with multiple resolutions, wherein the resolutions of sample images in graphic data pairs in the same group are the same.
[0229] Step S605 , sorting the multiple groups in order from low to high resolution to obtain a sorted sequence.
[0230] In the embodiment of the present disclosure, multiple groups may be arranged in order from low to high resolution to obtain a sorted sequence.
[0231] Step S606 , based on the sorting sequence, sequentially executing multiple resolution improvement sub-stages in the resolution improvement stage on the large model of the Wensheng graph to be trained in the resolution improvement stage.
[0232] In the embodiment of the present disclosure, multiple resolution enhancement sub-stages in the resolution enhancement stage can be executed sequentially on the large model of the Vincent graph to be trained in the resolution enhancement stage based on the sorting sequence, that is, the large model of the Vincent graph can be trained according to the principle of low to high resolution.
[0233] In any embodiment of the present disclosure, for the first resolution improvement sub-stage among multiple resolution improvement sub-stages, the first group with the lowest resolution in the sorting sequence can be used to execute the first resolution improvement sub-stage on the Vincent graph large model to be trained in the resolution improvement stage, and obtain multiple sets of model parameters of the Vincent graph large model trained in the first resolution improvement sub-stage.
[0234] For the j-th (j is a positive integer greater than 1) resolution enhancement sub-stage among the multiple resolution enhancement sub-stages, a second evaluation indicator of the multiple sets of model parameters of the Vincent graph large model trained in the j-1-th resolution enhancement sub-stage can be obtained (wherein, a method for obtaining the second evaluation indicator is similar to a method for obtaining the first evaluation indicator and is not further described herein). Based on the second evaluation indicator, model parameters of the Vincent graph large model to be trained in the j-th resolution enhancement sub-stage are determined from the multiple sets of model parameters of the Vincent graph large model trained in the j-1-th resolution enhancement sub-stage (the implementation principle is similar to step S603 and is not further described herein). Subsequently, the grouping at the j-th position in the sorted sequence can be used as training data for the j-th resolution enhancement sub-stage. Finally, the training data of the j-th resolution enhancement sub-stage can be used to adjust the model parameters of the Vincent graph large model to be trained in the j-th resolution enhancement sub-stage, thereby obtaining the multiple sets of model parameters of the Vincent graph large model trained in the j-th resolution enhancement sub-stage.
[0235] It should be understood that the multiple sets of model parameters of the Vincent image large model obtained by training in the last resolution improvement sub-stage are the multiple sets of model parameters of the Vincent image large model obtained by training in the resolution improvement stage.
[0236] Therefore, following the principle of increasing resolution from low to high, the resolution of the image is continuously improved to perform resolution improvement training on the large-scale cultural image model, which can improve the learning effect of the large-scale cultural image model and improve the prediction accuracy of the large-scale cultural image model.
[0237] Step S607 , determining the model parameters to be fine-tuned from the multiple sets of model parameters of the large Wensheng graph model obtained from the last training phase.
[0238] Step S608 : determining a first image-text data pair from the plurality of image-text data pairs based on the first quality scores and aesthetic scores of the sample images in the plurality of image-text data pairs.
[0239] Step S609 : using the first image-text data pair, fine-tune the parameters of the model to be fine-tuned to obtain target model parameters corresponding to the large text-image model.
[0240] For explanations of steps S607 to S609 , reference may be made to the relevant descriptions in any embodiment of the present disclosure, and will not be repeated here.
[0241] The training method of the large-scale model of the Wensheng graph in the embodiment of the present disclosure follows the principle of increasing resolution from low to high, and performs resolution improvement training on the large-scale model of the Wensheng graph, which can improve the learning effect of the large-scale model of the Wensheng graph and improve the prediction accuracy of the large-scale model of the Wensheng graph.
[0242] In any embodiment of the present disclosure, an example is given in which a non-first training stage includes a multi-scale training stage. In order to clearly illustrate how a multi-scale training stage is performed on a large model of a culture graph in any embodiment of the present disclosure, the present disclosure also proposes a training method for a large model of a culture graph.
[0243] Figure 7 This is a flow chart of the training method of the large-scale Wensheng graph model provided in the seventh embodiment of the present disclosure.
[0244] like Figure 7 As shown, the training method of the large text-image model may include the following steps S701 to S708:
[0245] Step S701 , for a first training phase among multiple training phases, multiple image-text data pairs are used to perform the first training phase on the large model of text-graph data to obtain multiple sets of model parameters of the large model of text-graph data trained in the first training phase.
[0246] Exemplarily, the first training stage may be a semantic tonality alignment stage. It should be noted that the first training stage may also be other training stages, such as a resolution enhancement stage, etc., and the embodiments of the present disclosure do not impose any limitation on this.
[0247] For explanation of step S701, please refer to the relevant description in any embodiment of the present disclosure, and will not be repeated here.
[0248] Step S702 : For a multi-size training stage other than the first training stage in the multiple training stages, obtain third evaluation indicators of multiple sets of model parameters of the Vincent graph large model trained in a previous training stage.
[0249] The method for obtaining the third evaluation indicator is similar to that for obtaining the first evaluation indicator, and will not be elaborated here.
[0250] Step S703 : Based on the third evaluation indicator, the model parameters of the Vincent graph large model to be trained in the multi-scale training phase are determined from the multiple sets of model parameters of the Vincent graph large model trained in the previous training phase.
[0251] For example, the set of model parameters with the highest third evaluation index may be used as the model parameters of the large Wensheng graph model to be trained in the multi-scale training phase.
[0252] Step S704 , obtaining buckets of multiple sizes obtained by bucketing multiple image-text data pairs; wherein the sizes of sample images in the image-text data pairs in the same bucket are the same.
[0253] In an embodiment of the present disclosure, multiple image-text data pairs can be divided into multi-size buckets based on the sizes of sample images in the multiple image-text data pairs to obtain buckets of multiple sizes; wherein the sizes of sample images in the image-text data pairs in the same bucket are the same.
[0254] Step S705 : Based on the bucketing of multiple sizes, multi-size learning is performed on the large model of the Wensheng graph to be trained in the multi-size training phase.
[0255] In the embodiment of the present disclosure, multi-scale learning is performed on the large model of the Wensheng graph to be trained in the multi-scale training stage based on bucketing of multiple sizes.
[0256] As an example, in each step of the multi-scale training stage, a bucket is randomly selected, and then a text-image data pair is randomly selected in the bucket to train the large text-image model to be trained in the multi-scale training stage.
[0257] Step S706 , determining the model parameters to be fine-tuned from the multiple sets of model parameters of the large Wensheng graph model obtained through training in the last training phase.
[0258] Step S707 : determining a first image-text data pair from the plurality of image-text data pairs based on the first quality scores and aesthetic scores of the sample images in the plurality of image-text data pairs.
[0259] Step S708 : using the first image-text data pair, fine-tune the parameters of the model to be fine-tuned to obtain target model parameters corresponding to the large text-image model.
[0260] For explanations of steps S706 to S708 , reference may be made to the relevant descriptions in any embodiment of the present disclosure and will not be repeated here.
[0261] The training method of the large-scale culture graph model of the disclosed embodiment performs multi-scale learning on the large-scale culture graph model by using image and text data pairs in buckets of multiple sizes, which can further improve the learning effect of the large-scale culture graph model and improve the prediction accuracy of the large-scale culture graph model.
[0262] In any embodiment of the present disclosure, taking Chinese as an example, the present disclosure proposes a systematic training method for a large model of a Chinese language graph, which mainly includes the following steps: Figure 8 The five steps shown are: preparing training data, selecting a model, aligning the model, training the model, and evaluating the model. First, you prepare the training data. Then, based on application requirements, you select the appropriate architecture and model. Finally, you train and align the model, and then you train the entire system. During this process, you continuously evaluate the model's performance to facilitate the next stage of training and select the final model.
[0263] Among them, the model training part includes four sub-modules, namely: semantic tone alignment, resolution improvement, multi-size training, and aesthetic improvement.
[0264] The following sections describe each section in detail:
[0265] The first part is preparing training data.
[0266] To train a large Chinese text-to-graph model, you need to prepare sufficient text-to-graph data pairs and perform the following operations on the data pairs:
[0267] 1. Filter out image-text data pairs containing low-quality images, such as those with low clarity, black edges, mosaics, and visual illusions;
[0268] 2. Filter out image-text data pairs with poor image-text relevance;
[0269] 3. Clean the text in the image and text data, such as the HTML tags of the web page.
[0270] The second part is model selection.
[0271] Choose a suitable base model: Currently, there are a large number of open-source cultural image models online. For example, the English cultural image models include the SD (Stable Diffusion) series and the Flux series; the Chinese cultural image models include Ketu and Hunyuan. You can choose a suitable base model based on application requirements, taking into account both model size and model effect.
[0272] The third part is model alignment.
[0273] If it is an English text-image model, the text encoder in the text-image model needs to be replaced with the Chinese version of the text encoder. Then, a large number of low-resolution images are used to quickly align the semantics of the text and image sides.
[0274] If it is a large model of a Chinese cultural image, this step is not required.
[0275] The fourth part is model training. The model training process can be divided into the following training stages:
[0276] 1) Semantic-Tonal Alignment: Fully learn the semantics and stylistic tonality of the image. To improve model training speed and efficiency, you can adjust model hyperparameters (such as learning rate, batch size, and image size).
[0277] 2) Resolution improvement stage: During model training, the principle of increasing resolution from low to high should be followed. First, train at 256-resolution. When semantic learning is sufficient, select the optimal model parameters and enter the 512-resolution learning stage. Continue in this way to continuously improve the image resolution.
[0278] 3) Multi-scale training: After training to the highest resolution, the model needs to be trained at multiple scales. For example, for a target resolution of 1024*1024, sample images can be bucketed at multiple scales, while maintaining the same length-width product, such as 512*2048, 256*4096, 2048*512, and 4096*256. Before training, sample images are pre-assigned to each bucket. At each training step, a random bucket is first selected, and then a sample image is randomly selected from that bucket. This training process repeats.
[0279] 4) Fine-tuning, or aesthetic enhancement, phase: First, based on the training data, sample images are screened for "high-quality" images based on their inherent properties and image quality operators. For example, images with resolution greater than x, clarity greater than y, and correlation greater than z are selected. These images are then sent to an outsourced annotation service to identify aesthetically pleasing images (e.g., in terms of color, lighting, and contrast). Finally, these aesthetic images are used to fine-tune the model parameters. Training should be stopped promptly to prevent overfitting. For example, if the model's evaluation metrics do not improve after a certain step, training can be stopped.
[0280] Part V, Model Evaluation.
[0281] During model training, it is necessary to continuously select the optimal model parameters. This serves as the starting model for the next stage of training and also allows the final model to be selected for use. Model evaluation methods can include the following two methods, which can also be used in combination:
[0282] (1) Selection based on loss: Construct an evaluation set or test set. After training x steps, look at the loss of the model on the evaluation set. Select the model parameters with the lowest loss on the evaluation set, which are the possible optimal model parameters. The implementation principle is as follows: Figure 6 As shown in step X to step Z in step S602.
[0283] (2) Selection based on machine review operators: Construct machine review operators, such as image-text relevance, image aesthetics, and image clarity, and comprehensively select the optimal model parameters based on the scores of different machine review operators. The implementation principle is as follows: Figure 6 As shown in steps W' to Z' in step S602.
[0284] In summary, the solution provided by this disclosure enables systematic training of a large-scale Chinese cultural image model, efficiently and rapidly generating high-quality image data. AI painting products that utilize this trained cultural image model can increase clicks and downloads, increase user engagement, and attract more users, thereby improving product DAU (Daily Active Users) and retention.
[0285] The above are various embodiments corresponding to the training method of the large-scale culture graph model. The present disclosure also proposes an application method of the large-scale culture graph model, namely, a culture graph method.
[0286] Figure 9 This is a flow chart of the Wensheng diagram method provided in Example 8 of the present disclosure.
[0287] like Figure 9 As shown, the text-generated graph method may include the following steps S901 to S902:
[0288] Step S901: Obtain input text.
[0289] There is no restriction on the method for obtaining the input text. For example, the input text may be text input by a user, or the input text may be text collected online, and so on.
[0290] Among them, the input method of inputting text includes but is not limited to touch input (such as sliding, clicking, etc.), keyboard input, voice input, etc.
[0291] Step S902 : calling a large text-image model with built-in target model parameters to process the input text and obtain a target image.
[0292] Among them, the Wenshengtu model is adopted Figures 1 to 8 The method shown in any one of the embodiments is used for training.
[0293] In the embodiment of the present disclosure, a large text-image model with built-in target model parameters may be called to process the input text to obtain a target image output by the large text-image model.
[0294] The text-based image method of the disclosed embodiment uses a trained text-based image large model to automatically generate images required by users, which can improve the quality and efficiency of image generation. In addition, users only need to provide input text to automatically generate images that semantically match the input text, which can reduce the user's operation steps and lower the technical and usage thresholds.
[0295] With the above Figures 1 to 7Corresponding to the training method of the large model of the Wensheng graph provided in the embodiment, the present disclosure also provides a training device for the large model of the Wensheng graph. Figures 1 to 7 The training method of the large model of the Wensheng graph provided in the embodiment corresponds to the embodiment, so the implementation method of the training method of the large model of the Wensheng graph is also applicable to the training device of the large model of the Wensheng graph provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0296] Figure 10 This is a structural diagram of the training device for the large-scale Wensheng graph model provided in Example 9 of the present disclosure.
[0297] like Figure 10 As shown, the training device 1000 for the large-scale model of the culture graph may include: an execution module 1010 , a screening module 1020 , a determination module 1030 and a fine-tuning module 1040 .
[0298] The execution module 1010 is configured to sequentially execute multiple training phases on the Wensheng graph large model using multiple image-text data pairs; wherein the model parameters of the Wensheng graph large model to be trained in the i-th training phase are obtained by screening based on the evaluation indicators of multiple sets of model parameters of the Wensheng graph large model trained in the i-1-th training phase, where i is a positive integer greater than 1;
[0299] A screening module 1020 is used to determine the model parameters to be fine-tuned from the multiple sets of model parameters of the Wensheng graph large model trained in the last training phase;
[0300] A determination module 1030 is configured to determine a first image-text data pair from the plurality of image-text data pairs based on the first quality scores and aesthetic scores of the sample images in the plurality of image-text data pairs;
[0301] The fine-tuning module 1040 is configured to fine-tune the parameters of the model to be fine-tuned using the first image-text data pair to obtain target model parameters corresponding to the large text-image model.
[0302] In a possible implementation of the embodiment of the present disclosure, the first training stage of the multiple training stages is the semantic tonal alignment stage, and the execution module 1010 executes the semantic tonal alignment stage, specifically: determining a second image-text data pair adapted to the semantic tonal alignment stage from multiple image-text data pairs; using the text-image large model to process the sample text in the second image-text data pair to obtain a first output image; training the text-image large model based on the style differences and semantic differences between the first output image and the sample image in the second image-text data pair to obtain multiple sets of model parameters corresponding to the text-image large model trained in the semantic tonal alignment stage.
[0303] In a possible implementation of the embodiment of the present disclosure, the execution module 1010 is further configured to adjust hyperparameters of the large-scale culture graph model during the training process of the large-scale culture graph model.
[0304] In a possible implementation of the embodiment of the present disclosure, a non-first training stage in the multiple training stages includes a resolution improvement stage, and the execution module 1010 executes the resolution improvement stage, specifically: obtaining a first evaluation index of multiple sets of model parameters of the Vincent image large model trained in the previous training stage; based on the first evaluation index, determining the model parameters of the Vincent image large model to be trained in the resolution improvement stage from the multiple sets of model parameters of the Vincent image large model trained in the previous training stage; dividing multiple image-text data pairs to obtain multiple resolution groups; wherein the resolutions of sample images in the image-text data pairs in the same group are the same; sorting the multiple groups in order from low to high resolution to obtain a sorted sequence; based on the sorted sequence, sequentially executing multiple resolution improvement sub-stages in the resolution improvement stage on the Vincent image large model to be trained in the resolution improvement stage.
[0305] In a possible implementation of the embodiment of the present disclosure, the execution module 1010 executes the j-th second training sub-stage, specifically: obtaining a second evaluation index of multiple sets of model parameters of the Vincent image large model obtained by training in the j-1-th resolution improvement sub-stage; wherein j is a positive integer greater than 1; based on the second evaluation index, determining the model parameters of the Vincent image large model to be trained in the j-th resolution improvement sub-stage from the multiple sets of model parameters of the Vincent image large model obtained by training in the j-1-th resolution improvement sub-stage; using the grouping at the j-th position in the sorted sequence as training data for the j-th resolution improvement sub-stage; and using the training data of the j-th resolution improvement sub-stage to adjust the model parameters of the Vincent image large model to be trained in the j-th resolution improvement sub-stage.
[0306] In one possible implementation of the embodiments of the present disclosure, execution module 1010 is configured to: obtain a test data pair, wherein the test data pair includes a test text and a test image; input the test text into the Vincent graph large model having any set of model parameters built therein for any set of model parameters obtained by training in a previous training phase, to obtain a second output image; and determine a first evaluation indicator for any set of model parameters based on a degree of difference between the second output image and the test image.
[0307] In a possible implementation of the embodiment of the present disclosure, the execution module 1010 is used to: obtain an evaluation text; for any set of model parameters of the Vincent graph large model trained in the previous training stage, input the evaluation text into the Vincent graph large model with any set of model parameters built in to obtain a third output image; obtain the image quality of the third output image; wherein the image quality is used to indicate the image aesthetics and / or image clarity; and determine a first evaluation indicator of any set of model parameters based on the image quality of the third output image and based on the semantic relevance between the third output image and the evaluation text.
[0308] In a possible implementation of the embodiment of the present disclosure, a non-first training stage in the multiple training stages includes a multi-size training stage, and the execution module 1010 executes the multi-size training stage, specifically: obtaining a third evaluation indicator of multiple sets of model parameters of the Vincent graph large model trained in the previous training stage; based on the third evaluation indicator, determining the model parameters of the Vincent graph large model to be trained in the multi-size training stage from the multiple sets of model parameters of the Vincent graph large model trained in the previous training stage; obtaining multiple-size buckets obtained by bucketing multiple image-text data pairs; wherein the sizes of sample images in the image-text data pairs in the same bucket are the same; and performing multi-size learning on the Vincent graph large model to be trained in the multi-size training stage based on the multiple-size buckets.
[0309] In a possible implementation of the embodiments of the present disclosure, the determination module 1030 is used to: determine a first quality score of a sample image in any image-text data pair based on the attribute information of the sample image in any image-text data pair and the semantic relevance between the sample image and the sample text in any image-text data pair; determine a candidate image-text data pair from a plurality of image-text data pairs based on the first quality scores of the sample images in the plurality of image-text data pairs; determine an aesthetic score of the sample image in the candidate image-text data pair based on the visual aesthetic elements in the sample image in the candidate image-text data pair; and determine a first image-text data pair from the candidate image-text data pairs based on the aesthetic score of the sample image in the candidate image-text data pair.
[0310] In one possible implementation of the embodiment of the present disclosure, the text-graph model is applicable to input text in a specified language adapted to the target business scenario, and multiple graph-text data pairs are obtained using the following modules:
[0311] A first acquisition module is configured to acquire a plurality of initial data pairs, wherein the initial data pairs include sample texts in a specified language and sample images corresponding to the sample texts;
[0312] The second acquisition module is used to obtain the semantic relevance between the sample text and the sample image in the same initial data pair;
[0313] a third acquisition module, configured to acquire a second quality score of the sample image in the same initial data pair;
[0314] The fourth acquisition module is configured to determine a plurality of image-text data pairs from the plurality of initial data pairs based on the semantic relevance and the second quality scores of the plurality of initial data pairs.
[0315] In a possible implementation of the embodiment of the present disclosure, the third acquisition module is used to: perform image recognition on the sample image in the same initial data pair to obtain a recognition result; wherein the recognition result is used to indicate whether the sample image has at least one of black edges, mosaics and visual illusions; obtain attribute information of the sample image in the same initial data pair; and determine a second quality score of the sample image in the same initial data pair based on the attribute information and recognition result of the sample image in the same initial data pair.
[0316] In a possible implementation of the embodiment of the present disclosure, the training device 1000 for the large-scale cultural graph model further includes:
[0317] The cleaning module is used to clean the sample texts in multiple image-text data pairs.
[0318] In a possible implementation of the embodiment of the present disclosure, the large model of the cultural graph is obtained using the following modules:
[0319] A selection module is used to determine an initial large model that is suitable for the target business scenario from multiple base large models;
[0320] A replacement module is used to replace the text encoder in the initial large model with a text encoder adapted to the specified language if the initial large model is not compatible with the specified language; the specified language refers to a language adapted to the target business scenario;
[0321] The alignment module is used to align the updated initial large model to obtain the Wensheng graph large model.
[0322] In a possible implementation of the embodiment of the present disclosure, the alignment module is used to: determine a third image-text data pair for model alignment from multiple image-text data pairs; use the updated initial large model to encode the sample text in the third image-text data pair to obtain text features; use the updated initial large model to encode the sample image in the third image-text data pair to obtain image features; based on the semantic similarity between the text features and the image features, perform model alignment on the updated initial large model to obtain a text-image large model.
[0323] In a possible implementation of an embodiment of the present disclosure, the alignment module is used to: query a resolution threshold for model alignment; determine a third image-text data pair from multiple image-text data pairs based on the resolution threshold; wherein the resolution of the sample image in the third image-text data pair is lower than the resolution threshold.
[0324] The training device for the large model of the Wensheng graph in the embodiment of the present disclosure uses multiple training stages to systematically train the large model of the Wensheng graph, which can improve the training effect of the large model of the Wensheng graph. In addition, in each training stage, the quality of the multiple sets of model parameters of the large model of the Wensheng graph obtained by training in the previous training stage is evaluated to obtain evaluation indicators. Based on the evaluation indicators, high-precision model parameters are screened from the multiple sets of model parameters of the large model of the Wensheng graph obtained by training in the previous training stage, and used as the model parameters of the large model of the Wensheng graph to be trained in this training stage. This can improve the prediction quality of the large model of the Wensheng graph obtained by the final training and improve the user experience in the Wensheng graph scene. In addition, by using high-quality and aesthetically pleasing image and text data pairs to fine-tune the model parameters of the large model of the Wensheng graph obtained by training in the last training stage, the aesthetics and quality of the image output by the large model of the Wensheng graph obtained by training can be improved, further improving the user experience in the Wensheng graph scene.
[0325] With the above Figure 9 Corresponding to the method for the text graph provided in the embodiment, the present disclosure also provides a text graph device. Figure 9 The embodiment provides a corresponding method for the Wensheng diagram, so the implementation of the Wensheng diagram method is also applicable to the Wensheng diagram device provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0326] Figure 11 This is a structural diagram of the Wenshengtu device provided in the tenth embodiment of the present disclosure.
[0327] like Figure 11 As shown, the text image device 1100 may include: an acquisition module 1110 and a calling module 1120.
[0328] Wherein, the acquisition module 1110 is used to acquire input text;
[0329] The calling module 1120 is used to call the Wenshengtu model with built-in target model parameters to process the input text and obtain the target image; wherein the Wenshengtu model is a Figure 10 The device shown is trained.
[0330] The text-based image device of the disclosed embodiment uses a trained text-based image large model to automatically generate images required by users, which can improve the quality and efficiency of image generation. In addition, users only need to provide input text to automatically generate images that semantically match the input text, which can reduce the user's operation steps and lower the technical and usage thresholds.
[0331] In order to implement the above embodiments, the present disclosure also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method of the large model of the Vincent graph or the Vincent graph method proposed in any of the above embodiments of the present disclosure.
[0332] In order to implement the above embodiments, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the training method of the Vincent graph large model or the Vincent graph method proposed in any of the above embodiments of the present disclosure.
[0333] In order to implement the above embodiments, the present disclosure also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the training method of the Wensheng graph large model or the Wensheng graph method proposed in any of the above embodiments of the present disclosure.
[0334] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0335] Figure 12 A schematic block diagram of an example electronic device that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0336] like Figure 12As shown, electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes based on computer programs stored in ROM (Read-Only Memory) 1202 or loaded from storage unit 1208 into RAM (Random Access Memory) 1203. RAM 1203 may also store various programs and data required for the operation of device 1200. Computing unit 1201, ROM 1202, and RAM 1203 are interconnected via bus 1204. An I / O (Input / Output) interface 1205 is also connected to bus 1204.
[0337] Various components in device 1200 are connected to I / O interface 1205, including an input unit 1206, such as a keyboard and mouse; an output unit 1207, such as various types of displays and speakers; a storage unit 1208, such as a magnetic disk and optical disk; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0338] Computing unit 1201 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1201 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. Computing unit 1201 performs the various methods and processes described above, such as the aforementioned training method for the large-scale Wensheng graph model or the Wensheng graph method. For example, in some embodiments, the training method for the large-scale Wensheng graph model or the Wensheng graph method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by computing unit 1201, one or more steps of the above-described method for training a large model of a Wensheng graph or the Wensheng graph method can be performed. Alternatively, in other embodiments, computing unit 1201 can be configured to perform the above-described method for training a large model of a Wensheng graph or the Wensheng graph method by any other suitable means (e.g., via firmware).
[0339] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0340] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0341] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0342] To provide for user interaction, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic input, voice input, or tactile input.
[0343] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0344] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS (Virtual Private Server) services. The server may also be a server in a distributed system or a server integrated with blockchain.
[0345] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0346] According to the technical solution of the embodiment of the present disclosure, a plurality of training stages are used to systematically train the large model of the Wensheng graph, which can improve the training effect of the large model of the Wensheng graph. In addition, in each training stage, the quality of the multiple sets of model parameters of the large model of the Wensheng graph obtained by training in the previous training stage is evaluated to obtain evaluation indicators. Based on the evaluation indicators, high-precision model parameters are screened from the multiple sets of model parameters of the large model of the Wensheng graph obtained by training in the previous training stage, and used as the model parameters of the large model of the Wensheng graph to be trained in this training stage. This can improve the prediction quality of the large model of the Wensheng graph obtained by the final training, and improve the user experience in the Wensheng graph scene. In addition, using high-quality and aesthetically pleasing image and text data pairs to fine-tune the model parameters of the large model of the Wensheng graph obtained by training in the last training stage can improve the aesthetics and quality of the image output by the large model of the Wensheng graph obtained by training, further improving the user experience in the Wensheng graph scene.
[0347] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0348] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A training method for a large-scale cultural graph model, comprising: Using multiple image-text data pairs, sequentially perform multiple training phases on the Wensheng graph large model; wherein the model parameters of the Wensheng graph large model to be trained in the i-th training phase are obtained by screening based on the evaluation indicators of multiple sets of model parameters of the Wensheng graph large model trained in the i-1-th training phase, where i is a positive integer greater than 1; Determine the model parameters to be fine-tuned from the multiple sets of model parameters of the large Wensheng graph model obtained from the last training phase; determining a first image-text data pair from the plurality of image-text data pairs based on first quality scores and aesthetic scores of sample images in the plurality of image-text data pairs; Using the first image-text data pair, fine-tune the parameters of the model to be fine-tuned to obtain target model parameters corresponding to the large text-image model; The multiple training stages include a semantic tonal alignment stage, and the semantic tonal alignment stage includes: Using the text-image macro model, processing the sample text in the second image-text data pair adapted to the semantic tone alignment stage among the multiple image-text data pairs to obtain a first output image; The large Wensheng-image model is trained based on the style differences and semantic differences between the first output image and the sample images in the second image-text data pair to obtain multiple sets of model parameters of the large Wensheng-image model trained in the semantic tone alignment stage.
2. The method according to claim 1, wherein The semantic-tonal alignment stage is the first training stage among the multiple training stages.
3. The method according to claim 2, wherein: The method further comprises: During the training process of the large-scale Wensheng graph model, hyperparameters of the large-scale Wensheng graph model are adjusted.
4. The method according to claim 1, wherein The non-first training phase of the plurality of training phases includes a resolution improvement phase, and the resolution improvement phase includes: Obtaining first evaluation indicators of multiple sets of model parameters of the large Wensheng graph model trained in the previous training phase; Determining, based on the first evaluation metric, the model parameters of the Vincent graph large model to be trained in the resolution enhancement phase from the multiple sets of model parameters of the Vincent graph large model trained in the previous training phase; Dividing the plurality of image-text data pairs to obtain groups of a plurality of resolutions; wherein the sample images in the image-text data pairs in the same group have the same resolution; Sorting the multiple groups in ascending order of resolution to obtain a sorted sequence; Based on the sorting sequence, multiple resolution enhancement sub-stages in the resolution enhancement stage are sequentially executed on the large Wensheng graph model to be trained in the resolution enhancement stage.
5. The method according to claim 4, wherein The j-th resolution improvement sub-stage includes: Obtaining second evaluation indicators of multiple sets of model parameters of the Vincent image large model trained in the j-1th resolution enhancement sub-stage; where j is a positive integer greater than 1; Determining, based on the second evaluation metric, the model parameters of the Vincent image large model to be trained in the j-th resolution enhancement sub-stage from the multiple sets of model parameters of the Vincent image large model trained in the j-1-th resolution enhancement sub-stage; Using the group at the j-th position in the sorted sequence as training data for the j-th resolution improvement sub-stage; The training data of the j-th resolution enhancement sub-stage is used to adjust the model parameters of the large Vincent image model to be trained in the j-th resolution enhancement sub-stage.
6. The method according to claim 4, wherein: The first evaluation index of obtaining multiple sets of model parameters of the Wensheng graph large model obtained by training in the previous training phase includes: Acquire a test data pair; wherein the test data pair includes a test text and a test image; For any set of model parameters of the Wenshengtu large model trained in the previous training phase, input the test text into the Wenshengtu large model with the any set of model parameters built in, to obtain a second output image; A first evaluation metric for any one set of model parameters is determined based on a degree of difference between the second output image and the test image.
7. The method according to claim 4, wherein: The first evaluation index of obtaining multiple sets of model parameters of the Wensheng graph large model obtained by training in the previous training phase includes: Get the evaluation text; For any set of model parameters of the Vincent graph large model trained in the previous training phase, input the evaluation text into the Vincent graph large model having the any set of model parameters built therein to obtain a third output image; Acquiring image quality of the third output image; wherein the image quality is used to indicate image aesthetics and / or image clarity; A first evaluation index of any one set of model parameters is determined based on the image quality of the third output image and the semantic relevance between the third output image and the evaluation text.
8. The method according to claim 1, wherein The non-first training phase of the plurality of training phases includes a multi-size training phase, and the multi-size training phase includes: Obtaining a third evaluation index of multiple sets of model parameters of the large Wensheng graph model trained in the previous training phase; Determining, based on the third evaluation metric, the model parameters of the Vincent graph large model to be trained in the multi-scale training phase from the multiple sets of model parameters of the Vincent graph large model trained in the previous training phase; Obtaining buckets of multiple sizes obtained by bucketing the multiple image-text data pairs; wherein the sample images in the image-text data pairs in the same bucket have the same size; Based on the bucketing of the multiple sizes, multi-size learning is performed on the large Wensheng graph model to be trained in the multi-size training phase.
9. The method according to any one of claims 1 to 8, wherein The determining a first image-text data pair from the plurality of image-text data pairs based on the first quality scores and aesthetic scores of the sample images in the plurality of image-text data pairs comprises: Determining a first quality score of the sample image in any image-text data pair based on attribute information of the sample image in any image-text data pair and a semantic relevance between the sample image and the sample text in the image-text data pair; determining a candidate image-text data pair from the plurality of image-text data pairs according to first quality scores of sample images in the plurality of image-text data pairs; Determining an aesthetic score of the sample image in the candidate image-text data pair based on visual aesthetic elements in the sample image in the candidate image-text data pair; The first image-text data pair is determined from the candidate image-text data pairs based on the aesthetic scores of the sample images in the candidate image-text data pairs.
10. The method according to any one of claims 1 to 8, wherein The text-to-graph model is applicable to input text in a specified language adapted to the target business scenario. The multiple graph-to-text data pairs are obtained by the following steps: Acquire a plurality of initial data pairs; wherein the initial data pairs include sample texts in the specified language and sample images corresponding to the sample texts; Obtaining the semantic relevance between the sample text and the sample image in the same initial data pair; Obtaining a second quality score of the sample image in the same initial data pair; The plurality of image-text data pairs are determined from the plurality of initial data pairs based on the semantic relevance and the second quality score of the plurality of initial data pairs.
11. The method according to claim 10, wherein: The obtaining of a second quality score of the sample image in the same initial data pair includes: Performing image recognition on the sample image in the same initial data pair to obtain a recognition result; wherein the recognition result is used to indicate whether the sample image has at least one of black edges, mosaics, and visual illusions; Acquiring attribute information of the sample image in the same initial data pair; A second quality score of the sample image in the same initial data pair is determined according to the attribute information and the recognition result of the sample image in the same initial data pair.
12. The method according to claim 10, wherein: The method further comprises: Perform text cleaning on the sample texts in the multiple image-text data pairs.
13. The method according to any one of claims 1 to 8, wherein The Wensheng graph model is obtained by following the steps below: Determine the initial large model that is suitable for the target business scenario from multiple base large models; If the initial large model is not compatible with the specified language, replacing the text encoder in the initial large model with a text encoder compatible with the specified language; wherein the specified language refers to a language compatible with the target business scenario; Model alignment is performed on the updated initial large model to obtain the Wensheng graph large model.
14. The method according to claim 13, wherein The step of performing model alignment on the updated initial large model to obtain the Wensheng graph large model includes: Determining a third image-text data pair for model alignment from the plurality of image-text data pairs; Encoding the sample text in the third image-text data pair using the updated initial large model to obtain text features; Encoding the sample image in the third image-text data pair using the updated initial large model to obtain image features; Based on the semantic similarity between the text features and the image features, the updated initial large model is aligned to obtain the text-image large model.
15. The method according to claim 14, wherein The determining of a third image-text data pair for model alignment from the plurality of image-text data pairs comprises: Query the resolution threshold used for model alignment; determining the third image-text data pair from the plurality of image-text data pairs according to the resolution threshold; Wherein, the resolution of the sample image in the third image-text data pair is lower than the resolution threshold.
16. A method for generating a Wensheng diagram, comprising: Get input text; Calling a large text-image model with built-in target model parameters to process the input text to obtain a target image; Wherein, the large model of the cultural graph is obtained by training using the method described in any one of claims 1-15.
17. A training device for a large-scale Wensheng graph model, comprising: an execution module, configured to sequentially execute multiple training phases on the Wensheng graph large model using multiple image-text data pairs; wherein the model parameters of the Wensheng graph large model to be trained in the i-th training phase are obtained by screening based on the evaluation indicators of multiple sets of model parameters of the Wensheng graph large model trained in the i-1-th training phase, where i is a positive integer greater than 1; A screening module is used to determine the model parameters to be fine-tuned from the multiple sets of model parameters of the Wenshengtu large model trained in the last training phase; a determining module, configured to determine a first image-text data pair from the plurality of image-text data pairs based on a first quality score and an aesthetic score of a sample image in the plurality of image-text data pairs; a fine-tuning module, configured to fine-tune the parameters of the model to be fine-tuned using the first image-text data pair, so as to obtain target model parameters corresponding to the large text-image model; The multiple training stages include a semantic tonal alignment stage, and the execution module executes the semantic tonal alignment stage, specifically: Using the text-image macro model, processing the sample text in the second image-text data pair adapted to the semantic tone alignment stage among the multiple image-text data pairs to obtain a first output image; The large Wensheng-image model is trained based on the style differences and semantic differences between the first output image and the sample images in the second image-text data pair to obtain multiple sets of model parameters of the large Wensheng-image model trained in the semantic tone alignment stage.
18. The device according to claim 17, wherein The semantic-tonal alignment stage is the first training stage among the multiple training stages.
19. The device according to claim 18, wherein The execution module is further configured to: During the training process of the large-scale Wensheng graph model, hyperparameters of the large-scale Wensheng graph model are adjusted.
20. The apparatus according to claim 17, wherein The non-first training phase among the multiple training phases includes a resolution improvement phase, and the execution module executes the resolution improvement phase, specifically: Obtaining first evaluation indicators of multiple sets of model parameters of the large Wensheng graph model trained in the previous training phase; Determining, based on the first evaluation metric, the model parameters of the Vincent graph large model to be trained in the resolution enhancement phase from the multiple sets of model parameters of the Vincent graph large model trained in the previous training phase; Dividing the plurality of image-text data pairs to obtain groups of a plurality of resolutions; wherein the sample images in the image-text data pairs in the same group have the same resolution; Sorting the multiple groups in ascending order of resolution to obtain a sorted sequence; Based on the sorting sequence, multiple resolution enhancement sub-stages in the resolution enhancement stage are sequentially executed on the large Wensheng graph model to be trained in the resolution enhancement stage.
21. The device according to claim 20, wherein The execution module executes the j-th resolution improvement sub-stage, specifically: Obtaining second evaluation indicators of multiple sets of model parameters of the Vincent image large model trained in the j-1th resolution enhancement sub-stage; where j is a positive integer greater than 1; Determining, based on the second evaluation metric, the model parameters of the Vincent image large model to be trained in the j-th resolution enhancement sub-stage from the multiple sets of model parameters of the Vincent image large model trained in the j-1-th resolution enhancement sub-stage; Using the group at the j-th position in the sorted sequence as training data for the j-th resolution improvement sub-stage; The training data of the j-th resolution enhancement sub-stage is used to adjust the model parameters of the large Vincent image model to be trained in the j-th resolution enhancement sub-stage.
22. The device according to claim 20, wherein The execution module is used to: Acquire a test data pair; wherein the test data pair includes a test text and a test image; For any set of model parameters of the Wenshengtu large model trained in the previous training phase, input the test text into the Wenshengtu large model with the any set of model parameters built in, to obtain a second output image; A first evaluation metric for any one set of model parameters is determined based on a degree of difference between the second output image and the test image.
23. The apparatus according to claim 20, wherein The execution module is used to: Get the evaluation text; For any set of model parameters of the Vincent graph large model trained in the previous training phase, input the evaluation text into the Vincent graph large model having the any set of model parameters built therein to obtain a third output image; Acquiring image quality of the third output image; wherein the image quality is used to indicate image aesthetics and / or image clarity; A first evaluation index of any one set of model parameters is determined based on the image quality of the third output image and the semantic relevance between the third output image and the evaluation text.
24. The apparatus according to claim 17, wherein The non-first training phase among the multiple training phases includes a multi-size training phase, and the execution module executes the multi-size training phase, specifically: Obtaining a third evaluation index of multiple sets of model parameters of the large Wensheng graph model trained in the previous training phase; Determining, based on the third evaluation metric, the model parameters of the Vincent graph large model to be trained in the multi-scale training phase from the multiple sets of model parameters of the Vincent graph large model trained in the previous training phase; Obtaining buckets of multiple sizes obtained by bucketing the multiple image-text data pairs; wherein the sample images in the image-text data pairs in the same bucket have the same size; Based on the bucketing of the multiple sizes, multi-size learning is performed on the large Wensheng graph model to be trained in the multi-size training phase.
25. The device according to any one of claims 17 to 24, wherein The determining module is configured to: Determining a first quality score of the sample image in any image-text data pair based on attribute information of the sample image in any image-text data pair and a semantic relevance between the sample image and the sample text in the image-text data pair; determining a candidate image-text data pair from the plurality of image-text data pairs according to first quality scores of sample images in the plurality of image-text data pairs; Determining an aesthetic score of the sample image in the candidate image-text data pair based on visual aesthetic elements in the sample image in the candidate image-text data pair; The first image-text data pair is determined from the candidate image-text data pairs based on the aesthetic scores of the sample images in the candidate image-text data pairs.
26. The device according to any one of claims 17 to 24, wherein The text-to-graph model is applicable to input text in a specified language that is adapted to the target business scenario. The multiple graph-to-text data pairs are obtained using the following modules: A first acquisition module is configured to acquire a plurality of initial data pairs, wherein the initial data pairs include sample texts in the specified language and sample images corresponding to the sample texts; The second acquisition module is used to obtain the semantic relevance between the sample text and the sample image in the same initial data pair; A third acquisition module, configured to acquire a second quality score of the sample image in the same initial data pair; The fourth acquisition module is configured to determine the plurality of image-text data pairs from the plurality of initial data pairs based on the semantic relevance and the second quality score of the plurality of initial data pairs.
27. The device according to claim 26, wherein The third acquisition module is used to: Performing image recognition on the sample image in the same initial data pair to obtain a recognition result; wherein the recognition result is used to indicate whether the sample image has at least one of black edges, mosaics, and visual illusions; Acquiring attribute information of the sample image in the same initial data pair; A second quality score of the sample image in the same initial data pair is determined according to the attribute information and the recognition result of the sample image in the same initial data pair.
28. The apparatus according to claim 26, wherein The device further comprises: The cleaning module is used to perform text cleaning on the sample texts in the plurality of image-text data pairs.
29. The device according to any one of claims 17 to 24, wherein The Wensheng graph model is obtained using the following modules: A selection module is used to determine an initial large model that is suitable for the target business scenario from multiple base large models; a replacement module, configured to replace the text encoder in the initial large model with a text encoder adapted to the specified language if the initial large model is not compatible with the specified language; wherein the specified language is a language adapted to the target business scenario; The alignment module is used to perform model alignment on the updated initial large model to obtain the Wensheng graph large model.
30. The apparatus according to claim 29, wherein The alignment module is used to: Determining a third image-text data pair for model alignment from the plurality of image-text data pairs; Encoding the sample text in the third image-text data pair using the updated initial large model to obtain text features; Encoding the sample image in the third image-text data pair using the updated initial large model to obtain image features; Based on the semantic similarity between the text features and the image features, the updated initial large model is aligned to obtain the text-image large model.
31. The device according to claim 30, wherein The alignment module is used to: Query the resolution threshold used for model alignment; determining the third image-text data pair from the plurality of image-text data pairs according to the resolution threshold; Wherein, the resolution of the sample image in the third image-text data pair is lower than the resolution threshold.
32. A Wensheng diagram device, comprising: Acquisition module, used to obtain input text; A calling module is used to call a large text-image model with built-in target model parameters to process the input text to obtain a target image; Wherein, the large model of the cultural graph is obtained by training using the device described in any one of claims 17-31.
33. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method of the Vincent graph large model according to any one of claims 1 to 15, or to execute the Vincent graph method according to claim 16.
34. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the training method for the Vincent graph large model according to any one of claims 1 to 15, or to execute the Vincent graph method according to claim 16.
35. A computer program product, comprising a computer program, wherein when executed by a processor, the computer program implements the steps of the method for training a large model of a Vincent graph according to any one of claims 1 to 15, or implements the Vincent graph method according to claim 16.
Citation Information
Patent Citations
Model training method, related system and storage medium
CN113407820A
Pre-training method and device for medical multi-modal model
CN114972929A