Visual content generation and target large model training method and device based on large model

By introducing a thinking phase into the large model and multimodal large model training, the problem of insufficient visual content generation quality was solved, and higher quality visual content and smoother interaction were achieved.

CN120893504AActive Publication Date: 2025-11-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510734116.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-11-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

Existing large models produce poor-quality results in visual content generation, failing to meet users' high-quality requirements.

Method used

A multimodal large model is adopted, introducing a thinking stage of target instruction information, and the target large model is trained through maximum likelihood estimation and reinforcement learning to improve the accuracy of the generated results.

Benefits of technology

By introducing a thinking phase and multimodal large model training, the generated visual content and response scripts are more in line with user needs, improving the accuracy of the generated results and the smoothness of the interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120893504A_ABST
    Figure CN120893504A_ABST
Patent Text Reader

Abstract

The invention provides a visual content generation and target large model training method and device based on a large model, and relates to the artificial intelligence fields of deep learning, large models, computer vision, natural language processing and the like. The visual content generation method based on the large model can comprise the steps of obtaining target instruction information; the target instruction information is input into the target large model, corresponding target result information is obtained and output, the target result information comprises the target visual content, the target result information is generated by the target large model according to the target thinking information, and the target thinking information is thinking process information generated by the target large model for the target instruction information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the fields of deep learning, large model, computer vision and natural language processing, and more particularly to a large model-based visual content generation and target large model training method and device. BACKGROUND

[0002] A large model refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of natural language text, and can simulate human language cognition and generation processes to some extent. At present, large models have been widely applied in different scenarios, such as visual content generation. Visual content generation refers to generating corresponding visual content using a large model according to user input instruction information, where the visual content can be a picture or a video. SUMMARY

[0003] The present disclosure provides a large model-based visual content generation and target large model training method and device.

[0004] A large model-based visual content generation method, comprising:

[0005] obtaining target instruction information;

[0006] inputting the target instruction information into a target large model to obtain and output corresponding target result information, wherein the target result information includes target visual content, and the target result information is generated by the target large model according to target thinking information, and the target thinking information is thinking process information generated by the target large model for the target instruction information.

[0007] A target large model training method, comprising:

[0008] obtaining a pre-trained basic large model;

[0009] obtaining first training data, wherein the first training data includes first sample instruction information, first sample result information corresponding to the first sample instruction information, and first sample thinking information, the first sample thinking information is thinking process information generated for the first sample instruction information, and the first sample result information includes first visual content;

[0010] training the basic large model according to the first training data, and determining the target large model according to a training result.

[0011] A large model-based visual content generation device, comprising an instruction obtaining module and a result generating module;

[0012] The instruction obtaining module is configured to obtain target instruction information;

[0013] The result generation module is configured to input the target instruction information into a target large model, obtain corresponding target result information, and output the target result information, wherein the target result information includes target visual content, and the target result information is generated by the target large model according to target thinking information, and the target thinking information is generated by the target large model for the target instruction information.

[0014] A target large model training apparatus includes a model acquisition module, a data acquisition module, and a model training module.

[0015] The model acquisition module is configured to acquire a pre-trained basic large model.

[0016] The data acquisition module is configured to acquire first training data, wherein the first training data includes first sample instruction information, first sample result information corresponding to the first sample instruction information, and first sample thinking information, the first sample thinking information is thinking process information generated for the first sample instruction information, and the first sample result information includes first visual content.

[0017] The model training module is configured to train the basic large model according to the first training data, and determine the target large model according to a training result.

[0018] An electronic device includes:

[0019] at least one processor; and

[0020] a memory communicatively connected to the at least one processor; wherein

[0021] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0022] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method described above.

[0023] A computer program product includes computer programs / instructions that are executed by a processor to implement the method described above.

[0024] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:

[0026] Figure 1 The flowchart of the large model-based visual content generation method embodiment of the present disclosure;

[0027] Figure 2 The interaction mode between the user and the target large model of the present disclosure is shown in the schematic diagram;

[0028] Figure 3 The first schematic diagram of the large model-based visual content generation process of the present disclosure;

[0029] Figure 4 The second schematic diagram of the large model-based visual content generation process of the present disclosure;

[0030] Figure 5 The flowchart of the first embodiment of the target large model training method of the present disclosure;

[0031] Figure 6 The flowchart of the second embodiment of the target large model training method of the present disclosure;

[0032] Figure 7 The schematic diagram of the component structure of the large model-based visual content generation device embodiment 700 of the present disclosure;

[0033] Figure 8 The schematic diagram of the component structure of the target large model training device embodiment 800 of the present disclosure;

[0034] Figure 9 A schematic block diagram of an electronic device 900 that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0035] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered only as exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.

[0036] In addition, it should be understood that the term "and / or" herein is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it.

[0037] Figure 1A flowchart of an embodiment of the method for generating visual content based on a large model according to the present disclosure is shown in FIG. 1. As shown in FIG. 1, the method includes the following specific implementations. Figure 1

[0038] In step 101, target instruction information (query) is obtained.

[0039] In step 102, the target instruction information is input into a target large model to obtain and output corresponding target result information, which includes target visual content and is generated by the target large model based on target thinking information generated by the target large model for the target instruction information.

[0040] Currently, although a large model can be used to generate visual content corresponding to instruction information input by a user, the generated visual content is usually of poor quality and cannot well meet the requirements of the user.

[0041] However, by using the scheme according to the method embodiment, a thinking stage for the target instruction information input by the user is explicitly added for the target large model, i.e., the thinking process information can be generated for the target instruction information, and then the target result information required can be generated based on the thinking process information, thereby improving the accuracy of the generated target result information and enabling the target result information to better meet the requirements of the user, etc.

[0042] In some embodiments of the present disclosure, the target large model can include a multi-modal large model, and the target instruction information can include first generation requirement description information, or the first generation requirement description information and a first picture corresponding to the first generation requirement description information.

[0043] The multi-modal large model is a model architecture that can simultaneously process input and output of multi-modal data (such as text, pictures, audio, video, etc.) and can realize cross-modal understanding and generation. Its core goal is to integrate the understanding and generation capabilities in traditional multi-modal large models through a unified framework, thereby improving the task generalization efficiency and interactive flexibility, etc. In the scheme according to the present disclosure, the multi-modal large model can be used as the target large model, thereby further improving the accuracy of the generated target result information, etc.

[0044] The target instruction information can only include the first generation requirement description information (such as corresponding to a text-to-picture task), or can include the first generation requirement description information and the first picture corresponding to the first generation requirement description information (such as corresponding to a picture editing task), i.e., the target instruction information can only include text information, or can include text information and picture information, which is very flexible and convenient.

[0045] ​Accordingly, the visual content generation process described in this disclosure can be divided into three stages: the instruction acquisition stage, the thinking stage, and the response stage. The instruction acquisition stage refers to the stage of acquiring the target instruction information input by the user; the thinking stage refers to the stage of the target large model's internal thinking process in response to the target instruction information; and the response stage refers to the stage of generating and outputting the target result information.

[0046] The target result information may include target visual content, which may be an image or a video.

[0047] In some embodiments of this disclosure, the target result information may further include: a response script that matches the target instruction information; in addition, target thinking information may also be output at the same time as the target result information is output.

[0048] In other words, while generating target visual content, it can also generate response scripts that match the target instruction information to enrich the information content returned to the user and improve the smoothness of the interaction between the user and the target large model. In addition, it can also return target thinking information to the user to further enrich the information content returned to the user.

[0049] Accordingly, Figure 2 This is a schematic diagram illustrating the interaction between the user and the target large model as described in this disclosure. Figure 2 As shown, users can input target instructions into the target model. The target model can sequentially execute the instruction acquisition stage, the thinking stage, and the response stage, and can return the target result information to the user. The target result information may include target visual content and response scripts that match the target instruction information.

[0050] in addition, Figure 3 This is a first schematic diagram of the visual content generation process based on a large model as described in this disclosure. Figure 3 As shown, assuming the target instruction information only includes the first generation requirement description information, specifically: "Draw me a futuristic clock to place on a wooden table," then the target large model can be generated using a left-to-right token-by-token generation method, such as... Figure 3The white square in the lowermost layer represents a text token, and the gray square represents a picture token. Whether to generate a text token or a picture token is determined by the target large model. For example, the target large model can first generate the text content "draw a futuristic suspended mechanical clock, cobalt blue metal...", and can generate a corresponding picture a. Specifically, the token of the picture a can be generated first, and then the picture a can be generated based on the token of the picture a by using an image decoder (Image Decoder). Further, the text content "I need a wooden table, brown wood grain with..." can be generated, and a corresponding picture b can be generated. Then, the text content "I need to put the clock on the wooden table and show it to the user" can be generated, and a corresponding picture c can be generated. The picture c is the target visual content. In addition, the text content "Hello, this is the clock and wooden table picture you requested" can also be generated as a response.

[0051] Figure 4 The second schematic diagram of the large model-based visual content generation process of the present disclosure is shown. As shown in Figure 4 the target instruction information includes the first generation requirement description information and the first picture corresponding to the first generation requirement description information. The first generation requirement description information is specifically "help me add a banana next to the apple in this picture". The first picture is the "picture" mentioned in the first generation requirement description information. The token of the first picture can be obtained by using an image encoder (Image Encoder). The target large model can first generate the text content "I need to draw a banana first", and can generate a corresponding picture a'. Then, the text content "the banana drawing is not good, redraw it" can be generated, and a corresponding picture b' can be generated. Then, the text content "I want to add the banana to the original picture" can be generated, and a corresponding picture c' can be generated. The picture c' is the target visual content. In addition, the text content "the banana has been added, do you have any other requirements" can also be generated as a response.

[0052] The target large model can be obtained by pre-training. The training process of the target large model is described below.

[0053] Figure 5 The flowchart of the first embodiment of the target large model training method of the present disclosure is shown. As shown in Figure 5 the following specific implementation methods are included.

[0054] In step 501, a pre-trained basic large model is obtained.

[0055] In step 502, first training data is obtained, the first training data including: first sample instruction information, first sample result information corresponding to the first sample instruction information, and first sample thinking information, the first sample thinking information being thinking process information generated for the first sample instruction information, and the first sample result information including first visual content.

[0056] In step 503, the base large model is trained according to the first training data, and a target large model is determined according to a training result.

[0057] Based on the above training data, the target large model can learn how to generate thinking process information, so as to generate target result information corresponding to target instruction information input by a user based on the thinking process information, thereby improving the accuracy of the generated target result information, and enabling the target result information to better meet user requirements, etc.

[0058] In some embodiments of the present disclosure, the target large model can include a multi-modal large model. Accordingly, the base large model can be a multi-modal large model, such as a pre-trained multi-modal large model, such as a Chameleon model, which can be directly reused, thereby improving training efficiency, and the powerful reasoning capability of the multi-modal large model can be used to improve the accuracy of the obtained target result information, etc.

[0059] How to obtain the first training data is not limited, for example, it can be manually collected and labeled. The first training data can include: first sample instruction information, first sample result information, and first sample thinking information, the first sample thinking information being thinking process information generated for the first sample instruction information, and the first sample result information including first visual content.

[0060] That is, the first training data can include: <query> ……< / query> <thinking> ……< / thinking> <response> ……< / response> , wherein, <query> ……< / query> represents the first sample instruction information, <thinking> ……< / thinking> <represents the first sample thinking information, <response> ……< / response> represents the first sample result information.

[0061] In some embodiments of the present disclosure, the first sample instruction information can include: second generation requirement description information, or the second generation requirement description information and a second picture corresponding to the second generation requirement description information. Correspondingly, any first sample thinking information can include one of the following: 1) refined requirement description information obtained by refining the second generation requirement description information; 2) step description information for generating the first sample result information; 3) initial result information and optimization description information, the first sample result information being obtained by performing optimization processing corresponding to the optimization description information on the initial result information; 4) M candidate result information corresponding to the first sample instruction information and selection reason information, M being a positive integer greater than 1, the M candidate result information including the first sample result information, and the selection reason information being used to explain reasons why the first sample result information is better than other candidate result information.

[0062] That is, at least the above four ways can be used to generate the thinking process information. The four ways are further described below by taking a picture as the first visual content.

[0063] In the way 1), the first sample instruction information can be enriched and rewritten in the form of text, that is, the second generation requirement description information is refined to obtain refined requirement description information. Compared with the second generation requirement description information, the refined requirement description information can perfect details, styles, layouts, etc. inside the picture to be generated.

[0064] In the way 2), the second generation requirement description information can be disassembled to obtain step description information for generating the first sample result information, such as which text content is generated first, which picture is generated next, and how to finally combine to obtain the final required picture.

[0065] In the way 3), the generated picture content can be repeatedly modified. For example, the initial result information and the optimization description information (text thinking information) can be given, and the picture in the first sample result information can be obtained by performing optimization processing corresponding to the optimization description information on the picture in the initial result information, that is, the picture in the first sample result information can be obtained by performing detail modification on the picture in the initial result information according to the optimization description information.

[0066] In the way 4), M candidate result information corresponding to the first sample instruction information can be given, M being a positive integer greater than 1, and the specific value can be determined according to actual needs. The selection reason information can also be given, the M candidate result information including the first sample result information, and the selection reason information being used to explain reasons why the first sample result information is better than other candidate result information, that is, the selection reason information is used to explain why the first sample result information is selected as the final required result from the M candidate result information.

[0067] It can be seen that, through the above processing, the target large model can learn various different thinking modes, thereby improving the learning effect of the target large model, that is, improving the performance of the target large model. Subsequently, when the target large model is used for actual reasoning application, the target large model can determine the specific thinking process information by itself.

[0068] In some embodiments of the present disclosure, when the base large model is trained according to the first training data, the base large model can be trained by maximum likelihood estimation according to the first training data.

[0069] Maximum likelihood estimation is a mature training method. Accordingly, the base large model can be trained by maximum likelihood estimation according to the process of the first sample instruction information, the first sample thinking information and the first sample result information, thereby improving the training efficiency and learning effect of the target large model.

[0070] After the base large model is trained according to the first training data, the target large model can be determined according to the training result.

[0071] In some embodiments of the present disclosure, after the base large model is trained according to the first training data, an intermediate large model can be obtained. Then, the intermediate large model can be directly determined as the target large model, or second training data can be obtained, the second training data can include second sample instruction information, and the intermediate large model can be trained by reinforcement learning according to the second training data, thereby obtaining the target large model.

[0072] Training the base large model according to the first training data means supervised fine-tuning (SFT) training of the base large model. The base large model is pre-trained, and then combined with pre-training and fine-tuning to obtain the target large model required, or to further improve the performance of the target large model, the intermediate large model can be trained by reinforcement learning using the second training data after obtaining the intermediate large model.

[0073] Specifically, the reinforcement learning can use a human feedback reinforcement learning (RLHF) algorithm and the like.

[0074] In some embodiments of the present disclosure, the reinforcement learning training method of the intermediate large model according to the second training data can include: inputting the second sample instruction information into the intermediate large model to obtain the output intermediate result information, the intermediate result information including the second visual content, determining the comprehensive evaluation result according to the intermediate result information and the second sample instruction information, and updating the intermediate large model according to the principle of improving the comprehensive evaluation result.

[0075] The comprehensive evaluation result can be a comprehensive score, i.e., a reward model comprehensive score, and an optimization goal of reinforcement learning is to improve the comprehensive score of the output result. Accordingly, after the comprehensive score is determined according to the intermediate result information and the second sample instruction information, the intermediate large model can be updated (i.e., optimized) according to the principle of improving the comprehensive score.

[0076] In some embodiments of the present disclosure, the second sample instruction information can include third generation requirement description information, or the third generation requirement description information and a third picture corresponding to the third generation requirement description information. Accordingly, in response to the second visual content being a picture, the manner of determining the comprehensive evaluation result according to the intermediate result information and the second sample instruction information can include obtaining a similarity score between the second visual content and the third generation requirement description information, and obtaining an aesthetic score of the second visual content. In response to determining that the second sample instruction information does not include the third picture, the comprehensive score is determined according to the similarity score and the aesthetic score. In response to determining that the second sample instruction information includes the third picture, a sum of squares of differences between corresponding pixel points in the second visual content and the third picture is obtained, the corresponding pixel points are pixel points with the same coordinate position, and the comprehensive score is determined according to the similarity score, the aesthetic score, and the sum of squares.

[0077] That is, in the scheme described in the present disclosure, a multi-objective reinforcement learning manner including model scoring and rule calculation can be adopted. The model scoring refers to the similarity score and the aesthetic score described above. For example, a pre-trained text-image similarity model can be used to determine the similarity score between the second visual content and the third generation requirement description information, and a pre-trained picture aesthetic evaluation model can be used to determine the aesthetic score of the second visual content. The rule calculation refers to calculating a sum of squares of differences between corresponding pixel points in the second visual content and the third picture (differences between individual pixel points). The similarity score is used to reflect the degree to which the target large model complies with the user instruction. The higher the similarity score, the higher the degree to which the target large model complies with the user instruction. The aesthetic score is used to reflect the aesthetic degree of the generated second visual content. The higher the aesthetic score, the higher the aesthetic degree of the second visual content. The sum of squares is used to reflect whether the original picture is followed during the picture editing process. The larger the value of the sum of squares, the higher the degree of compliance. Accordingly, the accuracy of the comprehensive score obtained can be improved by determining the comprehensive score in combination with the similarity score, the aesthetic score, and the sum of squares, and the optimization efficiency of the intermediate large model can be improved.

[0078] How to determine the comprehensive score in combination with the similarity score, the aesthetic score, and the sum of squares is not limited. For example, the comprehensive score can be calculated according to a predetermined calculation formula.

[0079] In combination with the above description, Figure 6A flowchart of a second embodiment of the target large model training method according to the present disclosure. As shown in Figure 6 The following specific implementations are included.

[0080] In step 601, a pre-trained base large model is obtained.

[0081] The base large model can be a multi-modal large model.

[0082] In step 602, first training data is obtained, which includes first sample instruction information, first sample result information corresponding to the first sample instruction information, and first sample thinking information. The first sample thinking information is the thinking process information generated for the first sample instruction information, and the first sample result information includes first visual content.

[0083] The first sample instruction information can include second generation requirement description information, or second generation requirement description information and a second picture corresponding to the second generation requirement description information.

[0084] In addition, any first sample thinking information can include one of the following: refined requirement description information obtained by refining the second generation requirement description information; step description information for generating the first sample result information; initial result information and optimization description information, the first sample result information being obtained by performing optimization processing corresponding to the optimization description information on the initial result information; M candidate result information corresponding to the first sample instruction information and selection reason information, M being a positive integer greater than 1, the M candidate result information including the first sample result information, and the selection reason information being used to explain the reason why the first sample result information is better than other candidate result information.

[0085] In step 603, the base large model is trained according to the first training data to obtain an intermediate large model.

[0086] For example, the base large model can be trained using maximum likelihood estimation in an autoregressive manner according to the first training data.

[0087] In step 604, second training data is obtained, which includes second sample instruction information.

[0088] The second sample instruction information includes third generation requirement description information, or third generation requirement description information and a third picture corresponding to the third generation requirement description information.

[0089] In step 605, the intermediate large model is trained using reinforcement learning according to the second training data to obtain a target large model.

[0090] For example, the second sample instruction information can be input into the intermediate large model to obtain output intermediate result information, the second visual content is included in the intermediate result information, and then the comprehensive evaluation result can be determined according to the intermediate result information and the second sample instruction information, and the intermediate large model can be updated according to the principle of improving the comprehensive evaluation result.

[0091] After obtaining the target large model, it can be applied to actual inference application, such as can be applied to Figure 1 the visual content generation method shown in the figure, for generating corresponding target result information based on input target instruction information.

[0092] In addition, in the inference application process, after generating the target result information by using the target large model, the target large model can be further trained by reinforcement learning by using the target result information and the corresponding target instruction information, so as to further improve the performance of the target large model.

[0093] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the disclosure is not limited by the described action sequence, because according to the disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the disclosure. In addition, the parts not described in detail in a certain embodiment can refer to the related description in other embodiments.

[0094] The above is the introduction of the method embodiment, and the scheme described in the disclosure will be further described through the device embodiment.

[0095] Figure 7 The constituent structure diagram of the large model based visual content generation device embodiment 700 of the disclosure is shown in the figure. Figure 7 As shown, it includes an instruction acquisition module 701 and a result generation module 702.

[0096] The instruction acquisition module 701 is configured to acquire target instruction information.

[0097] The result generation module 702 is configured to input the target instruction information into the target large model to obtain and output corresponding target result information, the target visual content is included in the target result information, the target result information is generated by the target large model according to the target thinking information, and the target thinking information is the thinking process information generated by the target large model for the target instruction information.

[0098] In some embodiments of the present disclosure, the target large model can include a multi-modal large model, and the target instruction information can include first generation requirement description information, or the first generation requirement description information and a first picture corresponding to the first generation requirement description information.

[0099] In some embodiments of the present disclosure, the target result information can further include a response script matched with the target instruction information, and / or the result generation module 702 can output the target thinking information while outputting the target result information.

[0100] Figure 8 A schematic diagram of a constituent structure of the target large model training apparatus embodiment 800 of the present disclosure is shown in FIG. 8. As shown in FIG. 8, the target large model training apparatus embodiment 800 of the present disclosure includes a model acquisition module 801, a data acquisition module 802, and a model training module 803. Figure 8

[0101] The model acquisition module 801 is configured to acquire a pre-trained basic large model.

[0102] The data acquisition module 802 is configured to acquire first training data, and the first training data includes first sample instruction information, first sample result information corresponding to the first sample instruction information, and first sample thinking information, the first sample thinking information being thinking process information generated for the first sample instruction information, and the first sample result information including first visual content.

[0103] The model training module 803 is configured to train the basic large model according to the first training data, and determine a target large model according to a training result.

[0104] In some embodiments of the present disclosure, the target large model can include a multi-modal large model, and the first sample instruction information can include second generation requirement description information, or the second generation requirement description information and a second picture corresponding to the second generation requirement description information.

[0105] In some embodiments of the present disclosure, any first sample thinking information can include one of the following: refined requirement description information obtained by refining the second generation requirement description information; step description information of generating the first sample result information; initial result information and optimization description information, the first sample result information being obtained by performing optimization processing corresponding to the optimization description information on the initial result information; M candidate result information corresponding to the first sample instruction information and selection reason information, M being a positive integer greater than 1, the M candidate result information including the first sample result information, and the selection reason information being used to explain reasons why the first sample result information is better than other candidate result information.

[0106] ​In some embodiments of the present disclosure, the model training module 803 can perform autoregressive training on the base large model according to the first training data in a maximum likelihood estimation manner when training the base large model according to the first training data.

[0107] In some embodiments of the present disclosure, the model training module 803 can obtain an intermediate large model after training the base large model according to the first training data, and then directly determine the intermediate large model as the target large model, or obtain second training data, which can include second sample instruction information, and perform reinforcement learning training on the intermediate large model according to the second training data to obtain the target large model.

[0108] In some embodiments of the present disclosure, the model training module 803 can perform reinforcement learning training on the intermediate large model according to the second training data in the following manner: input the second sample instruction information into the intermediate large model to obtain output intermediate result information, the intermediate result information including the second visual content, determine a comprehensive evaluation result according to the intermediate result information and the second sample instruction information, and update the intermediate large model according to the principle of improving the comprehensive evaluation result.

[0109] In some embodiments of the present disclosure, the second sample instruction information can include third generation requirement description information, or the third generation requirement description information and a third picture corresponding to the third generation requirement description information, and the comprehensive evaluation result can include a comprehensive score. Accordingly, in response to the second visual content being a picture, the model training module 803 can determine the comprehensive evaluation result according to the intermediate result information and the second sample instruction information in the following manner: obtain a similarity score between the second visual content and the third generation requirement description information, and obtain an aesthetic score of the second visual content, determine the comprehensive score according to the similarity score and the aesthetic score in response to determining that the second sample instruction information does not include the third picture, and obtain a sum of squares of differences between corresponding pixel points in the second visual content and the third picture, the corresponding pixel points being pixel points with the same coordinate position, and determine the comprehensive score according to the similarity score, the aesthetic score and the sum of squares in response to determining that the second sample instruction information includes the third picture.

[0110] The specific working processes of the above-mentioned device embodiments can refer to the related descriptions in the foregoing method embodiments, which will not be described herein again.

[0111] In summary, by using the thinking chain technology of the multi-modal large model, the accuracy of the visual content generation result can be improved, and the method is applicable to different visual content generation scenarios and has wide applicability.

[0112] The scheme described in the present disclosure can be applied to the field of artificial intelligence, and in particular relates to the fields of deep learning, large models, computer vision, and natural language processing. Artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) of humans, and includes both hardware and software technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, special-purpose artificial intelligence chips, cloud computing, distributed storage, big data processing, etc., and artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc.

[0113] In addition, the instruction information and result information in the embodiments of the present disclosure are not for a specific user and cannot reflect the personal information of a specific user. In the technical scheme of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0114] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0115] Figure 9 A schematic block diagram of an electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, servers, servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections, and their functions, as described herein, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0116] As shown in Figure 9 The electronic device 900 includes a computing unit 901 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 902 or a computer program loaded into a random access memory (RAM) 903 from a storage unit 908. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0117] A plurality of components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0118] The computing unit 901 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphic processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded to the RAM 903 and executed by the computing unit 901, one or more steps of the methods described in the present disclosure can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the methods described in the present disclosure by other any appropriate means, such as by means of firmware.

[0119] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0120] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as part of a separate software package, and partially on a remote machine or server.

[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a linearly-programmed electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0122] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0123] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0124] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0125] It should be understood that the various forms of flow shown above can be re-ordered, added to, or have steps deleted, using the steps shown above. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.

[0126] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A visual content generation method based on a large model, comprising: Obtain target instruction information; The target instruction information is input into the target large model to obtain the corresponding target result information and output it. The target result information includes target visual content. The target result information is generated by the target large model based on the target thinking information. The target thinking information is the thinking process information generated by the target large model in response to the target instruction information.

2. The method according to claim 1, wherein, The target large model includes: a multimodal large model; The target instruction information includes: first generation requirement description information, or the first generation requirement description information and a first image corresponding to the first generation requirement description information.

3. The method according to claim 1, wherein, The target result information also includes: a response script that matches the target instruction information; And / or, the method further includes: outputting the target thinking information while outputting the target result information.

4. A method for training a large target model, comprising: Obtain the pre-trained base model; Acquire first training data, which includes: first sample instruction information, first sample result information corresponding to the first sample instruction information, and first sample thinking information. The first sample thinking information is thinking process information generated in response to the first sample instruction information, and the first sample result information includes first visual content. The basic large model is trained based on the first training data, and the target large model is determined based on the training results.

5. The method according to claim 4, wherein, The target large model includes: a multimodal large model; The first sample instruction information includes: second generation requirement description information, or the second generation requirement description information and a second image corresponding to the second generation requirement description information.

6. The method according to claim 5, wherein, Each of the first sample's thinking information includes one of the following: The refined requirement description information is obtained by refining the second generated requirement description information; Description of the steps for generating the first sample result information; Initial result information and optimized description information, wherein the first sample result information is obtained by performing optimization processing corresponding to the optimized description information on the initial result information; The first sample instruction information includes M candidate result information and selection reason information, where M is a positive integer greater than 1. The M candidate result information includes the first sample result information, and the selection reason information is used to explain why the first sample result information is superior to other candidate result information.

7. The method according to claim 4, wherein, The step of training the basic large model based on the first training data includes: Based on the first training data, the basic large model is trained using the maximum likelihood estimation method through autoregression.

8. The method according to claim 4, wherein, The step of training the basic large model based on the first training data and determining the target large model based on the training results includes: The basic large model is trained based on the first training data to obtain the intermediate large model; The intermediate large model is determined as the target large model, or second training data is obtained, the second training data including second sample instruction information, and reinforcement learning training is performed on the intermediate large model based on the second training data to obtain the target large model.

9. The method according to claim 8, wherein, The step of training the intermediate large model using reinforcement learning based on the second training data includes: The second sample instruction information is input into the intermediate large model to obtain the output intermediate result information, which includes the second visual content; The comprehensive evaluation result is determined based on the intermediate result information and the second sample instruction information; The intermediate large model is updated in accordance with the principle of improving the overall evaluation results.

10. The method according to claim 9, wherein, The second sample instruction information includes: third generation requirement description information, or the third generation requirement description information and the third image corresponding to the third generation requirement description information; The comprehensive evaluation results include: a comprehensive score; In response to determining that the second visual content is an image, determining the comprehensive evaluation result based on the intermediate result information and the second sample instruction information includes: Obtain a similarity score between the second visual content and the third generated requirement description information, and obtain an aesthetic score for the second visual content; In response to determining that the third image is not included in the second sample instruction information, the comprehensive score is determined based on the similarity score and the aesthetic score; In response to determining that the second sample instruction information includes the third image, the sum of squares of the differences between the second visual content and the corresponding pixels in the third image is obtained, wherein the corresponding pixels are pixels with the same coordinate position, and the comprehensive score is determined based on the similarity score, the aesthetic score and the sum of squares.

11. A visual content generation device based on a large model, comprising: Instruction acquisition module and result generation module; The instruction acquisition module is used to acquire target instruction information; The result generation module is used to input the target instruction information into the target large model, obtain the corresponding target result information and output it. The target result information includes target visual content. The target result information is generated by the target large model based on the target thinking information. The target thinking information is the thinking process information generated by the target large model in response to the target instruction information.

12. A target large model training device, comprising: Model acquisition module, data acquisition module, and model training module; The model acquisition module is used to acquire the pre-trained basic large model; The data acquisition module is used to acquire first training data, which includes: first sample instruction information, first sample result information corresponding to the first sample instruction information, and first sample thinking information. The first sample thinking information is thinking process information generated in response to the first sample instruction information, and the first sample result information includes first visual content. The model training module is used to train the basic large model based on the first training data, and determine the target large model based on the training results.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

15. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Text-to-image generation model optimization method and device, equipment and storage medium

    CN116611496A

  • Image generation method and data processing method for image generation

    CN117409109A

  • Information processing method, device and equipment and computer readable storage medium

    CN117494722A

  • Data processing method, text and graph generation method and related devices

    CN117671055A

  • Private knowledge visual content generation method and device based on general recognition large model

    CN117786473A