Method and apparatus for generating visual content based on a big model and training a target big model.

By integrating a thinking phase and advanced training methods, the visual content generation method addresses the issue of low-quality outputs, enhancing the accuracy and user satisfaction of generated content through a multimodal large model.

JP2026071280APending Publication Date: 2026-04-28BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2026-01-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing visual content generation methods using large models often produce low-quality results that fail to meet user requirements due to a lack of explicit consideration of the thinking process behind the generated content.

Method used

A method and apparatus that incorporates a thinking phase into the visual content generation process, utilizing a multimodal large model to generate visual content based on user instructions, and a training process that includes pre-training, fine-tuning, and reinforcement learning to enhance the accuracy and flexibility of the model.

Benefits of technology

Improves the quality and accuracy of generated visual content by explicitly considering the thinking process, enabling the model to better meet user requirements and enhance interaction smoothness through enriched information feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071280000001_ABST
    Figure 2026071280000001_ABST
Patent Text Reader

Abstract

This invention provides a method for generating visual content based on a big model, which improves the accuracy of target result information and enables the target result information to meet user requirements, as well as a method for training a target big model. [Solution] A method for generating visual content based on a big model includes the steps of: acquiring target instruction information; inputting the target instruction information into a target big model; and obtaining and outputting corresponding target result information, which includes target visual content, and is generated based on target thinking information, which is thought process information generated by the target big model in relation to the target instruction information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to fields such as deep learning, large models, computer vision, and natural language processing. In particular, it relates to a method and apparatus for visual content generation based on a large model and training of a target large model.

Background Art

[0002] A large model refers to a deep learning model trained using a large amount of text data, which can generate natural language text, understand the meaning of natural language text, and can simulate the processes of human language cognition and generation to a certain extent. Currently, large models are widely applied in various scenarios such as visual content generation. Visual content generation means using a large model to generate corresponding visual content such as pictures and videos based on the instruction information input by the user.

Summary of the Invention

Problems to be Solved by the Invention

[0003] The present disclosure solves the technical problems related to a method and apparatus for visual content generation based on a large model and training of a target large model.

Means for Solving the Problems

[0004] The present disclosure provides a method and apparatus for visual content generation based on a large model and training of a target large model.

[0005] The method for generating visual content based on a large model includes the steps of obtaining target instruction information, inputting the target instruction information into a target large model, and obtaining and outputting corresponding target result information including target visual content, wherein the target result information is generated based on target thinking information, which is the thinking process information generated by the target large model for the target instruction information.

[0006] A method for training a target big model includes the steps of: acquiring a base big model after pre-training; acquiring first training data which includes first sample instruction information, first sample result information including first visual content corresponding to the first sample instruction information, and first sample thinking information which is thought process information generated for the first sample instruction information; and training the base big model based on the first training data and determining the target big model based on the training results.

[0007] A visual content generation device based on a big model comprises an instruction acquisition module and a result generation module. The instruction acquisition module acquires target instruction information, and the result generation module inputs the target instruction information into a target big model and outputs corresponding target result information, which includes target visual content, and is generated based on target thinking information, which is thought process information generated by the target big model in response to the target instruction information.

[0008] The training device for the target big model comprises a model acquisition module, a data acquisition module, and a model training module. The model acquisition module acquires a pre-trained base big model. The data acquisition module acquires first training data, which includes first sample instruction information, first sample result information including first visual content corresponding to the first sample instruction information, and first sample thinking information which is thought process information generated for the first sample instruction information. The model training module trains the base big model based on the first training data and determines the target big model based on the training results.

[0009] The electronic device comprises at least one processor and a memory communicated with the at least one processor, wherein the memory stores instructions that can be executed by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is able to execute the method.

[0010] A non-temporary, computer-readable storage medium on which computer instructions are stored, wherein the computer instructions cause a computer to execute the method.

[0011] This is a computer program product that, when executed by a processor, includes a computer program / instruction that implements the aforementioned method.

[0012] It should be understood that the content described in this section neither marks any important or essential features of the embodiments of this disclosure nor limits the scope of this disclosure. Other features of this disclosure will be readily apparent from the following specification. [Brief explanation of the drawing]

[0013] The attached drawings are for the purpose of better understanding this embodiment and do not limit the disclosure. [Figure 1] This is a flowchart of an embodiment of a method for generating visual content based on a big model related to this disclosure. [Figure 2] This diagram illustrates the interaction method between the user and the target big model related to this disclosure. [Figure 3] This is the first schematic diagram of the visual content generation process based on the big model related to this disclosure. [Figure 4] This is the second schematic diagram of the visual content generation process based on the big model related to this disclosure. [Figure 5] This is a flowchart of the first embodiment of the training method for the target big model related to this disclosure. [Figure 6]This is a flowchart of the second embodiment of the training method for the target big model related to this disclosure. [Figure 7] This is a schematic diagram of the configuration of Embodiment 700 of a visual content generation device based on a big model relating to this disclosure. [Figure 8] This is a schematic diagram of the configuration of Embodiment 800 of the training device for the target big model related to this disclosure. [Figure 9] This is a schematic block diagram of an electronic device 900 that may be used to carry out embodiments of the present disclosure. [Modes for carrying out the invention]

[0014] Illustrative embodiments of this application will be described below with reference to the drawings. For ease of understanding, various details of the embodiments of this application are included and should be considered as illustrative only. Accordingly, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for the sake of brevity, well-known functions and structures will be omitted in the following description.

[0015] Furthermore, the term "and / or" in this specification simply describes a relationship between related objects, meaning that three such relationships may exist. For example, A and / or B can mean three situations: A exists alone, A and B exist together, and B exists alone. Also, the letter " / " in this specification generally means that the preceding and succeeding related objects are in an "or" relationship.

[0016] FIG. 1 is a flowchart of an embodiment of a method for generating visual content based on a large model according to the present disclosure. As shown in FIG. 1, the following specific embodiments are included. In step 101, target instruction information (query) is acquired. In step 102, the target instruction information is input into a target large model, and corresponding target result information including target visual content is obtained and output, where the target result information is generated based on target thinking (thinking) information, which is the thinking process information generated by the target large model for the target instruction information.

[0017] Currently, it is possible to generate visual content corresponding to the instruction information input by a user using a large model. However, the generated visual content generally has poor quality and cannot fully meet the user's requirements.

[0018] In the above method embodiment, for the target large model, the thinking stage for the target instruction information input by the user is explicitly added. That is, first, thinking process information is generated for the target instruction information, and further, desired target result information can be generated based on the thinking process information. Therefore, the accuracy of the generated target result information is improved, and the target result information can better meet the user's requirements.

[0019] In some embodiments of the present disclosure, the target large model includes a multimodal large model, and the target instruction information may include first generation requirement description information, or first generation requirement description information and a first picture corresponding to the first generation requirement description information.

[0020] A multimodal large model is a model structure that can simultaneously process the input and output of multimodal data (such as text, pictures, audio, video, etc.) and realize cross-modal understanding and generation. Its core goal is to integrate the understanding and generation capabilities in traditional multimodal large models through a unified framework, thereby improving the generalization efficiency of tasks and the flexibility of interactions, etc. In the method of the present disclosure, a multimodal large model can be adopted as the target large model, so the accuracy of the generated target result information can be further improved.

[0021] The target instruction information may only include the first generation requirement description information (for example, corresponding to the task of generating pictures from text), or may simultaneously include the first generation requirement description information and the first picture corresponding to the first generation requirement description information (for example, corresponding to the picture editing task). That is, the target instruction information may only include text information, or may simultaneously include text information and picture information, so it is very flexible and convenient.

[0022] Accordingly, the visual content generation process according to the present disclosure can be divided into three phases: an instruction acquisition phase, a thinking phase, and a response phase. The instruction acquisition phase is the phase of acquiring the target instruction information input by the user, the thinking phase is the internal thinking process phase of the target large model for the target instruction information, and the response phase is the phase of generating and outputting the target result information.

[0023] The target result information may include the target visual content, and the target visual content may be a picture or a video.

[0024] In some embodiments of the present disclosure, the target result information may further include a response utterance expression adapted to the target instruction information. Also, target thinking information may be output simultaneously with the target result information.

[0025] In other words, by generating target visual content and corresponding response utterances that match the target instruction information, the information returned to the user can be enriched, improving the smoothness of interaction between the user and the target big model. Furthermore, by returning target thinking information to the user, the information returned to the user can be further enriched.

[0026] Accordingly, Figure 2 is a schematic diagram of the interaction method between the user and the target big model related to this disclosure. As shown in Figure 2, the user inputs target instruction information into the target big model, and the target big model sequentially executes the instruction acquisition stage, thinking stage, and response stage, and can return target result information to the user, including target visual content and response utterances that match the target instruction information.

[0027] Figure 3 is a first schematic diagram of the visual content generation process based on the big model related to this disclosure. As shown in Figure 3, assuming that the target instruction information includes only the first generation requirement description information, and that the first generation requirement description information specifically states "Draw a picture of a technological clock placed on a wooden table," the target big model can adopt a token generation method from left to right. As shown in the bottom layer of Figure 3, the white squares represent text tokens, and the gray squares represent picture tokens. Whether to generate text tokens or picture tokens is determined by the target big model itself. For example, the target big model may first generate text content such as "Draw a futuristic hanging mechanical clock made of cobalt blue metal," and then generate the corresponding picture a. Specifically, it may first generate a token for picture a, and then generate picture a using an image decoder based on the token for picture a. Furthermore, the text content "We need a wooden table with a brown wood grain pattern" may be generated, and a corresponding picture b may be generated. Next, the text content "We need to place a clock on the wooden table and show it to the user" may be generated, and a corresponding picture c may be generated. Picture c is the target visual content. In addition, as a response utterance, the text content "Hello, this is a diagram of the clock and wooden table you requested" can be generated simultaneously.

[0028] Figure 4 is a second schematic diagram of the visual content generation process based on the big model relating to this disclosure. As shown in Figure 4, the target instruction information simultaneously includes first generation requirement description information and a first picture corresponding to the first generation requirement description information. Assuming that the first generation requirement description information is specifically "Add a banana next to the apple in this picture," the first picture is "this picture" mentioned in the first generation requirement description information. The token of the first picture can be obtained using an image encoder. The target big model may first generate text content that says "We need to draw a banana first," and generate a corresponding picture a'. Next, it may generate text content that says "The banana isn't very good, so draw it again," and generate a corresponding picture b'. Next, it may generate text content that says "I want to add a banana to the original picture," and generate a corresponding picture c'. Picture c' is the target visual content. It may also simultaneously generate text content as a response utterance, such as "A banana has been added. Do you have any other requests?"

[0029] The target big model is obtained through pre-training, and the training process for the target big model is described below.

[0030] Figure 5 is a flowchart of a first embodiment of the method for training a target big model according to this disclosure. As shown in Figure 5, the following specific embodiments are included. In step 501, a pre-trained base big model is acquired. In step 502, first training data is acquired, which includes first sample instruction information, first sample result information including first visual content corresponding to the first sample instruction information, and first sample thinking information which is thought process information generated for the first sample instruction information. In step 503, the base big model is trained based on the first training data, and a target big model is determined based on the training results.

[0031] Based on the aforementioned training data, the goal big model can be trained to generate thought process information so that it generates goal result information corresponding to the goal instruction information entered by the user based on the thought process information. This improves the accuracy of the generated goal result information and allows the goal result information to better meet the user's requirements.

[0032] In some embodiments of this disclosure, the target big model may include a multimodal big model. Accordingly, the base big model may be a multimodal big model, and for example, an existing pre-trained multimodal big model, such as a mixed modal base model (Chameleon), can be used directly, thereby improving training efficiency and enhancing the accuracy of the obtained target result information due to the powerful inference capabilities of the multimodal big model.

[0033] There are no restrictions on how the first training data is obtained; for example, it can be obtained by manually collecting and labeling it. The first training data may include first sample instruction information, first sample result information, and first sample thought information, where the first sample thought information is the thought process information generated in response to the first sample instruction information. The first sample result information includes first visual content.

[0034] In other words, the first training data includes: <query> …< / query> <thinking> …< / thinking> <response> …< / response> It may include, <query> …< / query> This represents the first sample instruction information, <thinking> …< / thinking> < represents the first sample thought information, <response> …< / response> This represents the results information for the first sample.

[0035] In some embodiments of this disclosure, the first sample instruction information may include second generation requirement description information, or second generation requirement description information and a second picture corresponding to the second generation requirement description information. Accordingly, any one of the first sample thought information may include any one of the following: 1) subdivided requirement description information obtained by subdividing the second generation requirement description information, 2) step description information for generating the first sample result information, 3) initial result information and optimization description information, and 4) M candidate result information and selection reason information corresponding to the first sample instruction information. Here, the first sample result information is obtained by performing optimization processing corresponding to the optimization description information on the initial result information, M is a positive integer greater than 1, the M candidate result information includes the first sample result information, and the selection reason information explains why the first sample result information is superior to the other candidate result information.

[0036] In other words, thought process information can be generated using at least the four methods described above. Below, these four methods will be further explained, with the first visual content being a picture.

[0037] In method 1), the first sample instruction information is enriched and rewritten in text format, that is, the second generation requirement description information is subdivided to obtain subdivided requirement description information. Compared to the second generation requirement description information, the subdivided requirement description information can improve the details, style, layout, etc., of the generated picture.

[0038] Method 2) involves decomposing the second generation requirements description information to obtain step description information for generating the first sample result information, such as which text content to generate first, which pictures to generate next, and how to combine them to obtain the final desired picture.

[0039] In method 3), the content of the generated picture can be repeatedly modified by providing initial result information and optimization description information (textual correction information). The picture in the first sample result information is obtained by performing optimization processing corresponding to the optimization description information on the picture in the initial result information. In other words, the picture in the first sample result information can be obtained by modifying the picture in the initial result information in detail using the optimization description information.

[0040] In method 4), M candidate result information corresponding to the first sample instruction information can be provided simultaneously. M is a positive integer greater than 1, and its specific value can be set according to actual needs. Furthermore, selection reason information may be provided. The M candidate result information includes the first sample result information. The selection reason information explains why the first sample result information is superior to the other candidate result information. That is, the selection reason information explains why the first sample result information is ultimately selected as the desired result from among the M candidate result information.

[0041] In this way, the above process allows the target big model to learn various different thinking methods, thereby improving the learning effect of the target big model, i.e., the performance of the target big model. Subsequently, when using the target big model for actual inference applications, the target big model can independently determine specific thought process information.

[0042] In some embodiments of this disclosure, when training the underlying big model on the first training data, the underlying big model may be autoregressively trained on the first training data using a maximum likelihood estimation method.

[0043] Maximum likelihood estimation is a mature training method. Accordingly, by adopting the maximum likelihood estimation method and performing autoregressive training on the base big model in the order of first sample instruction information, first sample thought information, and first sample result information, it is possible to improve the training efficiency and learning effect of the target big model.

[0044] After training a base big model based on the first training data, a target big model can be determined based on the training results.

[0045] In some embodiments of this disclosure, a base big model may be trained based on first training data, an intermediate big model may be obtained, and then the intermediate big model may be directly determined as the target big model; or a second training data may be obtained which may include second sample instruction information, and the intermediate big model may be reinforced learning trained based on the second training data to obtain the target big model.

[0046] Training a base big model based on the first training data refers to performing supervised fine-tuning (SFT) training on the base big model. Since the base big model was obtained through pre-training, the desired target big model can be obtained by combining pre-training and fine-tuning. Alternatively, to further improve the performance of the target big model, reinforcement learning training can be performed using the second training data after obtaining an intermediate big model.

[0047] Specifically, the aforementioned reinforcement learning can employ algorithms such as Human Feedback Reinforcement Learning (RLHF).

[0048] In embodiments of the present disclosure, a method for training an intermediate big model using reinforcement learning based on second training data may include inputting second sample instruction information into the intermediate big model, obtaining intermediate result information including the output second visual content, determining an overall evaluation result based on the intermediate result information and the second sample instruction information, and updating the intermediate big model in accordance with the principle of improving the overall evaluation result.

[0049] The overall evaluation result can refer to the overall score, i.e., the overall score of the reinforcement model. The optimization goal of reinforcement learning is to improve the overall score of the output result. Accordingly, after determining the overall score based on the intermediate result information and the second sample instruction information, the intermediate big model can be updated (i.e., optimized) according to the principle of improving the overall score.

[0050] In some embodiments of this disclosure, the second sample instruction information may include third generation requirement description information, or third generation requirement description information and a third picture corresponding to the third generation requirement description information. Accordingly, depending on whether the second visual content is a picture, the method for determining the overall evaluation result from the intermediate result information and the second sample instruction information may include obtaining a similarity score between the second visual content and the third generation requirement description information, obtaining an aesthetic score for the second visual content, determining an overall score based on the similarity score and aesthetic score depending on whether the second sample instruction information does not contain a third picture, and obtaining the sum of squares of the difference values ​​of corresponding pixel points in the second visual content and the third picture that have the same coordinate position depending on whether the second sample instruction information contains a third picture, and determining an overall score from the similarity score, aesthetic score and the sum of squares.

[0051] In other words, the method of this disclosure can use a multi-objective reinforcement learning method that includes model scores and rule calculations. Model scores refer to the similarity score and aesthetic score mentioned above. For example, a similarity model of text and pictures obtained through pretraining can be used to determine the similarity score between the second visual content and the third generation requirements description information, and a picture aesthetic evaluation model obtained through pretraining can be used to determine the aesthetic score of the second visual content. Rule calculation refers to calculating the sum of squares of the difference values ​​(difference values ​​for each pixel) between corresponding pixel points in the second visual content and the third picture. Of these, the similarity score is used to reflect the degree to which the target big model follows user instructions, with a higher similarity score indicating a higher degree to which the target big model follows user instructions. The aesthetic score is used to reflect the aesthetic quality of the generated second visual content, with a higher aesthetic score indicating a higher aesthetic quality of the second visual content. The sum of squares is used to reflect how closely the edited picture followed the original picture; a larger sum of squares indicates a higher degree of adherence. Accordingly, by combining similarity scores, aesthetic scores, and the sum of squares simultaneously to determine the overall score, the accuracy of the resulting overall score can be improved, and consequently, the optimization efficiency of the intermediate big model can be enhanced.

[0052] There are no restrictions on how similarity scores, aesthetic scores, sum of squares, etc., are combined to determine the overall score; for example, the overall score can be obtained by calculating it according to a predetermined formula.

[0053] Figure 6 is a flowchart of a second embodiment of the training method for the target big model according to this disclosure, in accordance with the above description. As shown in Figure 6, the following specific embodiments are included. In step 601, a pre-trained base big model is acquired. The base big model may be a multimodal big model. In step 602, first training data is acquired, which includes first sample instruction information, first sample result information including first visual content corresponding to the first sample instruction information, and first sample thinking information which is thought process information generated for the first sample instruction information.

[0054] Here, the first sample instruction information may include the second generation requirement description information, or the second generation requirement description information and a second picture corresponding to the second generation requirement description information.

[0055] Furthermore, each of the first sample thinking information may include one of the following: subdivided requirement description information obtained by subdividing the second generation requirement description information; step description information for generating the first sample result information; initial result information and optimization description information; M candidate result information corresponding to the first sample instruction information, including the first sample result information; and selection reason information for explaining why the first sample result information is superior to the other candidate result information. Here, M is a positive integer greater than 1. The first sample result information is obtained by performing an optimization process corresponding to the optimization description information on the initial result information.

[0056] In step 603, the base big model is trained based on the first training data to obtain an intermediate big model.

[0057] For example, based on the first training data, autoregressive training can be performed on a basic big model using the maximum likelihood estimation method.

[0058] In step 604, the second training data, including the second sample instruction information, is acquired.

[0059] The second sample instruction information includes the third generation requirement description information, or the third generation requirement description information and the third picture corresponding to the third generation requirement description information.

[0060] In step 605, reinforcement learning training is performed on the intermediate big model based on the second training data to obtain the target big model.

[0061] For example, by inputting second sample instruction information into the intermediate big model, intermediate result information can be obtained. The intermediate result information includes second visual content. Next, the overall evaluation result can be determined based on the intermediate result information and the second sample instruction information, and the intermediate big model can be updated according to the principles for further improving the overall evaluation result.

[0062] After obtaining a target big model, it can be applied to actual inference applications. For example, by applying it to the visual content generation method shown in Figure 1, it is possible to generate corresponding target result information based on the input target instruction information.

[0063] Furthermore, in the inference application process, after generating target result information using the target big model, the performance of the target big model can be further improved by performing reinforcement learning training on the target big model using the target result information and the corresponding target command information.

[0064] Furthermore, while the embodiments of the aforementioned method have been described as a series of operations for the sake of simplicity, as those skilled in the art will understand, some steps in this application can be performed in other orders or simultaneously, and therefore this application is not limited to the order of operations described. Next, those skilled in the art should understand that all embodiments described in the specification are preferred embodiments, and related operations and modules are not necessarily required by this application. Also, for parts not described in detail in one embodiment, the relevant descriptions in other embodiments can be referenced.

[0065] The above describes embodiments of the method; however, the embodiments of this disclosure will be further described below through embodiments of the apparatus.

[0066] Figure 7 is a schematic diagram of the configuration of an embodiment 700 of a visual content generation device based on a big model according to the present disclosure. As shown in Figure 7, it comprises an instruction acquisition module 701 and a result generation module 702. The instruction acquisition module 701 acquires target instruction information. The result generation module 702 inputs the target instruction information to a target big model and outputs corresponding target result information including target visual content, which is generated based on target thinking information, which is thought process information generated by the target big model in response to the target instruction information.

[0067] In some embodiments of this disclosure, the target big model may include a multimodal big model. The target instruction information may include first generation requirement description information, or first generation requirement description information and a first picture corresponding to the first generation requirement description information.

[0068] In some embodiments of this disclosure, the target result information may further include response statements tailored to the target instruction information, and / or the result generation module 702 may further output target thinking information simultaneously with the target result information.

[0069] Figure 8 is a schematic diagram of the configuration of Embodiment 800 of the training device for a target big model according to this disclosure. As shown in Figure 8, it comprises a model acquisition module 801, a data acquisition module 802, and a model training module 803. The model acquisition module 801 acquires a base big model after pre-training. The data acquisition module 802 acquires first training data which includes first sample instruction information, first sample result information including first visual content corresponding to the first sample instruction information, and first sample thinking information which is thinking process information generated for the first sample instruction information. The model training module 803 trains the base big model based on the first training data and determines a target big model based on the training results.

[0070] In some embodiments of this disclosure, the target big model may include a multimodal big model. The first sample instruction information includes second generation requirement description information, or second generation requirement description information and a second picture corresponding to the second generation requirement description information.

[0071] In some embodiments of this disclosure, each of the first sample thinking information may include, respectively, one of the following: subdivided requirement description information obtained by subdividing the second generation requirement description information; step description information for generating the first sample result information; initial result information and optimization description information; M candidate result information including the first sample result information corresponding to the first sample instruction information; and selection reason information for explaining why the first sample result information is superior to the other candidate result information, where M is a positive integer greater than 1. The first sample result information is obtained by performing an optimization process corresponding to the optimization description information on the initial result information.

[0072] In some embodiments of this disclosure, the model training module 803 may perform autoregressive training on the underlying big model using a maximum likelihood estimation method based on the first training data when training the underlying big model based on the first training data.

[0073] In embodiments of this disclosure, the model training module 803 can acquire an intermediate big model after training a base big model based on first training data. Subsequently, the intermediate big model may be directly determined as the target big model, or second training data including second sample instruction information may be acquired, and the intermediate big model may be reinforced learning trained based on the second training data to acquire the target big model.

[0074] In embodiments of this disclosure, a method by which the model training module 803 performs reinforcement learning training on an intermediate big model based on second training data may include inputting second sample instruction information to the intermediate big model, obtaining output intermediate result information including second visual content, determining an overall evaluation result based on the intermediate result information and the second sample instruction information, and updating the intermediate big model in accordance with the principle of improving the overall evaluation result.

[0075] In some embodiments of this disclosure, the second sample instruction information may include third generation requirement description information, or third generation requirement description information and a third picture corresponding to the third generation requirement description information. The overall evaluation result may include an overall score. Accordingly, in response to the determination that the second visual content is a picture, the method by which the model training module 803 determines the overall evaluation result based on the intermediate result information and the second sample instruction information may include obtaining a similarity score between the second visual content and the third generation requirement description information, obtaining an aesthetic score for the second visual content, determining an overall score based on the similarity score and aesthetic score in response to the determination that the second sample instruction information does not include a third picture, and obtaining the sum of squares of the difference values ​​of corresponding pixel points in the second visual content and the third picture that have the same coordinate position, in response to the determination that the second sample instruction information includes a third picture, and determining an overall score from the similarity score, aesthetic score, and sum of squares.

[0076] The specific workflows for each of the above-described apparatus embodiments can be found in the relevant descriptions in the method embodiments described above, so a detailed explanation will be omitted here.

[0077] In short, by adopting the method described in this disclosure, it is possible to improve the accuracy of visual content generation results by utilizing multimodal big model thinking chain technology. Because it can be applied to different visual content generation scenarios, it has broad applicability.

[0078] The technical solutions described in this application can be applied to the field of artificial intelligence, particularly to areas such as deep learning, big modeling, computer vision, and natural language processing. Artificial intelligence is the study of how computers can simulate human thought processes and intelligent behaviors (e.g., learning, reasoning, thinking, planning, etc.), and it encompasses both hardware-level and software-level technologies. Hardware technologies for artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Software technologies for artificial intelligence mainly include several areas such as computer vision technologies, speech recognition technologies, natural language processing technologies, and machine learning / deep learning, big data processing technologies, and knowledge mapping technologies.

[0079] The instruction information and result information in the embodiments described in this disclosure are not intended for any specific user and do not reflect the personal information of any specific user. In the proposed technology described in this disclosure, the acquisition, storage, application, processing, transmission, provision, and distribution of the personal information of the users involved all comply with the provisions of relevant laws and regulations and do not violate public order and morals.

[0080] According to embodiments of this disclosure, the disclosure further provides electronic devices, readable storage media, and computer program products.

[0081] Figure 9 is a schematic block diagram of an electronic device 900 that may be used to carry out embodiments of the present disclosure. The electronic device represents various forms of digital computers, such as laptops, desktop computers, workbenches, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may further represent various forms of mobile devices, such as PDAs, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure as described and / or requested herein.

[0082] As shown in Figure 9, the device 900 includes an arithmetic means 901 that can perform various appropriate operations and processes according to a computer program stored in a read-only memory (ROM) 902, or a computer program loaded from a storage means 908 into a random access memory (RAM) 903. The RAM 903 may store various programs and data necessary for the operation of the electronic device 900. The arithmetic means 901, ROM 902, and RAM 903 are connected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0083] Multiple components of device 900, including, for example, input means 906 such as a keyboard and mouse, output means 907 such as various types of displays and speakers, storage means 908 such as magnetic disks and optical disks, and communication means 909 such as a network card, modem, and wireless communication transceiver, are connected to the I / O interface 905. The communication means 909 enables the electronic device 900 to exchange information / data with other devices, for example, via computer networks such as the Internet and / or various telecommunication networks.

[0084] The arithmetic means 901 may be a variety of general-purpose and / or dedicated processing components having processing and computational capabilities. Some examples of the arithmetic means 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), a variety of dedicated artificial intelligence (AI) computing chips, a variety of computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The arithmetic means 901 performs the various methods and processes described above, for example, the methods of the present disclosure. For example, in some embodiments, the methods of the present disclosure may be implemented as a computer software program physically embedded in a machine-readable medium, such as a storage means 908. In some embodiments, part or all of the computer program can be loaded and / or installed into the electronic device 900 via a ROM 902 and / or a communication means 909. Once the computer program is loaded into the RAM 903 and executed by the arithmetic means 901, one or more steps of the methods of the present disclosure can be performed. Alternatively, in other embodiments, the computing means 901 may be configured in any other suitable way (e.g., via firmware) to perform the method described herein.

[0085] Various embodiments of the systems and technologies described herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented in one or more computer programs. The one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor is a dedicated or general-purpose programmable processor capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, at least one input device, and at least one output device.

[0086] Program code for carrying out the methods of this disclosure can be written using any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a dedicated computer, or another programmable data processing device so that when the program code is executed by the processor or controller, it performs the functions / operations specified in the flowchart and / or block diagrams. The program code may run entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a standalone package, or entirely on a remote machine or server.

[0087] In the context of this disclosure, a machine-readable medium is a tangible medium that is used by or in conjunction with an instruction execution system, apparatus, or device, and may contain or store programs. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of machine-readable storage media include one or more line-based electrical connections, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0088] To provide user interaction, the systems and technologies described herein may be implemented on a computer equipped with a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) for providing input from the user to the computer. Other types of devices may also be used to provide user interaction. For example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback), and input from the user may be received in any form (including sound input, voice input, or haptic input).

[0089] The systems and technologies described herein can be implemented in computing systems including backend components (e.g., data servers), computing systems including middleware components (e.g., application servers), or computing systems including frontend components (e.g., client computers having a graphical user interface or a web browser, through which users can interact with embodiments of the systems and technologies described herein), or in computing systems including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication (e.g., communication networks) in any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and internetworks.

[0090] A computer system can include a client and a server. The client and server are generally geographically distant from each other and typically interact through a communication network. The client-server relationship arises from computer programs running on corresponding computers that have a client-server relationship with each other. A server may be a cloud server, a server in a distributed system, or a server integrated with blockchain technology.

[0091] It should be understood that steps can be rearranged, added, or deleted using the various forms of flows described above. For example, each step described in this application may be performed in a parallel or sequential order, or in a different order, as long as the desired results of the proposed technology disclosed herein are achieved.

[0092] The specific embodiments described above do not constitute limitations on the scope of protection of this application. Those skilled in the art will understand that various modifications, combinations, partial combinations, and substitutions are possible in accordance with design requirements and other factors. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for generating visual content based on a big model, Steps to obtain target instruction information, The steps include inputting the aforementioned target instruction information into a target big model, obtaining and outputting the corresponding target result information, which includes target visual content, and is generated by the target big model based on target thinking information, which is thought process information generated by the target big model in response to the aforementioned target instruction information; A method for generating visual content based on a big model, including...

2. The method according to claim 1, The aforementioned target big model includes a multimodal big model, The aforementioned target instruction information includes first generation requirement description information, or the first generation requirement description information and a first picture corresponding to the first generation requirement description information. A method for generating visual content based on a big model.

3. The method according to claim 1, The aforementioned target result information further includes a response utterance that matches the aforementioned target instruction information. and / or, The method further includes the step of outputting the target result information and the target thinking information, A method for generating visual content based on a big model.

4. This is a training method for the target big model. Steps to acquire a basic big model after pre-training, A step of acquiring first training data which includes first sample instruction information, first sample result information including first visual content corresponding to the first sample instruction information, and first sample thinking information which is thinking process information generated for the first sample instruction information. The steps include training the basic big model based on the first training data and determining the target big model based on the training results, Training methods for target big models, including...

5. The method according to claim 4, The aforementioned target big model includes a multimodal big model, The first sample instruction information includes the second generation requirement description information, or the second generation requirement description information and a second picture corresponding to the second generation requirement description information. Training methods for target big models.

6. The method according to claim 5, Each of the above-mentioned sample thought information items contains, The subdivided requirements description information obtained by subdividing the second generation requirements description information, Step description information for generating the first sample result information, Initial result information and optimization description information, M candidate result information (where M is a positive integer greater than 1) that include the first sample result information corresponding to the first sample instruction information, and selection reason information for explaining why the first sample result information is superior to the other candidate result information, It includes any one of the following: The first sample result information is obtained by performing an optimization process corresponding to the optimization description information on the initial result information. Training methods for target big models.

7. The method according to claim 4, The step of training the basic big model based on the first training data is: The process includes the step of performing autoregressive training on the basic big model using a maximum likelihood estimation method based on the first training data, Training methods for target big models.

8. The method according to claim 4, The steps of training the base big model based on the first training data and determining the target big model based on the training results are as follows: A step of training the basic big model based on the first training data to obtain an intermediate big model, The steps include determining the intermediate big model as the target big model, or acquiring second training data including second sample instruction information, and performing reinforcement learning training on the intermediate big model based on the second training data to obtain the target big model. Training methods for target big models, including...

9. The method according to claim 8, The step of performing reinforcement learning training on the intermediate big model based on the second training data is: The steps include inputting the second sample instruction information into the intermediate big model to obtain intermediate result information including the output second visual content, A step of determining the overall evaluation result based on the aforementioned intermediate result information and the aforementioned second sample instruction information, The steps include updating the intermediate big model in accordance with the principles for improving the overall evaluation results, Training methods for target big models, including...

10. The method according to claim 9, The second sample instruction information includes the third generation requirement description information, or the third generation requirement description information and the third picture corresponding to the third generation requirement description information. The aforementioned overall evaluation results include the overall score, In response to the determination that the second visual content is a picture, the step of determining an overall evaluation result based on the intermediate result information and the second sample instruction information is as follows: The steps include obtaining a similarity score between the second visual content and the third generation requirements description information, and obtaining an aesthetic score for the second visual content, In response to the determination that the second sample instruction information does not include the third picture, the steps include determining the overall score based on the similarity score and the aesthetic score, In response to the determination that the second sample instruction information includes the third picture, the sum of squared differences between corresponding pixel points in the second visual content and the third picture that have the same coordinate position is obtained, and the overall score is determined from the similarity score, the aesthetic score, and the sum of squares. Training methods for target big models, including...

11. A visual content generation device based on a big model, It comprises an instruction acquisition module and a result generation module, The instruction acquisition module acquires target instruction information, The result generation module inputs the target instruction information to the target big model and outputs the corresponding target result information, which includes target visual content, and is generated by the target big model based on the target thinking information, which is the thinking process information generated by the target big model in response to the target instruction information. A device for generating visual content.

12. It is a training device for target big models. It comprises a model acquisition module, a data acquisition module, and a model training module. The aforementioned model acquisition module acquires a pre-trained base big model, The data acquisition module acquires first training data which includes first sample instruction information, first sample result information which includes first visual content corresponding to the first sample instruction information, and first sample thinking information which is thinking process information generated for the first sample instruction information. The model training module trains the base big model based on the first training data and determines the target big model based on the training results. A training device for targeting large-scale models.

13. It is an electronic device, At least one processor, The system comprises at least one processor and a memory that is communicated with it, The memory stores instructions that can be executed by the at least one processor, and when an instruction is executed by the at least one processor, the at least one processor performs the method according to any one of claims 1 to 10.

14. A non-temporary, computer-readable storage medium storing computer instructions, wherein the computer instructions cause a computer to perform the method according to any one of claims 1 to 10.

15. A computer program, when executed by a processor, that implements the method described in any one of claims 1 to 10.