Method, device and equipment for determining training data of large model for generating cue words

By combining the graph-to-text model and the auxiliary large model, high-quality prompts are automatically generated, which solves the problem of resource consumption of manually constructed prompts and insufficient training data for deep learning models. It achieves the efficient generation of high-quality text-to-graph model prompts, which is suitable for a variety of image generation tasks.

CN120708600APending Publication Date: 2025-09-26BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202410353727.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, manually constructing prompts consumes a lot of human resources and the quality is unstable. Deep learning model training requires a large amount of labeled data and computing resources, making it difficult to efficiently generate high-quality text-based image model prompts.

Method used

By inputting the sample image into the image-to-text model to obtain the first description, and using the auxiliary large model to extract the main words, a second description is generated. The second description is used as a training sample and label to fine-tune the pre-trained large model to generate high-quality prompts.

Benefits of technology

It automatically generates a large number of high-quality prompts, reduces manual workload and training costs, is suitable for generalized text-to-image tasks, and improves the visual effect of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708600A_ABST
    Figure CN120708600A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method, a device and equipment for determining training data of a large model for generating prompt words. A specific embodiment of the method comprises the following steps: inputting a sample image into a graph-to-text model to obtain a first description of the sample image; inputting the first description into an auxiliary large model to obtain a subject word in the first description; generating a second description according to the subject word in the first description; an input sample is determined according to the second description, a training label of the input sample is determined according to the first description, and the input sample and the training label are used for training the pre-trained target large model again. By means of the method, a large number of high-quality prompt words used for the text graph model to generate the image can be automatically obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of large model technology, and in particular to a method, apparatus, and device for determining training data for a large model for generating prompts. Background Art

[0002] Prompts primarily refer to textual information entered by users, guiding the LargeModel (LargeModel) to generate corresponding artwork based on the prompts. In the application of large text-to-image models, a high-quality prompt can significantly enhance the visual quality of the model-generated images. Conversely, a low-quality prompt can also degrade the quality of the generated images. However, relying solely on manual construction of prompts consumes a significant amount of human resources. Furthermore, the quality of prompts is often unstable due to factors such as limited imagination and lack of knowledge. Summary of the Invention

[0003] The disclosed embodiments describe a method for fine-tuning a large model, an apparatus, a device, and a medium for generating prompts.

[0004] According to a first aspect, a method for determining training data for a large model for generating prompts is provided, comprising:

[0005] Inputting a sample image into the image-to-text model to obtain a first description of the sample image; inputting the first description into the auxiliary large model to obtain subject words in the first description; and generating a second description based on the subject words in the first description;

[0006] An input sample is determined according to the second description, and a training label of the input sample is determined according to the first description. The input sample and the training label are used to retrain the target large model that has completed pre-training.

[0007] According to a second aspect, a device for generating training data for a large model of prompts is provided, comprising:

[0008] The description acquisition unit is configured to input a sample image into the image-to-text model to obtain a first description of the sample image; input the first description into the auxiliary large model to obtain subject words in the first description; and generate a second description based on the subject words in the first description;

[0009] The retraining unit is configured to determine an input sample according to the second description and a training label of the input sample according to the first description, wherein the input sample and the training label are used to retrain the target large model that has completed pre-training.

[0010] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.

[0011] According to a fourth aspect, an electronic device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.

[0012] According to an embodiment of the present disclosure, a method, apparatus, device, and medium for determining training data for a large model for generating prompts are provided. First, a sample image can be input into a text-to-graph model to obtain a first description of the sample image; the first description can be input into an auxiliary large model to obtain the main words in the first description; and a second description can be generated based on the main words in the first description. Then, an input sample is determined based on the second description, and a training label for the input sample is determined based on the first description. The input sample and the training label are used to retrain the target large model that has completed pre-training. Through this method, a large number of high-quality prompts for generating images using a text-to-graph model can be automatically obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A schematic diagram showing the relationship between prompts and generated images;

[0014] Figure 2 A schematic diagram of a prompt generation scheme is shown;

[0015] Figure 3 A schematic diagram illustrating a method for determining training data for a large model for generating prompts according to an embodiment of the present disclosure is shown;

[0016] Figure 4 A schematic flow chart of a method for determining training data for a large model for generating prompts according to an embodiment of the present disclosure is shown;

[0017] Figure 5 A schematic diagram of adding a word limit in the second description according to an embodiment of the present disclosure is shown;

[0018] Figure 6 A schematic diagram of adding a style indication in the second description according to an embodiment of the present disclosure is shown;

[0019] Figure 7 A schematic block diagram of a device for determining training data for a large model for generating prompts according to an embodiment of the present disclosure is shown;

[0020] Figure 8 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown;

[0021] Figure 9A schematic diagram of the structure of a storage medium suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0022] The technical solutions provided in this specification are further described in detail below in conjunction with the accompanying drawings and embodiments. It will be understood that the specific embodiments described herein are merely for explaining the relevant inventions and are not intended to limit the inventions. It should also be noted that, for ease of description, only the portions relevant to the relevant inventions are shown in the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features therein may be combined with each other.

[0023] In the description of the implementations of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one / an implementation" or "the implementation" should be understood as "at least one / an implementation." The term "some implementations" should be understood as "at least some implementations." Other explicit and implicit definitions may be included below.

[0024] As mentioned earlier, prompts primarily refer to textual information entered by the user, guiding the Large Model to generate the corresponding artwork based on the prompt. In the application of the Text to Image Large Model, a high-quality prompt can significantly enhance the visual quality of the model-generated image, while a low-quality prompt can also degrade the visual quality of the generated image. Figure 1 A schematic diagram showing the relationship between the prompt and the generated image. Figure 1For example, let's say we want the Vincent image model to generate an image of a "waterfall." Compared to the image generated by inputting a relatively crude, low-quality description, such as "a waterfall falls from a height onto rocks," into the Vincent image model, inputting high-quality prompts, such as more specific prompts that describe the image entity and its environment, background, and effects in detail, such as "a spectacular waterfall scene, the waterfall falls from a height, splashing water onto rocks. The rocks are covered with green vegetation, and the rocks are irregularly shaped, with some areas forming depressions. The background of the waterfall is a mountain range, with clouds and mist shrouding the peaks, adding a mysterious atmosphere to the entire scene," typically produces an image effect with more specific details that meet the user's expectations, such as more obvious "clouds and mist" and "irregularly shaped rocks." Therefore, high-quality prompts often make the model-generated image closer to the user's desired visual effect. In other words, high-quality prompts have a significant positive effect on image generation quality. However, the existing method of manually constructing prompts not only consumes a large amount of human resources, but also suffers from inconsistent quality due to factors such as limited imagination and lack of knowledge. Therefore, some users or manufacturers using large-scale text-based image models hope to generate high-quality prompts through automated methods. Figure 2 A schematic diagram of a prompt generation scheme is shown. Figure 2 As shown, this type of method usually constructs an automatic prompt generation system, which can automatically generate prompts based on preset prompt generation rules or generation templates. However, this type of system usually generates prompts based on generation rules or generation templates for specific image generation tasks, and is difficult to directly use for other generalized image generation tasks other than specific tasks. If used for other generation tasks, a large amount of manual work is required to set up and maintain the generation rules or generation templates for different tasks. Another solution is to learn to construct prompts from a large amount of training data through training of deep learning models. However, this solution also has the following problems: training deep learning models usually requires a large amount of labeled data. This labeled data is often difficult to obtain or costly. In addition, training a deep learning model from scratch often requires a large amount of hardware computing resources, which are difficult to obtain or costly to obtain for many users.

[0025] In order to solve the above technical problems, an embodiment of the present disclosure provides a method for determining training data for a large model for generating prompts. Figure 3 FIG. 1 is a schematic diagram showing a method for determining training data for a large model for generating prompts according to an embodiment of the present disclosure. Figure 3As shown, in some embodiments, for example, a sample image can be input into a text-to-image model to obtain an image description of the sample image. The image description is then input into an auxiliary large model to obtain the subject words in the image description. Based on the subject words and a preset template, a simple description of the image that the text-to-image large model wants to generate is determined. Since this simple description of the image is used as an input sample in the subsequent fine-tuning of the large model to provide a prompt for a more specific and detailed image description output by the large model, it can also be called a simple prompt. Then, using the simple prompt as the input sample and the image description of the sample image as the sample label, fine-tuning (Fine tuning) is performed on the target large model.

[0026] The advantages of this method are: on the one hand, it can automatically generate a large number of high-quality prompts for the large model of cultural images, and use them for generalized cultural image tasks. There is no need to set templates or generation rules for prompts for specific image tasks. Under the premise of being able to generate a large number of high-quality prompts, the manual workload is greatly reduced. On the other hand, through this method, the training samples and sample labels used for fine-tuning the large model itself for generating prompts can be easily obtained without a large amount of manual labeling work, which greatly reduces the manual workload consumed by fine-tuning the large model itself. Thirdly, compared with retraining the deep learning model, generating high-quality prompts by fine-tuning the trained large model can greatly reduce the training data and training costs consumed by model training.

[0027] The detailed process of this method is further described below.

[0028] Figure 4 FIG. 1 is a flow chart showing a method for determining training data for a large model for generating prompts according to an embodiment of the present disclosure. Figure 4 As shown, the method comprises at least the following steps:

[0029] Step S401: Input a sample image into the image-to-text model to obtain a first description of the sample image; input the first description into the auxiliary large model to obtain the main words in the first description; and generate a second description based on the main words in the first description.

[0030] Step S403: Determine an input sample according to the second description, and determine a training label of the input sample according to the first description. The input sample and the training label are used to retrain the pre-trained target large model.

[0031] First, in step S401, a sample image can be input into a graph-based text model to obtain a first description of the sample image. In this step, a sample image can be input into a graph-based text model to obtain a linguistic description of the sample image (first description). A graph-based text model is a neural network model based on deep learning technology, which can generate text descriptions related to an input image. In different embodiments, the specific structure and type of the neural network on which the graph-based text model is based may be different, and this specification does not limit this. In one embodiment, for example, it may be a graph-based text model based on a generative adversarial network (GAN). In another embodiment, for example, it may be a graph-based text model based on a Transformer model. In one embodiment, for example, it may specifically be a Qwen_VL (Qwen Vision-Language) model.

[0032] In different embodiments, the sample images may be images of different specific subjects. This specification is not limited to this. In one example, an image of a cat may be input into an image-to-text model, and the following description of the image may be obtained: "This is a picture of a Scottish straight cat. The cat has black and white fur, big eyes, and a pink nose. It is sitting on a white sofa with a white curtain in the background."

[0033] After obtaining the first description, the first description can be input into the auxiliary macro model to obtain the subject words in the first description. Subject words are words that indicate the primary entity in the image and are used to refer to concrete or abstract things in the real or virtual world. For example, they can represent people, places, objects, concepts, events, etc. The auxiliary macro model is a macro model that can be used to extract the subject words from the image description. Specifically, in one embodiment, the first description and a prompt for extracting the subject words from the first description can be input into the auxiliary macro model to obtain the subject words. In one example, for example, the image description "This is a picture of a Scottish straight cat. The cat has black and white fur, large eyes, and a pink nose. It is sitting on a white sofa with a white curtain in the background" and a prompt for extracting the subject words from the first description, such as "Extract the core object in this sentence," can be input into the auxiliary macro model to obtain the subject word "Scottish straight cat." In different embodiments, the prompt for extracting the subject words from the first description can have different specific forms, and this specification does not limit this.

[0034] After obtaining the subject words in the first description, a second description can be generated based on the subject words in the first description. In one embodiment, the second description can be generated based on the subject words and a preset description template. In different embodiments, the specific form of the preset template may be different, and this specification does not limit this. For example, in one example, the preset template may specifically include one or more of the preset prompt templates "Randomly generate a prompt related to...", "Describe a scene containing...", "Generate a prompt containing the keyword...", "Randomly generate a prompt related to...", etc. Thus, for example, the placeholder "..." in these prompt templates can be replaced with the subject words to obtain a description statement containing the subject words, i.e., the second description. For example, "Describe a scene containing a Scottish Straight-eared cat", "Randomly generate a prompt related to a Scottish Straight-eared cat", "Generate a prompt containing the keyword Scottish Straight-eared cat", etc. Since the second description is used as an input sample in the fine-tuning of the target large model in the subsequent steps to guide the target large model to output prompts for the text image model, in the above prompt template example, the second description can, for example, include instructions for generating prompts for the target large model. For example, "randomly generate a prompt", "generate a prompt", etc. Because the second description includes a simple description of the image and can be used as a training sample that matches the detailed description of the image (i.e., the first description, as a sample label) in the subsequent steps, the above method can conveniently obtain the training samples required for fine-tuning the target large model in the subsequent steps.

[0035] In some scenarios, more than one subject word may be obtained from the auxiliary large model. Therefore, in one embodiment, the first description may be input into the auxiliary large model to obtain multiple subject words in the first description; and a second description may be generated based on one or more of the multiple subject words. In one example, for example, multiple subject words "clouds", "houses", and "towns" may be obtained from the auxiliary large model, and based on multiple of these subject words, a second description "randomly generate a prompt containing the keywords of clouds, houses, and towns" may be obtained. In the above manner, a second description is generated based on multiple subject words, and used to fine-tune the sample labels of the large model in subsequent steps, which often achieves a better fine-tuning effect on the target large model, and enables the fine-tuned target large model to generate a higher-quality detailed description of the image based on the simple description of the image by multiple subject words.

[0036] Then, in step S403, the input sample can be determined according to the second description obtained in step S401, and the training label of the input sample can be determined according to the first description obtained in step S401. The input sample and the training label can be used to retrain the target large model that has completed pre-training, or to fine-tune the target large model.

[0037] In different embodiments, the specific method of determining the input sample based on the second description or determining the training label of the input sample based on the first description may be different. In one embodiment, the second description may be used as the input sample. In another embodiment, a word limit statement may be added to the second description to obtain the input sample, such as Figure 5 As shown. In one example, for example, a word limit statement "about 20 to 30 words" can be added to the second description "randomly generate a prompt related to the countryside" to obtain the input sample "randomly generate a prompt related to the countryside, about 20 to 30 words". Since the input sample is usually paired with the first description (used as a training label) for fine-tuning the target large model. Therefore, in a specific embodiment, the word limit statement can be determined based on the number of words in the first description. For example, the first description paired with the second description is "In a quiet countryside, a red cloud sets off the beautiful sky", which is 25 words and is in the range of 20 to 30 words. Therefore, the word limit statement can be determined as "about 20 to 30 words". In the above manner, the fine-tuned target large model can generate a detailed description of the image with the expected number of words based on the simple description of the image with the word limit statement added.

[0038] In another embodiment, the first description may be converted to a set style to obtain the training label; and an instruction statement for the set style may be inserted into the second description to obtain the input sample. Figure 6 FIG. 1 shows a schematic diagram of adding a style indication in the second description according to an embodiment of the present disclosure. Figure 6As shown, for example, the first description (image description) can be input into another macro model to transform the first description into a set description style, thereby obtaining a training label for the input sample. In different embodiments, the specific structure or type of the other macro model may vary, and this specification does not limit this. Furthermore, an indication of the set description style is added to the second description (simple description, or simple prompt) to obtain an input sample. In one example, the first description, "A small town hidden in the mountains, surrounded by misty peaks, appears very mysterious. The town has many houses and a bridge spanning the town." can be added with the style indication "cyberpunk style" and then input into another macro model. The resulting image description "cyberpunk style" is used as a sample label: "A small town hidden in the mountains, surrounded by illusory clouds and towering technological peaks, exudes a mysterious atmosphere. The town is dotted with futuristic houses and a bridge spanning the town, shimmering with a cold light." Then, the style indication "cyberpunk style" is added to the second description, "Generate prompt, about the town in the mountains," to obtain an input sample.

[0039] After determining the input samples and the training labels of the input samples, the target large model can be trained (fine-tuned) again based on the input samples and the training labels. The target large model is a large model used to generate a detailed description of the image based on the simple prompts of the image. In different specific embodiments, the specific structure or specific type of other large models may be different, and this specification does not limit this. Fine-tuning refers to making some targeted adjustments and optimizations based on the existing pre-trained large model to adapt to specific tasks or data sets. Specifically, for example, the input samples and sample labels can be used to perform supervised learning on the pre-trained large model, such as using the back-propagation algorithm to update the parameters of the large model.

[0040] After fine-tuning the target large model, a simple description containing different subject words can be input into the target large model to generate more detailed descriptions of the image, which can be used as high-quality input prompts (prompts) for the text graph model. Therefore, in one embodiment, after fine-tuning the target large model, a third description containing subject words can be input into the retrained target large model to obtain a fourth description, which is used to input the text graph model to obtain the image corresponding to the fourth description. In one embodiment, multiple different fourth descriptions can also be obtained based on the same third description based on the retrained target large model. In the above manner, a large number of high-quality prompts can be easily obtained for the text graph model to generate high-quality images.

[0041] In various embodiments, high-quality images generated by inputting high-quality prompts into a text graph model can be used for various specific purposes or businesses, without limitation in this specification. In one example, they can be used for user interface (UI) design. In another example, the high-quality input prompts and high-quality images can also be used to train other text graph models of different specific types.

[0042] Figure 7 A schematic block diagram of a device for determining training data for a large model for generating prompts according to an embodiment of the present disclosure is shown. The device is used to perform the following steps: Figure 4 As shown in the method. Figure 7 As shown, the apparatus 700 includes:

[0043] The description acquisition unit 701 is configured to input a sample image into the image-to-text model to obtain a first description of the sample image; input the first description into the auxiliary large model to obtain subject words in the first description; and generate a second description based on the subject words in the first description;

[0044] The retraining unit 702 is configured to determine an input sample according to the second description and a training label of the input sample according to the first description, wherein the input sample and the training label are used to retrain the target large model that has completed pre-training.

[0045] The present disclosure also provides an electronic device including a memory and a processor. The memory stores executable code. When the processor executes the executable code, the following is achieved: Figure 4 The method shown.

[0046] You can also refer to the following Figure 8 , which shows a structural diagram of an electronic device 800 suitable for implementing an embodiment of the present application. Figure 8 The electronic device 800 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0047] like Figure 8As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801. The processing device 801 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0048] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Figure 8 The electronic device 800 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 8 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0049] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method for determining the training data of a large model for generating prompts provided in the embodiment of the present application are executed.

[0050] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the following steps provided in the present disclosure: Figure 4 A method for determining training data for a large model for generating prompts is shown in . Figure 9 A schematic diagram of a storage medium for implementing an embodiment of the present application. Figure 9 As shown, the storage medium 900 can be a non-transitory computer-readable storage medium for storing non-transitory computer-executable instructions 901. When the non-transitory computer-executable instructions 901 are executed by the processor, a method for determining training data for a large model for generating prompts provided in an embodiment of the present application can be implemented. For example, when the non-transitory computer-executable instructions 901 are executed by the processor, one or more steps in a method for determining training data for a large model for generating prompts provided in an embodiment of the present application can be executed. For example, the storage medium 900 can be applied to the above-mentioned electronic device. For example, the storage medium 900 may include a memory in the electronic device. For the description of the storage medium 900, reference can be made to the description of the memory in the embodiment of the electronic device, and the repeated parts will not be repeated here. The specific functions and technical effects of the storage medium 900 can be referred to the description of the method for determining training data for a large model for generating prompts provided in an embodiment of the present application, and will not be repeated here.

[0051] It should be noted that the computer-readable medium of the embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a memory card of a smart phone, a storage component of a tablet computer, a portable computer disk, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. In the embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0052] The computer-readable medium may be included in the electronic device, or may exist independently and not be incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the server, the one or more programs enable the electronic device to implement the method for determining training data for a large model for generating prompts provided in an embodiment of the present application.

[0053] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0054] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions. The units involved in the embodiments described in the present disclosure may be implemented using software or hardware. The name of the unit does not, in some cases, constitute a limitation on the unit itself. The functions described above in this document may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0055] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences between the other embodiments. In particular, the storage medium and computing device embodiments are described briefly because they are generally similar to the method embodiments. For relevant portions, refer to the description of the method embodiments.

[0056] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned disclosed concepts. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the present disclosure to form a technical solution. In addition, although the operations are described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination.

[0057] The above specific implementation methods further describe in detail the purpose, technical solutions and beneficial effects of the embodiments of the present invention. Although the subject matter has been described in a language specific to structural features and / or method logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims. It should be understood that the above are only specific implementation methods of the embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for determining training data for a large model for generating prompts, comprising: Inputting a sample image into the image-to-text model to obtain a first description of the sample image; Inputting the first description into the auxiliary large model to obtain the subject words in the first description; Generate a second description based on the subject words in the first description; An input sample is determined according to the second description, and a training label of the input sample is determined according to the first description. The input sample and the training label are used to retrain the target large model that has completed pre-training.

2. The method according to claim 1, wherein Generating a second description according to the subject words in the first description includes: generating the second description according to the subject words and a preset description template.

3. The method according to claim 1, wherein Input the first description into the auxiliary large model to obtain the subject words in the first description, including: The first description and a prompt indicating the extraction of a subject word in the first description are input into the auxiliary macro model to obtain the subject word.

4. The method according to claim 1, wherein Determining an input sample according to the second description includes: adding a word limit statement to the second description to obtain the input sample.

5. The method according to claim 1, wherein The word limit statement is determined according to the word count of the first description.

6. The method according to claim 1, wherein Determining an input sample based on the second description, and determining a training label of the input sample based on the first description, including: converting the first description into a set style to obtain the training label; An instruction statement for the set style is inserted into the second description to obtain the input sample.

7. The method according to claim 1, wherein Inputting the first description into the auxiliary macro model to obtain subject words in the first description, and generating a second description based on the subject words in the first description, including: inputting the first description into the auxiliary macro model to obtain multiple subject words in the first description; A second description is generated based on one or more of the plurality of subject words.

8. The method according to claim 1, further comprising: The third description containing the subject word is input into the retrained target large model to obtain a fourth description. The fourth description is used to input into the text-image model to obtain an image corresponding to the fourth description.

9. A device for determining training data for a large model for generating prompts, comprising: a description acquisition unit configured to input a sample image into the image-to-text model to obtain a first description of the sample image; Inputting the first description into the auxiliary large model to obtain the subject words in the first description; generating a second description based on the subject words in the first description; The retraining unit is configured to determine an input sample according to the second description and a training label of the input sample according to the first description, wherein the input sample and the training label are used to retrain the target large model that has completed pre-training.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 8.

11. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multi-task classification model training method, classification method, model and device

    CN116933120A

  • Text generation model training method and training device

    CN117273107A

  • Generation method and device of cue word information of large language model, equipment and medium

    CN117539975A

  • Method and device for training cue word generation model, equipment and medium

    CN117557885A

  • Text image generation method and device based on intelligent assistance

    CN117689755A