Training Method and Training Device for Text Generation Model
By acquiring sample data sets and using large language models to generate predicted text, and parameter adjustment of the text generation model is solved, the problem that the image-guided text generation model is difficult to meet the customized needs, and efficient customized copy generation is achieved.
Patent Information
- Application Number
- CN202311285513.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-09-28
AI Technical Summary
The prior art is difficult to meet the needs of customized copy generation in image-guided text generation models, and there is a lack of multimodal instruction fine-tuning data, which makes the model only able to perform basic tasks and difficult to generalize to vertical tasks of different visual languages.
By obtaining the sample dataset, including the sample image and the copy content generated by the large language model, the characteristics of the sample image are determined and inputted to the large language model to generate predicted text, the text generation model is adjusted parameterically based on the difference between the sample text and the predicted text.
Customized copy generation is realized, the model's ability to generate text in specified styles is improved, the dependence on manual annotation is reduced, and the training speed and efficiency is improved.
Smart Images

Figure CN117273107B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to technical fields such as computer vision, deep learning, and large models, and can be applied to scenarios such as AIGC. Specifically, it relates to a training method, device, electronic device, computer-readable storage medium, and computer program product for a text generation model. Background Art
[0002] Artificial intelligence is a discipline that studies how to make computers simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] The methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0004] The present disclosure provides a training method, device, electronic device, computer-readable storage medium, and computer program product for a text generation model.
[0005] According to one aspect of the present disclosure, there is provided a training method for a text generation model, including: obtaining a sample data set, where the sample data set includes sample images and sample texts corresponding to the sample images, and the sample texts corresponding to the sample images include text content generated for the sample images using a large language model; determining sample image features of the sample images; inputting the sample image features into the large language model to obtain predicted texts corresponding to the sample images; and adjusting parameters of the text generation model based on the difference between the sample texts and the predicted texts.
[0006] According to another aspect of the present disclosure, there is provided a training apparatus for a text generation model, including: a sample data acquisition unit configured to acquire a sample data set, where the sample data set includes sample images and sample texts corresponding to the sample images, and the sample texts corresponding to the sample images include text contents generated for the sample images by a large language model; an image feature acquisition unit configured to determine sample image features of the sample images; a predicted text acquisition unit configured to input the sample image features into the large language model to obtain predicted texts corresponding to the sample images; and a parameter adjustment unit configured to adjust parameters of the text generation model based on the difference between the sample texts and the predicted texts.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method as described above.
[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the method as described above.
[0010] According to one or more embodiments of the present disclosure, the ability of a large language model can be utilized to quickly and efficiently generate corresponding text contents for a given image, so as to quickly obtain a multi-modal training data set for fine-tuning the instructions of the text generation model.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0012] The drawings exemplarily show embodiments and constitute a part of the specification, and are used together with the written description of the specification to explain the exemplary implementation manners of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0013] Figure 1A schematic diagram of an exemplary system in which various methods described herein can be implemented according to an embodiment of the present disclosure;
[0014] Figure 2 A method for training an image-guided text generation model according to an embodiment of the present disclosure;
[0015] Figure 3 An exemplary process for obtaining a sample data set according to an embodiment of the present disclosure;
[0016] Figure 4 An exemplary process for image-guided text generation according to an embodiment of the present disclosure;
[0017] Figure 5 An exemplary block diagram of an apparatus for training an image-guided text generation model according to an embodiment of the present disclosure;
[0018] Figure 6 A block diagram of an exemplary electronic device that can be used to implement an embodiment of the present disclosure. Detailed Description of the Embodiment
[0019] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0020] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0021] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0022] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0023] Figure 1FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein may be implemented according to embodiments of the present disclosure. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more application programs.
[0024] In embodiments of the present disclosure, the server 120 may run one or more services or software applications that enable the execution of methods for training a text generation model according to embodiments of the present disclosure.
[0025] In certain embodiments, the server 120 may also provide other services or software applications, which may include non-virtual environments and virtual environments. In certain embodiments, these services may be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software-as-a-service (SaaS) model.
[0026] In Figure 1 the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof executable by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0027] Users may use the client devices 101, 102, 103, 104, 105, and / or 106 to obtain image and text information used in the methods of embodiments of the present disclosure. The client devices may provide an interface that enables users of the client devices to interact with the client devices. The client devices may also output information to users via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.
[0028] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0029] Network 110 can be any type of network known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0030] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.
[0031] The computing units in server 120 can run one or more operating systems including any of the above operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0032] In some embodiments, server 120 can include one or more applications to analyze and combine data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0033] In some embodiments, server 120 can be a server of a distributed system, or a server incorporating a blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to address the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.
[0034] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be relational databases, for example. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.
[0035] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0036] Figure 1The system 100 can be configured and operated in various ways to enable the application of various methods and devices described in the present disclosure.
[0037] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0038] Figure 2 A method for training an image-guided text generation model according to an embodiment of the present disclosure is shown.
[0039] In step S202, a sample data set is obtained. The sample data set includes sample images and sample texts corresponding to the sample images. The sample texts corresponding to the sample images include the text content generated for the sample images using a large language model.
[0040] In step S204, the sample image features of the sample images are determined.
[0041] In step S206, the sample image features are input into a large language model to obtain predicted texts corresponding to the sample images.
[0042] In step S208, the parameters of the text generation model are adjusted based on the difference between the sample texts and the predicted texts.
[0043] Using the method for training a text generation model provided by the embodiment of the present disclosure, the ability of a large language model can be utilized to quickly and efficiently generate corresponding text content for a given image, so as to quickly obtain a multi-modal training data set for fine-tuning the text generation model.
[0044] The principle of the present disclosure will be described in detail below.
[0045] In step S202, a sample data set can be obtained. The sample data set can include sample images and sample texts corresponding to the sample images. The sample texts corresponding to the sample images can include the text content generated for the sample images using a large language model.
[0046] Image-guided text generation is widely applied to text generation tasks such as automatic advertising, comments, and messages. In Natural Language Processing (NLP), fine-tuning has been proven to be an effective method for text generation. In the fine-tuning scheme, the parameters of the model are adjusted through labeled input-output pairs so that the model can learn the output ability for specific tasks.
[0047] Instruction-tuned training methods have gradually been applied to vision-language tasks. In image-guided text generation scenarios, image features are extracted based on a frozen image feature extractor, and the image features are mapped to a text embedding space that can be understood by a large language model (LLM) through a learnable translation module. Then, the frozen large language model (LLM) is used to generate text.
[0048] However, compared with traditional text generation tasks, due to the inherently greater diversity of visual inputs, it is difficult for models to generalize to different vision-language vertical tasks. In addition, limited by the lack of multi-modal (image-text pair) instruction fine-tuning data, image-guided text generation models can only perform basic tasks such as visual scene understanding and reasoning, knowledge-based image description, and multi-turn visual dialogue, and it is difficult to meet users' customized copywriting generation needs, such as copywriting writing in a specified style, advertisement generation, comment generation with different language habits, etc.
[0049] In related technologies, general instruction fine-tuning data is collected by converting existing NLP datasets into an instruction format using fixed-template prompts. By injecting visual information into the LLM, the instruction-tuned LLM is adapted to vision-to-language generation tasks. The LLaVA model directly projects the output of the visual encoder as the input of the LLaMA / Vinuca LLM model, and the LLM is fine-tuned according to the visual-language dialogue data generated by GPT-4.
[0050] Currently, multi-modal instruction fine-tuning data is quite scarce. The annotation process for image-text pairs is very time-consuming, and the definition of the correspondence between image-text pairs is not clear, which makes it difficult to generate a large number of high-quality multi-modal instruction fine-tuning datasets.
[0051] Therefore, the present disclosure provides an image-guided customized text generation method based on a large language model. In step S202, the copywriting content corresponding to the sample image is obtained through the large language model. By using the text generation ability of the large language model to generate semi-automated annotations for customized text, parameter-efficient fine-tuning of the large language model can be achieved. In this way, a large language model for customized writing style and personalized image-guided copywriting generation can be realized. In this process, the text generation ability of the large language model can be used to quickly generate input-output pairs of images and texts without manual annotation of image-text pairs.
[0052] In some embodiments, step S202 can be implemented through the following process: Use a pre-trained second text generation model to process the sample image to obtain a text description of the sample image; Input the text description of the sample image and the prompt information into a large language model to obtain a sample text corresponding to the sample image, where the prompt information specifies the style of the sample text. In some examples, the same prompt information is used for all sample images, so that a sample dataset for a specified text generation task can be obtained.
[0053] Among them, the second text generation model can be a model different from the text generation model trained in this disclosure. The second text generation model can have the ability to obtain a description text for the image content based on the input image. In some examples, the second text generation model can be a BLIP-2 model. The sample image can be input into the pre-trained second text generation model, and the description text of the sample image can be determined based on the output of the second text generation model (such as the caption output by the BLIP-2 model). Then, the description text output by the second text model and the prompt information for the specified task can be input into the large language model, so that the large language model can output a sample text for the specified task. In some examples, the prompt information can specify the style of the sample text. For example, the prompt information can specify a social platform, so that the large language model outputs a sample text that conforms to the style of the specified social platform. Another example is that the prompt information can restrict format information such as the text length and paragraphing method of the sample text.
[0054] In step S204, the sample image features of the sample image are determined.
[0055] Step S204 can include using a feature extraction unit to extract the sample visual features of the sample image, and using a linear layer to map the sample visual features to obtain the sample image features. Among them, the sample visual features can include at least one of high-frequency image features, low-frequency image features or a combination thereof in the sample image. The specific form of the sample visual features is not limited here. Various forms of feature extraction units can be used to extract the image features in the sample image. In some embodiments, image features can be extracted by means of edge detection, corner detection, texture analysis, color histogram, etc. In other embodiments, image features can also be extracted by means of deep learning. After obtaining the sample visual features of the sample image, a linear layer can be used to map the sample visual features to obtain the sample image features. Using a linear layer can combine the feature extraction unit for the image with the large language model, so that the result output by the feature extraction unit can be aligned with the input of the large language model, enabling the large language model to better understand the relationship between the image and the text.
[0056] In some embodiments, the feature extraction unit may include a vision encoder (such as a Vision Transformer (ViT)) and a query transformer (such as a Q-Former). The vision encoder can be used to process the input image to obtain image encoding features, and the query transformer can be used to process the image encoding features to obtain image visual features.
[0057] In some examples, extracting the sample visual features of a sample image using the feature extraction unit may include: extracting the image encoding features of the sample image using the vision encoder; and processing the image encoding features and query tokens using the query transformer to obtain the sample visual features. Among them, the Q-Former can extract the features most relevant to the image from the image encoding features.
[0058] In step S206, the sample image features are input into the large language model to obtain the predicted text corresponding to the sample image.
[0059] Using the text processing ability of the large language model, the sample image features can be processed to obtain the corresponding predicted text. In some embodiments, the sample image features output in step S204 and the prompt information can be input into the large language model together to obtain the predicted text related to the prompt information. In some examples, the prompt information can indicate the style of the predicted text. For example, the prompt information used in step S206 can be the same as the prompt information used in generating the sample text above.
[0060] In step S208, the parameters of the text generation model are adjusted based on the difference between the sample text and the predicted text.
[0061] By adjusting the parameters of the text generation model, the text generation model can have better text generation ability for the specified text generation task. For example, when training the text generation model using a sample data set obtained with prompt information of a specified style, the predicted text and the sample text can be made closer by adjusting the parameters based on the difference between the sample text and the predicted text, that is, the text generation model can learn the ability to output text of the specified style.
[0062] During the training process of a general deep learning model, all parameters in the model can be adjusted, for example, by backpropagation, so that the model acquires the desired data processing ability. However, in the solution provided in this disclosure, the output of a large language model is used to implement a text generation task based on a specified image. A large language model usually has billions, tens of billions, or more parameters. Adjusting the parameters of such a large language model will consume extremely high computing resources, and adjusting the parameters of the large language model to adapt to the image-guided text generation task will also cause the large language model to forget the original language knowledge.
[0063] Therefore, this disclosure provides a parameter-efficient instruction fine-tuning method.
[0064] In step S208, the parameters of the large language model can be frozen, and only the parameters for feature extraction of the image are adjusted.
[0065] In some embodiments, when step S204 includes using a feature extraction unit to extract sample visual features of a sample image, and using a linear layer to map the sample visual features to obtain sample image features, step S208 can include freezing the parameters of the feature extraction unit and the large language model; and adjusting the parameters of the linear layer.
[0066] In other embodiments, when the feature extraction unit includes a visual encoder and a query transformer, step S208 can include freezing the parameters of the visual encoder, the query transformer, and the large language model, and adjusting at least one of the linear layer and the query token. Among them, a token vector can be randomly generated as the query token, and the parameters of the token vector are adjusted in step S208.
[0067] By freezing the parameters of the visual encoder, the query transformer, and the large language model in the text generation model, the training parameters can be kept at 1% of the overall model parameters. At the same time, only tens of thousands of stylized instruction fine-tuning data are required for model training. This greatly improves the model training speed and reduces the model training cost.
[0068] Figure 3 An exemplary process of obtaining a sample data set according to an embodiment of this disclosure is shown.
[0069] As Figure 3As shown, the model 302 can be used to process the input image 301 to obtain the descriptive text 303 of the input image 301. Among them, the model 302 can be the BLIP-2 model. Then, the descriptive text 303 and the prompt information 304 can be input into the large language model 305 together, where the prompt information 304 is used to indicate the style of the text 306 output by the large language model 305. The large language model 305 outputs the text 306 of the specified style based on the descriptive text 303 and the prompt information 304. The input-output pair composed of the input image 301 and the text 306 can be used as a sample data for the training process of the text generation model combined with Figure 2 the description. The process 300 described can be used to process multiple input images to quickly construct customized text labels that meet user needs for multiple images, so as to obtain a sufficient number of high-quality sample data for the training process of the text generation model. Figure 3
[0070] Figure 4 FIG. shows an exemplary process of image-guided text generation according to an embodiment of the present disclosure.
[0071] As Figure 4 shown, in the process 400, the input image 401 can be input into the visual encoder 402 to obtain the image encoding features of the input image. Then, the image encoding features can be used as the input of the key (K) or value (V) of the query transformer 403. The query transformer 403 outputs the sample visual features of the input image 401 based on the input key (K), value (V), and query vector (Q). The linear layer 404 can map the sample visual features. Before inputting the features into the large language model 405, the mapped sample visual features can be subjected to image embedding so that the features input into the large language model 405 conform to the input format of the large language model. The large language model can be used to process the (mapped) sample visual features to obtain the text generation result for the input image 401. When obtaining the text generation result using the large language model, the prompt information of the specified style can also be combined with the large language model, so as to obtain the text generation result of the specified style.
[0072] Figure 5 FIG. shows an exemplary block diagram of an apparatus for training an image-guided text generation model according to an embodiment of the present disclosure.
[0073] As Figure 5 shown, the apparatus 500 includes a sample data acquisition unit 510, an image feature acquisition unit 520, a predicted text acquisition unit 530, and a parameter adjustment unit 540.
[0074] The sample data acquisition unit 510 can be configured to acquire a sample data set. The sample data set includes sample images and sample texts corresponding to the sample images. The sample texts corresponding to the sample images include the copywriting content generated for the sample images using a large language model.
[0075] The image feature acquisition unit 520 can be configured to determine the sample image features of the sample images.
[0076] The predicted text acquisition unit 530 can be configured to input the sample image features into a large language model to obtain the predicted text corresponding to the sample images.
[0077] The parameter adjustment unit 540 can be configured to adjust the parameters of the text generation model based on the difference between the sample text and the predicted text.
[0078] In some embodiments, the image feature acquisition unit can be configured to: extract the sample visual features of the sample images using a feature extraction unit; and map the sample visual features using a linear layer to obtain the sample image features.
[0079] In some embodiments, the parameter adjustment unit can be configured to: freeze the parameters of the feature extraction unit and the large language model; and adjust the parameters of the linear layer.
[0080] In some embodiments, the feature extraction unit can include a visual encoder and a query transformer. Wherein, extracting the sample visual features of the sample images using the feature extraction unit includes: extracting the image encoding features of the sample images using the visual encoder; and processing the image encoding features and query tokens using the query transformer to obtain the sample visual features.
[0081] In some embodiments, the parameter adjustment unit can be configured to: freeze the parameters of the visual encoder, the query transformer, and the large language model; and adjust the parameters of at least one of the linear layer and the query tokens.
[0082] In some embodiments, the predicted text acquisition unit can be configured to: input the sample image features together with the prompt information into the large language model to obtain the predicted text related to the prompt information.
[0083] In some embodiments, the prompt information can indicate the style of the predicted text.
[0084] In some embodiments, the sample data acquisition unit can be configured to: process the sample images using a pre-trained second text generation model to obtain the text description of the sample images; and input the text description of the sample images together with the prompt information into the large language model to obtain the sample text corresponding to the sample images, where the prompt information specifies the style of the sample text.
[0085] In some embodiments, the second text generation model may be a BLIP-2 model.
[0086] Using the apparatus for training a text generation model provided by the embodiments of the present disclosure, the ability of a large language model can be utilized to quickly and efficiently generate corresponding text content for a given image, so as to quickly obtain a multi-modal training data set for instruction fine-tuning of the text generation model.
[0087] According to an embodiment of the present disclosure, there is also provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to the embodiment of the present disclosure.
[0088] According to an embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method according to the embodiment of the present disclosure.
[0089] According to an embodiment of the present disclosure, there is also provided a computer program product, including a computer program, wherein when the computer program is executed by a processor, the method according to the embodiment of the present disclosure is implemented.
[0090] Reference Figure 6 , the structural block diagram of an electronic device 600 that can be used as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0091] As Figure 6As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 602 or the computer program loaded from the storage unit 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0092] A plurality of components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device capable of inputting information into the electronic device 600. The input unit 606 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 607 can be any type of device capable of presenting information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 can include but is not limited to a magnetic disk, an optical disk. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth TM device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0093] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method 200 described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute method 200 in any other suitable manner (e.g., by means of firmware).
[0094] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0095] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0096] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0097] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0098] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0099] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0100] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0101] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after the present disclosure.
Claims
1. A training method for a text generation model, the text generation model including a feature extraction unit, a linear layer, and a large language model, the method comprises: Obtaining a sample data set, wherein the sample data set includes sample images and corresponding sample texts for the sample images, and the sample texts corresponding to the sample images include copywriting content generated for the sample images using a large language model; Extracting sample visual features of the sample images using the feature extraction unit; Mapping the sample visual features using the linear layer to obtain the sample image features; Inputting the sample image features into the large language model to obtain a predicted text corresponding to the sample image; and Adjusting the parameters of the feature extraction unit or the linear layer based on the difference between the sample text and the predicted text, wherein obtaining the sample data set includes: Processing the sample images using a pre-trained second text generation model to obtain text descriptions of the sample images; and Inputting the text descriptions of the sample images and prompt information into the large language model to obtain corresponding sample texts for the sample images, wherein the prompt information specifies the style of the sample texts.
2. The method according to claim 1, wherein, Adjusting the parameters of the text generation model includes: Freezing the parameters of the feature extraction unit and the large language model; and Adjusting the parameters of the linear layer.
3. The method according to claim 1, wherein, The feature extraction unit includes a visual encoder and a query transformer, wherein extracting sample visual features of the sample images using the feature extraction unit includes: Extracting image encoding features of the sample images using the visual encoder; Processing the image encoding features and query tokens using the query transformer to obtain the sample visual features.
4. The method according to claim 3, wherein, Adjusting the parameters of the text generation model includes: Freezing the parameters of the visual encoder, the query transformer, and the large language model; and Adjusting the parameters of at least one of the linear layer and the query tokens.
5. The method according to any one of claims 1-4, wherein, Inputting the sample image features into the large language model to obtain a predicted text corresponding to the sample image includes: Inputting the sample image features and prompt information into the large language model to obtain a predicted text related to the prompt information.
6. The method according to claim 5, wherein, The prompt information indicates the style of the predicted text.
7. The method according to claim 1, wherein, The second text generation model is a BLIP-2 model.
8. A training device for a text generation model, the text generation model including a feature extraction unit, a linear layer, and a large language model, the device comprises: A sample data acquisition unit, configured to acquire a sample data set, where the sample data set includes sample images and sample texts corresponding to the sample images, and the sample texts corresponding to the sample images include text content generated for the sample images using a large language model; An image feature acquisition unit, configured to extract sample visual features of the sample images using a feature extraction unit and map the sample visual features using a linear layer to obtain the sample image features; A predicted text acquisition unit, configured to input the sample image features into a large language model to obtain a predicted text corresponding to the sample image; And A parameter adjustment unit, configured to adjust parameters of the feature extraction unit or the linear layer based on the difference between the sample text and the predicted text, wherein the sample data acquisition unit is configured to: Process the sample image using a pre-trained second text generation model to obtain a text description of the sample image; and Input the text description of the sample image and a prompt message into a large language model to obtain a sample text corresponding to the sample image, where the prompt message specifies the style of the sample text.
9. The apparatus according to claim 8, wherein, the parameter adjustment unit is configured to: Freeze the parameters of the feature extraction unit and the large language model; and Adjust the parameters of the linear layer.
10. The apparatus according to claim 8, wherein, the feature extraction unit includes a visual encoder and a query converter, where extracting sample visual features of the sample images using the feature extraction unit includes: Extracting image encoding features of the sample images using the visual encoder; Processing the image encoding features and query tokens using the query converter to obtain the sample visual features.
11. The apparatus according to claim 10, wherein, the parameter adjustment unit is configured to: Freeze the parameters of the visual encoder, the query converter, and the large language model; and Adjust parameters of at least one of the linear layer and the query tokens.
12. The apparatus according to any one of claims 8-11, wherein, the predicted text acquisition unit is configured to: Input the sample image features and a prompt message into a large language model to obtain a predicted text related to the prompt message.
13. The apparatus according to claim 12, wherein, the prompt message indicates the style of the predicted text.
14. The apparatus according to claim 8, wherein, the second text generation model is a BLIP-2 model.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to perform the method according to any one of claims 1-7.
17. A computer program product comprising a computer program, wherein, the computer program, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Text generation method and device based on multiple modes and model training method and device based on multiple modes
CN114298121A
Stylized image description generation method based on transfer learning
CN115294427A