Pentograph model training method and device, electronic equipment and storage medium
By introducing a linear layer of attention mechanism into the literary and biographical graph model and performing distillation training, the problem of the existing literary and biographical graph model taking time and insufficient image quality is solved, and efficient and high-speed literary and biographical graph model training and generation effect is achieved.
Patent Information
- Application Number
- CN202411885947.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-09
AI Technical Summary
Existing large-scale text-to-graph diffusion generation models require a large number of inference sampling steps when generating high-quality images, resulting in high computational costs and long delays, limiting their application in real-world scenarios or resource-constrained environments.
By obtaining the distillation training data set from the pretrained data set and training the second vernacular graph model based on the data set and the pretrained first vernacular graph model. The second literary graph model adds a linear layer of attention mechanism based on the first literary graph model, and freezes the structural parameters of the same part during the training process to speed up the training process.
It effectively accelerates the training time-consuming of literary and genomic graphics models, improves training efficiency, and ensures high quality of generated images, suitable for real-life scenarios and resource-constrained environments.
Smart Images

Figure CN119962623A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the technical fields of computer vision, deep learning, large models, etc., and can be applied to scenarios such as artificial intelligence generated content (AIGC); specifically, it relates to a training method, device, electronic device and storage medium for a text-based graph model. Background Art
[0002] With the development of deep learning technology, various neural network models have achieved remarkable results in a series of downstream tasks. They provide many convenient services and help for human life through real-time human-computer interaction.
[0003] For example, the diffusion generation model has also achieved remarkable results in the field of text-to-image generation, and can accurately and efficiently generate images based on text description information. Summary of the invention
[0004] The present invention provides a training method, device, electronic device and storage medium for a text graph model.
[0005] According to one aspect of the present disclosure, a method for training a text graph model is provided, comprising:
[0006] Obtaining a distillation training data set from a pre-training data set; the distillation training data set includes a plurality of samples;
[0007] Based on the distilled training data set and the pre-trained first text graph model, a second text graph model is trained; the second text graph model adds an attention mechanism linear layer on the basis of the first text graph model; during the training process, the parameters of the structure of the second text graph model and the first text graph model that are the same are frozen.
[0008] According to another aspect of the present disclosure, a training device for a text graph model is provided, comprising:
[0009] An acquisition module is used to acquire a distillation training data set from a pre-training data set; the distillation training data set includes a plurality of samples;
[0010] A training module is used to train a second text graph model based on the distilled training data set and the pre-trained first text graph model; the second text graph model adds an attention mechanism linear layer on the basis of the first text graph model; during the training process, the parameters of the structure of the second text graph model and the first text graph model are frozen.
[0011] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0012] at least one processor; and
[0013] a memory communicatively connected to the at least one processor; wherein,
[0014] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any possible implementation manner and the aspects described above.
[0015] According to yet another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method of the above-mentioned aspect and any possible implementation manner.
[0016] According to yet another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements the above-mentioned aspects and any possible implementation method.
[0017] According to the technology disclosed in the present invention, the training time of the cultural graph model can be effectively accelerated and the training efficiency can be improved.
[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0020] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0021] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0022] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0023] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0024] Figure 5 The block diagram is a block diagram of an electronic device for implementing the method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] Obviously, the described embodiments are only part of the embodiments of the present disclosure, but not all of them. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present disclosure.
[0027] It should be noted that the terminal devices involved in the embodiments of the present disclosure may include but are not limited to mobile phones, personal digital assistants (PDAs), wireless handheld devices, tablet computers and other smart devices; display devices may include but are not limited to personal computers, televisions and other devices with display functions.
[0028] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0029] Existing large-scale diffusion generation models for text-to-image processing usually require a large number of inference sampling steps to generate high-quality images, which results in high computational costs and long delays, limiting their application in real-world scenarios or resource-constrained environments. Based on this, in the prior art, there are also acceleration schemes for diffusion generation models that accelerate inference by reducing the number of sampling steps, but this will result in significantly lower quality of generated images than before acceleration.
[0030] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure; Figure 1 As shown, this embodiment provides a training method for a text graph model, which may specifically include the following steps:
[0031] S101, obtaining a distillation training data set from a pre-training data set; the distillation training data set includes a plurality of samples;
[0032] The subject of executing the training method of the text graph model of this embodiment is a training device of the text graph model, which may be an electronic entity or an application integrated by software.
[0033] S102: Training a second text-generated graph model based on the distilled training data set and the pre-trained first text-generated graph model.
[0034] The second text image model in this embodiment adds an attention mechanism linear layer on the basis of the first text image model. During the training process, the parameters of the structure of the same part of the second text image model and the first text image model are frozen, that is, the training of this embodiment only adjusts the parameters of the attention mechanism linear layer added to the second text image model. Since the first text image model, as a teacher model, already has a strong text image capability and can generate very high-quality pictures. Therefore, as a student model, the second text image model can inherit the high-quality image generation capability of the teacher model. By introducing the attention mechanism linear layer, the image generation effect and efficiency of the text image model are further improved. At the same time, during the training process, since only the parameters of the attention mechanism linear layer are adjusted, and the attention mechanism linear layer has very few parameters added relative to the first text image model, the training scheme of this embodiment can effectively speed up the training time of the text image model and improve the training efficiency.
[0035] The training method of the text graph model of this embodiment, by adopting the above technical solution, can quickly and efficiently train the text graph model, and the trained text graph model can also follow the characteristics of the teacher model and have high-quality graph generation capabilities.
[0036] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure; the training method of the Wensheng graph model of this embodiment, in the above Figure 1 Based on the technical solutions of the embodiments shown in the figure, the technical solutions of the present disclosure are further described in more detail. Figure 2 As shown, the training method of the cultural graph model of this embodiment may specifically include the following steps:
[0037] S201, obtaining at least one of the image aesthetics, image clarity and image resolution of the images in each training data in the pre-training data set;
[0038] S202, distilling a distilled training data set from the pre-training data set based on at least one of the image aesthetics, image clarity, and image resolution of the images in each training data in the pre-training data set;
[0039] Specifically, the picture beauty calculation operator and the picture clarity calculation operator may be used to respectively calculate the picture beauty and picture clarity of the pictures in each training data in the pre-training data set, and extract the resolution of each picture.
[0040] Then, based on at least one of the image aesthetics, image clarity and image resolution of each image, the distillation training data set is screened. Preferably, in order to improve the quality of the distillation training data set, the three features of image aesthetics, image clarity and image resolution can all participate in the distillation of the distillation data set. During the specific distillation, it can be detected whether the image aesthetics of each photo meets the preset aesthetics threshold requirement, whether the image clarity meets the preset clarity threshold requirement, and whether the image resolution meets the preset resolution threshold requirement. Only when all three requirements are met can they be distilled into the distillation data set.
[0041] Of course, according to actual needs, only one or two of the requirements for image aesthetics, image clarity, and image resolution can be selected to distill the distilled data set. In practical applications, other distillation methods can also be used to distill the training data set, as long as a high-quality training data set can be obtained, and this is not limited here.
[0042] The pre-trained data set of this embodiment can be a training data set used in the pre-training of the first text graph model. In this embodiment, the pre-training data set used in the training of the teacher model can be used to distill the training data set, which can ensure the accuracy of the first text graph model as the teacher model in solving the ordinary differential equation (ODE) during the training process.
[0043] Although the above distillation mainly refers to the pictures in the training data, it should be noted that each training data includes not only pictures but also text description information of the pictures.
[0044] Through the above method, a high-quality distillation training data set can be obtained, which provides effective data support for the subsequent training of the text-generated graph model and can effectively improve the quality of the raw images of the trained text-generated graph model.
[0045] S203, obtaining samples from the distillation training data set;
[0046] Next, we extract samples from the distillation training dataset to train the text graph model. Since the training is performed using samples from the distillation training dataset, this training process can also be called distillation training. The distillation training dataset includes multiple samples, each of which includes a training image and corresponding text description information.
[0047] S204, adding noise to the training image in the sample to obtain a first noise image;
[0048] For example, the training image in the sample is denoised at a random sampling time step to obtain a first noise image, that is, the first noise image is a noise image at a random sampling time step.
[0049] S205, using the pre-trained first text image model to solve the first noise image to obtain a second noise image;
[0050] The first noise image is solved by the first text image model to obtain a second noise image, but the noise is not completely removed.
[0051] Specifically, the first text-generated graph model can be used to solve the first noise image of the random sampling time step for a preset number of steps to obtain the second noise image, which is the noise image of the random sampling time step minus the preset number of steps.
[0052] For example, in practical applications, the random sampling time step can be taken as the tth step, where t is any time step of random sampling and is not limited here. Correspondingly, the first noise image is the noise image x after the tth step of noise addition. t The preset number of steps in this embodiment may be k steps, and correspondingly, the second noise image may be represented as x t-k , that is, the noise image x after adding noise in the tth step t The preset number of steps k is solved. In this embodiment, k can be 15, 20, 25, etc., preferably, 20.
[0053] For example, at this time, the input of the first text-generated image model is the text description information in the sample and the noise image of the random sampling time step, such as the noise image x after the t-th step noise addition. t , at this time, the first text-generated image model can be based on the text description information in the sample, and the noise image x after the t-th step of noise addition t Solve the preset number of steps k and get x t-k .
[0054] S206, training a second noise image model based on the first noise image and the second noise image;
[0055] In this embodiment, multiple samples in the distillation training data set can be used to train the second text raw image model. During the training process, the first text raw image model is the teacher model, and the second text raw image model is the student model. During the entire training process, the parameters of the first text raw image model and the network part with the same structure as the first text raw image model are frozen, that is, fixed, and only the LORA parameters of the linear layer of the attention mechanism can be adjusted, which can ensure the stability of the raw image quality.
[0056] During the training process of this embodiment, the teacher model, i.e., the first text graph model, is hot-started using the text graph model that has completed training, and the parameters are frozen during the training process. The student model, i.e., the second text graph model, is based on the first text graph model that has completed training, and a very small amount of additional LORA parameters, such as about 4%, are added. Only the LORA parameters are updated during training, while other parameters remain frozen, which can effectively speed up training time and improve training efficiency.
[0057] For example, in this embodiment, when step S206 is specifically implemented, it may include the following steps:
[0058] (1) using the second text image model to predict the first noise-free image based on the first noise image;
[0059] (2) using the second text image model to predict a second noise-free image based on the second noise image;
[0060] (3) Based on the first noise-free image and the second noise-free image, the low-rank adaptation (LORA) parameters of the linear layer of the attention mechanism in the second text-based graph model are updated.
[0061] For example, specifically, the second text image model is based on the first noise image x t and the second noise image x t-k , predict f(x t ,t) and f(x t-k ,tk), that is, for input data points x with different noise levels t and x t-k , the noise-free sample points x0 predicted by the student model, i.e. the second-language graph model (t) and x0 (t-k) , x0 (t) That is the first noise-free image, x0 (t-k) This is the second noise-free image.
[0062] During prediction, the text description information in the sample and the first noise image x can be input into the second text image model. t At this time, the second text-generated image model is based on the text description information, and the first noise image x t Decode and get the first noise-free image x0 (t) .
[0063] Then, the text description information in the sample and the second noise image x are input to the second text image model. t-k At this time, the second text image model is based on the text description information, and the second noise image x t-k , decoded to obtain the second noise-free image x0 (t-k) .
[0064] Finally, based on x0 (t) and x0 (t-k) , update the low-rank adaptation (LORA) parameters of the linear layer of the attention mechanism in the second text graph model. Specifically, we can first calculate x0 (t) and x0 (t-k) The distance between them is used as the loss function of consistency distillation to update the student model, i.e., the second text graph model. When updating, the LORA parameters of the linear layer of the attention mechanism in the second text graph model are adjusted with the goal of convergence of the loss function.
[0065] By adopting the loss function of consistency distillation constructed in this way to guide the adjustment of the LORA parameters of the linear layer of the attention mechanism in the second-text raw image model, the second-text raw image model can quickly restore the noise-free image based on the noisy image of any number of steps, which can effectively shorten the image generation steps of the second-text raw image model, that is, reduce the image generation time of the second-text raw image model, and thus effectively improve the image generation efficiency of the second-text raw image model.
[0066] S207. During the training process, based on the LORA parameters of the linear layer of the attention mechanism, update the Exponential Moving Average (EMA) parameters;
[0067] For example, specifically, the updated EMA parameter can be taken as weight 1*EMA parameter + weight 2*LORA parameter, wherein weight 1 and weight 2 can be set based on experience and needs, and are not limited here.
[0068] Optionally, in one embodiment of the present disclosure, during the training process, the EMA parameters do not need to be updated in every training step. During the training process, the EMA parameters can be updated based on the LORA parameters of the linear layer of the attention mechanism at every preset training step, which can further improve the quality of the images generated after distillation training.
[0069] S208: Based on the second text graph model and EMA parameters obtained after the training, a target text graph model is constructed, and the target text graph model is used to perform the reasoning task.
[0070] Specifically, in this embodiment, after the training is completed, the second text graph model obtained through training can be directly used as the target text graph model in the inference stage. Alternatively, after the training is completed, the LORA parameters of the attention mechanism linear layer of the second text graph model obtained after the training are replaced with EMA parameters to obtain the target text graph model.
[0071] Compared with LORA parameters, EMA parameters aggregate the information of the entire training process and can perform the reasoning task of the Wensheng graph more accurately. Therefore, in this embodiment, the target Wensheng graph model is constructed based on the trained second Wensheng graph model and EMA parameters. Compared with the second Wensheng graph model, the Wensheng graph can be realized more accurately and efficiently, and the quality of the Wensheng graph model generated pictures can be improved. At the same time, the target Wensheng graph model also has the advantages of the second Wensheng graph model, which can effectively shorten the image generation steps, reduce the time consumption of image generation, and improve the image generation efficiency. Moreover, due to the use of EMA parameters, the image generation quality of the target Wensheng graph model can be further improved.
[0072] Existing consistency distillation training schemes usually take the guiding parameter w of the text condition as a condition and input it into the student model. During the training process, w is uniformly sampled within a range. However, in this embodiment, the parameters in the student model with the same structure as the teacher model are frozen, including the guiding parameter w of the fixed text condition. That is, during training, the parameter w is set to a value consistent with the normal reasoning of the hot start model, which can help the distillation training to be more stable. Because in the distillation training of this embodiment, both the teacher model and the student model contain the parameters of the hot start original text raw image model, if the independent teacher model and student model are loaded at the same time, the video memory requirements of the training machine are relatively high. In this embodiment, during training, only the same frozen original text raw image model parameters can be retained for sharing by the teacher and the student. The newly added LORA parameters in the student model are dynamically inserted into the original model parameters according to the training process for training, such as solving the ODE x in the teacher model. t-k When , the LORA parameters are not inserted in the calculation, and the student model predicts x0 (t) and x0 (t-k) When LORA parameters are inserted for calculation. Through this engineering optimization solution, the disclosed embodiment can reduce the video memory usage of distillation training by nearly half during the training process, further reducing the difficulty of distillation training of the Wensheng graph model.
[0073] The training method of the Wensheng graph model of this embodiment, by adopting the above-mentioned technical scheme and training based on the teacher model, can significantly shorten the time consumption of raw images on the basis of improving the raw image quality to a certain extent, and improve the training efficiency of the Wensheng graph model. Moreover, without affecting the raw image effect, the fused target Wensheng graph model will also have the advantages of improved raw image quality and reduced number of inference sampling steps. Therefore, the technical scheme of this embodiment can train the Wensheng graph model quickly and efficiently, and the trained Wensheng graph model can also inherit the characteristics of the teacher model and have high-quality raw image capabilities. In short, the above-mentioned technical scheme of this embodiment can not only accurately and efficiently train the Wensheng graph model; but also can effectively shorten the raw image steps of the Wensheng graph model obtained by training, reduce the raw image time consumption, and improve the raw image efficiency on the basis of improving the raw image quality of the Wensheng graph model obtained by training.
[0074] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure; Figure 3 As shown, this embodiment provides a training device 300 for a text graph model, including:
[0075] An acquisition module 301 is used to acquire a distillation training data set from a pre-training data set; the distillation training data set includes a plurality of samples;
[0076] The training module 302 is used to train a second text graph model based on the distilled training data set and the pre-trained first text graph model; the second text graph model adds an attention mechanism linear layer on the basis of the first text graph model; during the training process, the parameters of the structure of the second text graph model and the first text graph model are frozen.
[0077] The training device 300 of the culture graph model of this embodiment realizes the implementation principle and technical effect of the training of the culture graph model by adopting the above modules, which is the same as that of the above related method embodiments. For details, please refer to the records of the above related method embodiments, which will not be repeated here.
[0078] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure; Figure 4 As shown, the training device 400 of the text graph model of this embodiment, in the above Figure 3 Based on the technical solutions of the embodiments shown in the figure, the technical solutions of the present disclosure are further described in more detail. Figure 4 As shown, the training device 400 of the text graph model of this embodiment includes the above Figure 3 The modules with the same name and function are shown as: an acquisition module 401 and a training module 402 .
[0079] The training module 402 of this embodiment is used to:
[0080] Obtaining samples from the distillation training data set; the samples include training pictures and corresponding text description information;
[0081] Noise is added to the training image in the sample to obtain a first noise image; for example, noise is added to the training image in the sample at a random sampling time step to obtain the first noise image, that is, the first noise image is a noise image at a random sampling time step.
[0082] The first Wensheng graph model is used to solve the first noise image to obtain a second noise image; for example, the first Wensheng graph model can be used to solve the first noise image of the random sampling time step for a preset number of steps to obtain the second noise image, and the second noise image is the noise image of the random sampling time step minus the preset number of steps.
[0083] The second cultural graph model is trained based on the first noise picture and the second noise picture.
[0084] Further optionally, in one embodiment of the present disclosure, the training module 402 is used to:
[0085] Using the second cultural image model, based on the first noisy image, predict a first noise-free image;
[0086] Using the second cultural image model, based on the second noisy image, predicting the second noise-free image;
[0087] Based on the first noise-free image and the second noise-free image, a low-rank adaptive parameter of a linear layer of an attention mechanism in the second cultural graph model is updated.
[0088] Further optionally, in one embodiment of the present disclosure, the training module 402 is used to:
[0089] Calculating a distance between the first noise-free image and the second noise-free image as a loss function;
[0090] With the convergence of the loss function as a goal, the low-rank adaptive parameters of the linear layer of the attention mechanism in the second text-generated graph model are adjusted.
[0091] Further optionally, if Figure 4 As shown, in one embodiment of the present disclosure, the training device 400 of the culture graph model further includes:
[0092] The updating module 403 is used to update the exponential moving average parameters based on the low-rank adaptive parameters of the linear layer of the attention mechanism during the training process.
[0093] Further optionally, in one embodiment of the present disclosure, the updating module 403 is used to:
[0094] During the training process, the exponential sliding average parameter is updated based on the low-rank adaptive parameter of the linear layer of the attention mechanism at every preset training steps.
[0095] Further optionally, if Figure 4 As shown, in one embodiment of the present disclosure, the training device 400 of the culture graph model further includes:
[0096] The construction module 404 is used to construct a target text graph model based on the second text graph model and the exponential moving average parameter obtained after the training, and the target text graph model is used to perform the reasoning task.
[0097] Further optionally, in one embodiment of the present disclosure, the construction module 404 is used to:
[0098] The low-rank adaptive parameters of the linear layer of the attention mechanism of the second text-generated graph model obtained after the training are replaced with the exponential sliding average parameters to obtain the target text-generated graph model.
[0099] Further optionally, in one embodiment of the present disclosure, the acquisition module 401 is used to:
[0100] Obtaining at least one of picture aesthetics, picture clarity, and picture resolution of a picture in each training data in the pre-training data set;
[0101] The distilled training data set is distilled from the pre-training data set based on at least one of the image aesthetics, image clarity and image resolution of the images in each training data in the pre-training data set.
[0102] The training device 400 of the culture graph model of this embodiment realizes the training of the culture graph model by adopting the above modules, and the implementation principle and technical effect are the same as those of the above related method embodiments. For details, please refer to the records of the above related method embodiments, which will not be repeated here.
[0103] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0104] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0105] Figure 5A schematic block diagram of an example electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0106] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0107] A number of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 505, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0108] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as the above-mentioned methods of the present disclosure. For example, in some embodiments, the above-mentioned methods of the present disclosure may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the above-mentioned methods of the present disclosure described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the above-mentioned methods of the present disclosure in any other appropriate manner (e.g., by means of firmware).
[0109] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0110] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0111] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0112] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0113] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0114] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0115] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0116] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A training method for a text graph model, comprising: Get a distilled training dataset from the pre-training dataset; The distillation training data set includes multiple samples; Based on the distilled training data set and the pre-trained first text graph model, a second text graph model is trained; the second text graph model adds an attention mechanism linear layer on the basis of the first text graph model; During the training process, the parameters of the structure of the same part of the second text graph model and the first text graph model are frozen.
2. The method according to claim 1, wherein: Based on the distilled training data set and the pre-trained first text-generated graph model, training the second text-generated graph model includes: Obtaining samples from the distillation training data set; the samples include training pictures and corresponding text description information; Adding noise to the training image in the sample to obtain a first noise image; Using the first Wensheng graph model, solving the first noise image to obtain a second noise image; The second cultural graph model is trained based on the first noise picture and the second noise picture.
3. The method according to claim 2, wherein: Training the second cultural graph model based on the first noise picture and the second noise picture includes: Using the second cultural image model, based on the first noisy image, predict a first noise-free image; Using the second cultural image model, based on the second noisy image, predicting the second noise-free image; Based on the first noise-free image and the second noise-free image, a low-rank adaptive parameter of a linear layer of an attention mechanism in the second cultural graph model is updated.
4. The method according to claim 3, wherein: Based on the first noise-free image and the second noise-free image, updating the low-rank adaptive parameters of the attention mechanism linear layer in the second cultural graph model includes: Calculating a distance between the first noise-free image and the second noise-free image as a loss function; With the convergence of the loss function as a goal, the low-rank adaptive parameters of the linear layer of the attention mechanism in the second text-generated graph model are adjusted.
5. The method according to any one of claims 1 to 4, wherein: The method further comprises: During the training process, the exponential sliding average parameters are updated based on the low-rank adaptive parameters of the linear layer of the attention mechanism.
6. The method according to claim 5, wherein: During the training process, based on the low-rank adaptive parameters of the linear layer of the attention mechanism, the exponential sliding average parameters are updated, including: During the training process, the exponential sliding average parameter is updated based on the low-rank adaptive parameter of the linear layer of the attention mechanism at every preset training steps.
7. The method according to claim 5, wherein: The method further comprises: Based on the second culture graph model obtained after the training and the exponential moving average parameter, a target culture graph model is constructed, and the target culture graph model is applied to perform the reasoning task.
8. The method according to claim 7, wherein: Based on the second cultural graph model obtained after the training and the exponential sliding average parameter, a target cultural graph model is constructed, including: The low-rank adaptive parameters of the linear layer of the attention mechanism of the second text-generated graph model obtained after the training are replaced with the exponential sliding average parameters to obtain the target text-generated graph model.
9. The method according to any one of claims 1-4 and 6-8, wherein: Get the distilled training dataset from the pre-training dataset, including: Obtaining at least one of picture aesthetics, picture clarity, and picture resolution of a picture in each training data in the pre-training data set; The distilled training data set is distilled from the pre-training data set based on at least one of the image aesthetics, image clarity and image resolution of the images in each training data in the pre-training data set.
10. A training device for a text graph model, comprising: An acquisition module is used to obtain a distillation training data set from a pre-training data set; The distillation training data set includes multiple samples; A training module, used for training a second text-generated graph model based on the distilled training data set and the pre-trained first text-generated graph model; the second text-generated graph model adds an attention mechanism linear layer on the basis of the first text-generated graph model; During the training process, the parameters of the structure of the same part of the second text graph model and the first text graph model are frozen.
11. The device according to claim 10, wherein: The training module is used to: Obtaining samples from the distillation training data set; the samples include training pictures and corresponding text description information; Adding noise to the training image in the sample to obtain a first noise image; Using the first Wensheng graph model, solving the first noise image to obtain a second noise image; The second cultural graph model is trained based on the first noise picture and the second noise picture.
12. The device according to claim 11, wherein The training module is used to: Using the second cultural image model, based on the first noisy image, predict a first noise-free image; Using the second cultural image model, based on the second noisy image, predicting the second noise-free image; Based on the first noise-free image and the second noise-free image, a low-rank adaptive parameter of a linear layer of an attention mechanism in the second cultural graph model is updated.
13. The device according to claim 12, wherein: The training module is used to: Calculating a distance between the first noise-free image and the second noise-free image as a loss function; With the convergence of the loss function as a goal, the low-rank adaptive parameters of the linear layer of the attention mechanism in the second text-generated graph model are adjusted.
14. The device according to any one of claims 10 to 13, wherein: The device also includes: The updating module is used to update the exponential sliding average parameters based on the low-rank adaptive parameters of the linear layer of the attention mechanism during the training process.
15. The device according to claim 14, wherein: The update module is used to: During the training process, the exponential sliding average parameter is updated based on the low-rank adaptive parameter of the linear layer of the attention mechanism at every preset training steps.
16. The device according to claim 14, wherein: The device also includes: A construction module is used to construct a target text graph model based on the second text graph model and the exponential sliding average parameter obtained after the training is completed, and the target text graph model is used to perform the reasoning task.
17. The device according to claim 16, wherein: The building blocks are used to: The low-rank adaptive parameters of the linear layer of the attention mechanism of the second text-generated graph model obtained after the training are replaced with the exponential sliding average parameters to obtain the target text-generated graph model.
18. The device according to any one of claims 10-13 and 15-17, wherein: The acquisition module is used to: Obtaining at least one of picture aesthetics, picture clarity, and picture resolution of a picture in each training data in the pre-training data set; The distilled training data set is distilled from the pre-training data set based on at least one of the image aesthetics, image clarity and image resolution of the images in each training data in the pre-training data set.
19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
21. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.