Data processing, intelligent interaction, model training and development method, equipment and medium

By introducing residual structure and jump connection tuning subnetwork into the generative model, the problems of low training efficiency and insufficient reliability of the generative model are solved, and more efficient training and more reliable data processing are achieved.

CN120197650APending Publication Date: 2025-06-24ZHEJIANG ALIBABA ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311736419.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing generative models have problems of inefficiency and inadequate reliability in training and data processing, especially as the model size is growing, training is expensive and training is ineffective.

Method used

A pre-trained generative model is adopted, which contains a residual structure. The tuning subnet is formed through jump connections between the encoding layer and the decoding layer. The residual structure is efficiently trained through these tuning subnets, thereby improving the prediction reliability of the generative model.

Benefits of technology

Through this method, the training efficiency and data processing reliability of the generative model can be significantly improved, the training cost can be reduced, and the adaptability of the model in downstream tasks can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197650A_ABST
    Figure CN120197650A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, an intelligent interaction method, a model training method, a model development method, equipment and a medium. The data processing method comprises the steps that a pre-trained generative model is acquired, the generative model comprises a residual structure, the residual structure comprises jump connection between a coding layer and a decoding layer, and at least one tuning sub-network is formed between the coding layer and the decoding layer; and at least inputting the description data into the generative model to obtain generated data corresponding to the description data. In the embodiment of the invention, the residual structure comprises the jump connection between the coding layer and the decoding layer, and the at least one tuning sub-network is formed between the coding layer and the decoding layer, so that the at least one tuning sub-network can efficiently train the residual structure by being connected between the coding layer and the decoding layer which are connected in the jump manner; and a generative model with higher prediction reliability is obtained, so that more reliable data processing can be executed by adopting the generative model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and in particular, to a data processing, intelligent interaction, model training and development method, device, and medium. Background Art

[0002] In the field of artificial intelligence, generative models such as large language models are constructed through generative networks. For example, diffusion generative networks have effective network structures and can exhibit excellent generalization capabilities after being trained with samples. Based on the basic generative model pre-trained with large-scale data, fine-tuning training can be performed for a series of downstream tasks and applications such as image generation, image editing, and image transformation.

[0003] As the number of parameters of the generative model continues to grow, the cost of methods for adapting such models to various tasks becomes quite expensive. If a low-cost training method is adopted, the training effect is poor. Correspondingly, when using the trained generative model for data processing, the reliability is poor. Therefore, there is room for improvement in both the training efficiency of the generative model and the reliability of using the generative model for data processing. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a data processing, intelligent interaction, model training and development method, device, and medium to at least partially solve the above problems.

[0005] According to a first aspect of the embodiments of the present invention, a data processing method is provided, including: obtaining a pre-trained generative model, where the generative model includes a residual structure, the residual structure includes a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer; inputting at least description data into the generative model to obtain generated data corresponding to the description data.

[0006] According to a second aspect of the embodiments of the present invention, an intelligent interaction method is provided, including: obtaining first description data input in an interaction interface; inputting at least the first description data into a pre-trained generative model to obtain first generated data corresponding to the first description data, where the first generated data includes at least one of an image, a video frame, and a video, the generative model includes a residual structure, the residual structure includes a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer. Displaying the first generated data in the question-and-answer interface.

[0007] According to a third aspect of an embodiment of the present invention, there is provided a model training method, including: obtaining a generative network for initial training, the generative network including an encoder and a decoder, and a skip connection being formed between an encoding layer of the encoder and a decoding layer of the decoder; connecting at least one tuning sub-network of the skip connection between the encoding layer and the decoding layer of the generative network for initial training to obtain a generative network to be fine-tuned, wherein an output of the encoding layer is connected to an input of the at least one tuning sub-network, and an output of the at least one tuning sub-network is connected to an input of the decoding layer; based on training samples, performing fine-tuning training on the generative network to be fine-tuned to obtain a generative model.

[0008] According to a fourth aspect of an embodiment of the present invention, there is provided an application development method, including: creating a user interface module of an application program, the user interface module being configured to generate description data at least based on user operation data, and return presentation data based on generated data of the description data; obtaining a call interface of a generative model, the generative model being obtained according to the model training method described in the third aspect, the call interface being configured to return the generated data when being called; embedding at least the call interface of the generative model into the user interface module.

[0009] According to a fifth aspect of an embodiment of the present invention, there is provided a data processing device including: an acquisition module, acquiring a pre-trained generative model, the generative model including a residual structure, the residual structure including a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network being formed between the encoding layer and the decoding layer; a generation module, inputting at least description data into the generative model to obtain generated data corresponding to the description data.

[0010] According to a sixth aspect of an embodiment of the present invention, there is provided an intelligent interaction device including: an acquisition module, acquiring first description data input in an interaction interface; a generation module, inputting at least the first description data into a pre-trained generative model to obtain first generated data corresponding to the first description data, wherein the first generated data includes at least one of an image, a video frame, and a video, the generative model including a residual structure, the residual structure including a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network being formed between the encoding layer and the decoding layer; a display module, displaying the first generated data in the question-and-answer interface.

[0011] According to a seventh aspect of an embodiment of the present invention, there is provided a model training device including: a module acquisition module that acquires a generative network trained initially, the generative network including an encoder and a decoder, and a skip connection being formed between an encoding layer of the encoder and a decoding layer of the decoder; a model adjustment module that connects at least one tuning sub-network of the skip connection between the encoding layer and the decoding layer of the initially trained generative network to obtain a generative network to be fine-tuned, wherein an output of the encoding layer is connected to an input of the at least one tuning sub-network, and an output of the at least one tuning sub-network is connected to an input of the decoding layer; and a model training module that performs fine-tuning training on the generative network to be fine-tuned based on training samples to obtain a generative model.

[0012] According to an eighth aspect of an embodiment of the present invention, there is provided an application development device including: a creation module that creates a user interface module of an application program, the user interface module being configured to generate description data based at least on user operation data and return presentation data based on generated data of the description data; an acquisition module that acquires a call interface of a generative model obtained according to the model training method described in the third aspect, the call interface being configured to return the generated data when being called; and an embedding module that embeds at least the call interface of the generative model into the user interface module.

[0013] According to a ninth aspect of an embodiment of the present invention, there is provided an electronic device including: a processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used for storing at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method described in the first aspect.

[0014] According to a tenth aspect of an embodiment of the present invention, there is provided a computer storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements the method described in the first aspect.

[0015] In this embodiment, the pre-trained generative model includes a residual structure, the residual structure includes a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer. Therefore, the at least one tuning sub-network can efficiently train the residual structure by being connected between the encoding layer and the decoding layer of the skip connection, and further obtain a generative model with higher prediction reliability, so that a more reliable data process can be performed by using the generative model. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0017] Figure 1 It is a schematic structural diagram of a generative network according to some examples.

[0018] Figure 2A It is a flowchart of the steps of a data processing method according to some other embodiments of the present invention.

[0019] Figure 2B It is a partial schematic structural diagram of a generative network for some examples of the embodiment in FIG. 2.

[0020] Figure 2C It is a partial schematic structural diagram of a generative network for some other examples of the embodiment in FIG. 2.

[0021] Figure 2D It is a flowchart of the steps of a data processing method according to some other embodiments of the present invention.

[0022] Figure 2E It is a partial schematic structural diagram of a generative network for some other examples of the embodiment in FIG. 2.

[0023] Figure 3 It is a flowchart of the steps of an intelligent interaction method according to some other embodiments of the present invention.

[0024] Figure 4A It is a flowchart of the steps of a model training method according to some other embodiments of the present invention.

[0025] Figure 4B It is a flowchart of the steps of a model training method according to some other embodiments of the present invention.

[0026] Figure 5 It is a flowchart of the steps of an application development method according to some other embodiments of the present invention.

[0027] Figure 6 It is a schematic block diagram of a data processing device according to some other embodiments of the present invention.

[0028] Figure 7 It is a schematic block diagram of an intelligent interaction device according to some other embodiments of the present invention.

[0029] Figure 8 It is a schematic block diagram of a model training device according to some other embodiments of the present invention.

[0030] Figure 9 Schematic block diagram of an application development device according to other embodiments of the present invention.

[0031] Figure 10 Schematic structural diagram of an electronic device according to other embodiments of the present invention. Detailed implementation manners

[0032] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present invention.

[0033] The following further illustrates the specific implementation of the embodiments of the present invention with reference to the accompanying drawings of the embodiments of the present invention.

[0034] Generative models such as large language models are constructed through generative networks. For example, Figure 1 the generative network shown includes an encoder 110 and a decoder 120. The output of the encoder 110 can be directly connected to the input of the decoder 120, and the output of the encoder 110 can also be connected to the input of the decoder 120 through a context alignment layer (not shown) that adopts an attention mechanism. Figure 1 The generative network of can be a generative diffusion network, which has a structure of skip connections (SC) between the encoding layer and the decoding layer. The skip connections can be implemented as one or more. For example, Figure 1 the skip connections #1, skip connections #2,..., skip connections #N shown. Each skip connection forms a residual structure between the encoder 110 and the decoder 120, so that the generative network obtains better generalization ability. However, as the number of parameters of the generative model scale continues to grow, the cost of methods for adapting such models to various tasks becomes quite expensive. Each embodiment of the present invention provides a series of technical solutions that can improve the efficiency of fine-tuning training and reduce the fine-tuning training overhead of the generative network.

[0035] Figure 2AThe figure is a flowchart of the steps of a data processing method according to some other embodiments of the present invention. The solution of this embodiment can be applied to any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.) and PC machines, etc. For example, in the training phase, a computing device (such as a data center) configured with a CPU (an example of a processing unit) + GPU (an example of an acceleration unit) architecture can be used to train an encoder-decoder model based on training samples. A computing device such as a data center can be deployed in a cloud server such as a private cloud, a private cloud, or a hybrid cloud. Correspondingly, in the inference phase, a computing device configured with a CPU (an example of a processing unit) + GPU (an example of an acceleration unit) architecture can also be used for inference operations. Specifically, Figure 2A The data processing method includes:

[0036] S210: Obtain a pre-trained generative model, the generative model includes a residual structure, the residual structure includes a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer.

[0037] It should be understood that the generative model can be obtained by training a generative network, and the generative network can be a generative diffusion network. The generative model can be a generative model for generating images according to a description text, or a text processing model for generating text according to a text, for example, a knowledge Q&A model or a translation model.

[0038] It should also be understood that the generative model can include an encoder and a decoder, and there can be a residual structure with one or more skip connections (Skip Connection, SC) between the encoder and the decoder, for example, a U-Net structure. The network layer in the encoder is the encoding layer, and the network layer in the decoder is the decoding layer, and the encoding layer and the decoding layer connected by the skip connection correspond. An attention mechanism or the like can be adopted between the encoder and the decoder to form a context alignment structure including one or more network layers (for example, a fully connected layer).

[0039] S220: Input at least the description data into the generative model to obtain generated data corresponding to the description data.

[0040] It should be understood that the description data can describe text or can be converted into description data that describes text. For example, the voice description data can be converted into corresponding description text and input into the generative model. For the specific conversion method, the voice description data can be input into a pre-trained text conversion model to obtain the description text, or the text conversion model can be connected to the input side of the encoder of the generative model. Alternatively, the generative model can also be trained to receive voice description data as input, that is, when the generative model is trained end-to-end, the input side sample of the training sample is a voice sample.

[0041] It should also be understood that the generated data can be a descriptive text (for example, in a knowledge question answering scenario or a translation scenario); the generated data can also be a picture or a video frame, or a video based on a video frame sequence. This depends on the sample type of the end-to-end training used in the training phase of the generative model. If the output side sample is a video frame sequence, the generated data in the inference phase is a video; if the output side sample is a picture or a video frame, the generated data in the inference phase is a picture or a video frame.

[0042] In this embodiment, the pre-trained generative model includes a residual structure, and the residual structure includes a jump connection between the encoding layer and the decoding layer. At least one tuning sub-network is formed between the encoding layer and the decoding layer. Therefore, at least one tuning sub-network can efficiently train the residual structure by connecting to the encoding layer and the decoding layer of the jump connection, thereby obtaining a generative model with higher prediction reliability, so that more reliable data processing can be performed using the generative model.

[0043] In other examples, various conditional image data are input into each conditional sub-network respectively. The conditional image can be a description text and multiple additional conditions. Different image conditions are used to describe different types of image features. The multiple types of image feature data include but are not limited to different image features such as image edge detection features, image depth features, image segmentation features, image color features, image extension features, and image restoration features. Specifically, the description text is input into the encoder of the generative model, and each additional condition is input into the corresponding tuning sub-network of multiple jump connections to obtain the output generated image.

[0044] Specifically, in Figure 2B In the residual structure (U-Net structure) shown, the output of the encoding layer is connected to the input of each tuning sub-network 30, and the output of each tuning sub-network is connected to the input of the decoding layer. In other words, the output of each tuning sub-network is fused with the input of the encoding layer through the jump connection as the input of the decoding layer of the jump connection.

[0045] Right now:

[0046] Among them, Oj SC is the output of the j-th layer in the residual structure, and X N-j is the input of the skip connection, Tj is the output of the tuning sub-network, and N is the total number of layers of the encoder in the residual structure.

[0047] In some other examples, the residual structure further includes at least one conditional sub-network, and at least one conditional sub-network corresponds to at least one tuning sub-network. As an example of at least inputting description data into a generative model to obtain generated data corresponding to the description data, the description data can be input into the encoding layer of the generative model, the conditional data can be input into at least one conditional sub-network of the generative model, and the generated data corresponding to the description data can be obtained from the decoding layer of the generative model.

[0048] Specifically, at least one tuning sub-network respectively corresponds to at least one conditional sub-network. The first output of each conditional sub-network is connected between the input of the corresponding tuning sub-network and the input of the encoding layer, and the second output of each conditional sub-network is connected between the output of the corresponding tuning sub-network and the input of the decoding layer.

[0049] More specifically, after the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, it is processed by the conditional weight of the conditional sub-network in the multiple conditional sub-networks and output to the input of the decoding layer.

[0050] In some other examples, the description data is description text, and the generated data includes multiple video frames. The data processing method further includes: performing temporal encoding processing on the multiple video frames to obtain a generated video of the description text.

[0051] Figure 2D Illustrates a data processing method according to some other embodiments of the present invention.

[0052] S280: Obtain a pre-trained generative model.

[0053] S290: Input the description data into the encoding layer of the generative model, input the conditional data into at least one conditional sub-network of the generative model, and obtain the generated data corresponding to the description data from the decoding layer of the generative model.

[0054] It should be understood that Figure 2D the embodiments of Figure 2A are the same as or similar to those of

[0055] Further, the generative model of this embodiment may include a tuning sub-network with a conditional sub-network. For example, as Figure 2CAs shown, at least one tuning sub-network 30 corresponds to at least one conditional sub-network 40 respectively. The first output of each conditional sub-network 40 is connected between the input of the corresponding tuning sub-network 30 and the input of the encoding layer, and the second output of each conditional sub-network 40 is connected between the output of the corresponding tuning sub-network 30 and the input of the decoding layer.

[0056] That is: Where Oj CSC is the output processed by the tuning sub-network of the j-th layer in the residual structure; m is the m-th condition, traversing values between 1 and M; X N-j is the input; Tj m is the tuning sub-network of the m-th condition of the j-th layer, α m is the conditional weight coefficient of the m-th condition; Cj is the input encoded by the m-th condition of the j-th layer, and Cj is the conditional set of {Cj0…Cjm}.

[0057] Among them, the M conditions are supplementary conditional data for the input data of the encoder during fine-tuning training. For example, when the input of the encoder is descriptive text, the M conditions can be conditional images that supplement the descriptive text.

[0058] Furthermore, as Figure 2E shown, for each conditional sub-network (for example, M conditional sub-networks), the number of conditional sub-networks can be the same as the number of tuning sub-networks, that is to say, the m-th conditional sub-network corresponds to the m-th tuning sub-network.

[0059] More specifically, N skip connections are respectively formed between multiple encoding layers of the encoder and multiple decoding layers of the decoder. Each skip connection corresponds to M tuning sub-networks and M conditional sub-networks, and the m-th conditional sub-network corresponds to the m-th tuning sub-network. The m-th conditional sub-network includes N network layers sequentially arranged from the input side to the output side. The n-th network layer corresponds to the n-th skip connection, and the output of the n-th network layer is connected to the input of the m-th tuning sub-network of the n-th skip connection. Among them, N is an integer greater than 1, M is an integer greater than 1, 1 ≤ n ≤ N, and 1 ≤ m ≤ M.

[0060] Further, the output of the n-th network layer is connected to the encoding layer of the n-th skip connection. It should be understood that each network layer can be implemented by a linear layer and an activation layer such as a downsampling layer, and the data dimension of the output of the n-th network layer is aligned or consistent with the data dimension of the encoding layer of the n-th skip connection, thereby improving the reliability of the training process of the generative model and also improving the reliability of data processing using the generative model.

[0061] In some variant examples of the data processing method, the generated image of the generative model including the tuning sub-network without the conditional sub-network can be used as the conditional image of the generative model including the tuning sub-network with the conditional sub-network. Thus, the conditional image is generated through the description text, and further, the generative model including the tuning sub-network with the conditional sub-network is used to obtain the generated image from both the conditional image and the further description text.

[0062] In some other variant examples of the data processing method, the generated image of the generative model including the tuning sub-network without the conditional sub-network can be used as the input data of the encoder of the generative model including the tuning sub-network with the conditional sub-network, so as to further obtain the generated image that better matches the description text by using the attached conditional sub-network.

[0063] The following will be combined with Figure 3 Describe in detail the step flowchart of the intelligent interaction method based on the generative model of some other embodiments of the present invention. Figure 3 The intelligent interaction method includes:

[0064] S310: Obtain the first description data input in the interaction interface.

[0065] S320: Input at least the first description data into the pre-trained generative model to obtain the first generated data corresponding to the first description data, where the first generated data includes at least one of an image, a video frame, and a video. The generative model includes a residual structure, and the residual structure includes a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer.

[0066] S330: Display the first generated data in the Q&A interface.

[0067] In this embodiment, by receiving the description data input by the user through the user interface and displaying the data processing result to the user, fast interaction is realized. In addition, the type of the description data input by the user is not limited in this embodiment, which is convenient for the user to input various types of description data such as text description data or voice description data. In addition, the generated data presented by the user interface includes but is not limited to text, pictures, or videos. Therefore, this embodiment provides an intelligent information service for the user by combining the convenient interaction method of the user interface and the data processing ability of the generative model. For example, it can be applied to intelligent products such as intelligent assistants and virtual experts.

[0068] In some examples, the intelligent interaction method further includes: obtaining the second description data for modifying the first generated data from the interaction interface; inputting at least the second description data into the pre-trained generative model to generate the second generated data corresponding to the second description data as the modification result of the first generated data.

[0069] Through the above data modification process, after the initial generation process of the first generated data, using the first description data as context data and based on the second description data as the modification description data, the first generated data is adjusted to obtain the second generated description.

[0070] It should be understood that the types of the first generated data and the second generated data here can be the same. For example, modified speech data is obtained by modifying speech data, modified text data is obtained by modifying text data, and modified video data is obtained by modifying video data.

[0071] Furthermore, the types of the first description data and the second description data can be the same or different. For example, in some other examples, the first generated data is a picture or a video frame. The second description data for modifying the first generated data obtained from the interaction interface includes: obtaining a modification mark for the picture or video from the interaction interface; identifying the modification mark to generate the second description data, or combining the modification mark and the modification description text of the picture or video to generate the second description data.

[0072] That is to say, the second description data of the same type as the first description data can be used as a supplement and modification hint for the first description data, or the second description data of a different type from the first description data can be used as a supplement and modification hint for the first description data. For example, when the first description data is text data, text data can be further supplemented to modify the first generated data, or speech data / operation data different from the text data can be used as a supplement or modification hint for the text data. The operation data can be the target area marked in the above picture or video frame. That is to say, the target area is the area to be modified or the area with a larger modification weight. Through such an operation, the second description data combined with the operation data can further improve the accuracy and reliability of the modification because it provides more information about the modification weight.

[0073] Specifically, Figure 4A A model training method according to some embodiments of the present invention is shown. Figure 4A The model training method includes:

[0074] S410: Obtain a generative network for initial training. The generative network includes an encoder and a decoder, and a skip connection is formed between the encoding layer of the encoder and the decoding layer of the decoder.

[0075] It should be understood that the generative network for initial training can be a generative network obtained through pre-training. For example, a generative diffusion network.

[0076] S420: Connect at least one tuning sub-network of the skip connection between the encoding layer and the decoding layer of the initially trained generative network to obtain a generative network to be fine-tuned, wherein the output of the encoding layer is connected to the input of at least one tuning sub-network, and the output of at least one tuning sub-network is connected to the input of the decoding layer.

[0077] It should be understood that at least one tuning sub-network can be connected between the encoding layer and the decoding layer of some or all of the skip connections, and the structure formed by the tuning sub-network can be referred to as SC-Tuner. The tuning sub-network includes at least a linear layer and an activation layer. For example, the tuning sub-network can be formed by a first linear layer, an activation layer, and a second linear layer, and the activation layer is arranged between the first linear layer and the second linear layer. In addition, other activation layers can be added on the input side of the first thread layer and / or the output side of the second linear layer, and other linear layers can also be added at any position of the tuning sub-network. The above linear layer can be a fully connected layer or a convolutional layer. When the output of the generative network to be fine-tuned is image data, the linear layer in the tuning sub-network can be two-dimensional data. When the output of the generative network to be fine-tuned is text data, the linear layer in the tuning sub-network can be one-dimensional data. This is because the tuning sub-network is connected between the encoding layer and the decoding layer of the skip connection and only needs to have the same dimension as the output data of the decoder.

[0078] S430: Based on the training samples, perform fine-tuning training on the generative network to be fine-tuned to obtain a generative model.

[0079] It should be understood that the fine-tuning training of the generative network to be fine-tuned can be supervised training, and the training samples can include input-side samples and output-side samples. The input-side samples can be text data injected into a specific encoding layer of the encoder, and the input-side samples can also include image samples input to the encoder. The output-side samples can be text data or image data such as pictures or video frames.

[0080] In the solution of the embodiment of the present invention, at least one tuning sub-network of the generative network to be fine-tuned is connected between the encoding layer and the decoding layer connected by the skip structure. The output of the encoding layer is connected to the input of at least one tuning sub-network, and the output of at least one tuning sub-network is connected to the input of the decoding layer, so that when the generative network is fine-tuned and trained, the skip structure is further optimized through at least one tuning sub-network, improving the efficiency of the fine-tuning training. In addition, at least one tuning sub-network is not in the encoder or the decoder, and there is no need to adjust the parameters too much during the fine-tuning training, saving the computational overhead of the fine-tuning training.

[0081] Further, as an example of fine-tuning a generative network to be fine-tuned based on training samples, the generative network to be fine-tuned can be iteratively trained multiple times based on the training samples to obtain a generative model. Among them, in each iterative training, the difference between the output of the input-side samples in the training samples through the forward propagation of the generative network to be fine-tuned and the input-side samples in the training samples is determined, and while keeping the parameters of the encoder unchanged, the parameters in the decoder and at least one tuning sub-network are adjusted through the backpropagation of the difference in the decoder and at least one tuning sub-network. That is, in each iterative training, there is no need to train the parameters in the encoder, and the training processes of the encoder and at least one tuning sub-network are decoupled, thereby significantly reducing the computational overhead in fine-tuning training.

[0082] Specifically, the input-side samples can be descriptive texts, and the output-side samples can be generated images. For example, the descriptive texts are used to describe the generated images, so that the trained generative model can more effectively adapt to the downstream tasks of the initially trained generative network.

[0083] For example, taking the Stable Diffusion generative network as an example, residual structures such as U-Net often contain 12 skip connections. Tuning sub-networks are added to each skip connection, which can be efficiently trained for different tasks. Compared with other fine-tuning training methods such as LoRA, it can be trained with fewer tuning parameters and less memory consumption. That is to say, when applying each tuning sub-network in efficient generative tuning or few-shot task tuning, it can be applied to rapid customized transfer training under the condition of few samples, such as learning specific character features, specific image styles and other image features.

[0084] In the case where the tuning sub-network does not have an attached conditional sub-network (i.e., SC-Tuner), as an example of fine-tuning the generative network to be fine-tuned, the input-side samples can be used as the input of the encoder, and the output-side samples can be used as the output of the decoder to adjust the network parameters in at least one tuning sub-network to obtain a generative model, thereby reliably training the tuning sub-network without an attached conditional sub-network. That is to say, the constraints on the input side and the output side of the tuning sub-network come from the encoding layer and the decoding layer of the skip connection corresponding to the tuning sub-network. Correspondingly, the input-side samples at least include descriptive texts, and the output-side samples at least include image samples, and the image samples include but are not limited to pictures and video frames.

[0085] Further, after the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, it is processed by the conditional weights of the conditional sub-network among multiple conditional sub-networks and output to the input of the decoding layer. It should be understood that since the importance of the conditions corresponding to multiple conditional sub-networks can be different, the conditional weight of a conditional sub-network among multiple conditional sub-networks indicates the importance in that conditional sub-network. In addition, the conditional weights of the conditional sub-networks can also be implemented as linear layers such as fully connected layers, connected between each tuning sub-network and the corresponding encoding layer. At least one tuning sub-network and its conditional sub-network correspond to skip connections, and the tuning sub-network with the attached conditional sub-network can also be referred to as a CSC-Tuner.

[0086] Further, the functions of different conditional sub-networks are different, that is, multiple conditional sub-networks are respectively used to output various image feature data. For example, the various image feature data include but are not limited to different image features such as indicating image edge detection features, image depth features, image segmentation features, image color features, image expansion features, image restoration features, etc.

[0087] Without loss of generality, in the model training framework, components for fine-tuning the training residual structure can be configured. By inputting skip connections and related conditions into the SC-Tuner (i.e., the tuning sub-network without an attached conditional sub-network) and the CSC-Tuner (i.e., the tuning sub-network with an attached conditional sub-network), and concatenating them with the input in the decoding layer of the original corresponding residual structure and sending it to the next decoding layer.

[0088] That is, g j+1 = G j [O j (X N-j-1 + C j ) ; g j , where gj is the input to the decoding layer of the residual structure of the j-th layer; Gj is the output after the decoding layer operation of the residual structure of the j-th layer; Oj is the output after being processed by the tuning sub-network of the j-th layer in the residual structure.

[0089] Multiple skip connections 130 are respectively formed between multiple encoding layers of the encoder 120 and multiple decoding layers of the decoder 120. The conditional sub-network 40 includes multiple network layers sequentially arranged from the input side to the output side, and the multiple network layers respectively correspond to multiple skip connections.

[0090] Specifically, as Figure 2E shown, for each conditional sub-network (for example, M conditional sub-networks), the number of conditional sub-networks can be the same as the number of tuning sub-networks, that is, the m-th conditional sub-network corresponds to the m-th tuning sub-network.

[0091] More specifically, N skip connections are respectively formed between multiple encoding layers of the encoder and multiple decoding layers of the decoder. Each skip connection corresponds to M tuning sub-networks and M conditional sub-networks, and the m-th conditional sub-network corresponds to the m-th tuning sub-network. The m-th conditional sub-network includes N network layers sequentially arranged from the input side to the output side. The n-th network layer corresponds to the n-th skip connection, and the output of the n-th network layer is connected to the input of the m-th tuning sub-network of the n-th skip connection. Wherein, N is an integer greater than 1, M is an integer greater than 1, 1 ≤ n ≤ N, and 1 ≤ m ≤ M.

[0092] Furthermore, the output of the n-th network layer is connected to the encoding layer of the n-th skip connection. It should be understood that each network layer can be implemented by a linear layer such as a downsampling layer and an activation layer, etc. The data dimension of the output of the n-th network layer is aligned or consistent with the data dimension of the encoding layer of the n-th skip connection, thereby improving the reliability of the training process of the generative model.

[0093] That is to say, the output of the m-th conditional sub-network of the n + 1-th skip connection is connected to the input of the n + 1-th network layer to obtain the m-th conditional sub-network of the n-th skip connection. That is, the output of the n + 1-th network layer serves as the output of the m-th conditional sub-network of the n-th skip connection.

[0094] Figure 4B A model training method according to some other embodiments of the present invention is shown, including:

[0095] S470: Obtain a generative network in initial training. The generative network includes an encoder and a decoder, and a skip connection is formed between the encoding layer of the encoder and the decoding layer of the decoder.

[0096] S480: Connect at least one tuning sub-network and at least one conditional sub-network of the skip connection between the encoding layer and the decoding layer of the generative network in initial training to obtain a generative network to be fine-tuned. Wherein, the output of the encoding layer is connected to the input of at least one tuning sub-network, and the output of at least one tuning sub-network is connected to the input of the decoding layer. At least one tuning sub-network respectively corresponds to at least one conditional sub-network. The first output of each conditional sub-network is connected between the input of the corresponding tuning sub-network and the input of the encoding layer, and the second output of each conditional sub-network is connected between the output of the corresponding tuning sub-network and the input of the decoding layer.

[0097] S490: Use the input-side sample as the input of the encoder, the output-side sample as the output of the decoder, and the conditional sample as the input of at least one conditional sub-network to adjust the network parameters in at least one tuning sub-network and at least one conditional sub-network to obtain a generative model.

[0098] In this embodiment, the network structure of the generative network to be fine-tuned can refer to Figure 2Cand Figure 2E For example, step S470 is similar to step S410.

[0099] That is to say, in this embodiment, when the tuning subnetwork is attached with a conditional subnetwork (i.e., CSC-Tuner), as an example of fine-tuning the generative network to be fine-tuned, at least one tuning subnetwork corresponds to at least one conditional subnetwork respectively, and the input side sample can be used as the input of the encoder, the output side sample can be used as the output of the decoder, and the conditional sample can be used as the input of at least one conditional subnetwork, and the network parameters in at least one tuning subnetwork and at least one conditional subnetwork are adjusted to obtain a generative model. That is to say, in the conditional control generative image processing task, the condition is first encoded through the cascaded network layer (i.e., Cj), as an input of the tuning subnetwork, and is independently applied to the tuning subnetwork with another input of the tuning subnetwork from the encoding layer, thereby improving the training effect of the tuning subnetwork.

[0100] The following will be combined Figure 5 The application development methods according to other embodiments of the present invention are described in detail. Figure 5 The application development method can be applied to PAAS cloud service or IAAS cloud service, that is, the generative model executes the reasoning process on the server side of the PAAS cloud service or IAAS cloud service (its training process can be executed on the server side or executed on different server sides and then deployed to the server side), and the calling interface of the generative model is opened to the tenants of the PAAS cloud service or IAAS cloud service in the form of a virtual machine, that is, the calling interface of the generative model is provided to one virtual machine or multiple virtual machines, and the tenant can develop the application in the virtual based on the calling interface.

[0101] Specifically, application development methods include:

[0102] S510: Create a user interface module of the application, where the user interface module is configured to generate description data based at least on user operation data, and return presentation data based on the generated data of the description data.

[0103] S520: Acquire a calling interface of a generative model, where the generative model is obtained according to a model training method, and the calling interface is configured to return generated data when called;

[0104] S530: Embed at least the calling interface of the generative model into the user interface module.

[0105] In other examples, the user interface module is further configured to call the business module to return presentation data, the business module is configured to process the generated data to obtain presentation data, and the calling interface is configured to return the generated data when called by the business module.

[0106] Specifically, the user interface module is used to implement the interface for the application to interact with the user. It can be used to receive the input operations of the user and provide output results to the user. The application further includes at least one business module. The business module corresponds to the user interface module and is capable of performing data processing based on the user's input and obtaining a processing result. The processing process of the business module can be entirely executed by the generative model, that is, the input of the business module is directly or through simple preprocessing as the input of the generative model, and the output of the generative module is directly or through simple subsequent processing as the output of the business module. Alternatively, the processing process of the business module can be partially executed by the generative model, that is, the data processing ability of the business module is enhanced based on the inference ability of the generative model. At this time, the functional function of the business module is configured to call the call interface of the generative model, or the output data returned by the generative model through the call interface is used as the input of the functional function of the business module.

[0107] For example, the application is a content editing application, which includes business modules such as subtitle generation, copywriting generation, and dubbing generation. Correspondingly, a content editing interface can be provided in the user interface module. Further, by embedding the call interface of the generative model, generated subtitle data, copywriting data, or dubbing data, etc. can be returned from the call interface. The functional functions in the business module can further perform timestamp-based alignment processing on the subtitle data, copywriting data, or dubbing data, etc., and use the overall alignment processing result as the output of the business module, and return it to the user interface module to provide the alignment processing result to the content editing interface in a visual manner, enabling the user to achieve efficient development of the application. In addition, packaging the virtual machine and the call interface of the generative model and providing them to the tenant not only improves the development efficiency of the application but also reduces the resource requirements for local development and improves the flexibility of development.

[0108] Next, it will be combined with Figure 6 Describe a data processing device according to other embodiments of the present invention. Figure 6 The data processing device corresponds to the data processing method. The data processing device includes:

[0109] An acquisition module 610, which acquires a pre-trained generative model. The generative model includes a residual structure, and the residual structure includes a skip connection between an encoding layer and a decoding layer. At least one tuning sub-network is formed between the encoding layer and the decoding layer;

[0110] A generation module 620, which inputs at least the description data into the generative model to obtain generated data corresponding to the description data.

[0111] In other examples, the residual structure also includes at least one conditional sub-network, which corresponds to at least one tuning sub-network; the generation module is specifically used to: input the description data into the encoding layer of the generative model, input the conditional data into at least one conditional sub-network of the generative model, and obtain the generated data corresponding to the description data from the decoding layer of the generative model.

[0112] In other examples, the at least one tuning subnetwork corresponds to at least one conditional subnetwork, respectively, the first output of each conditional subnetwork is connected between the input of the corresponding tuning subnetwork and the input of the encoding layer, and the second output of each conditional subnetwork is connected between the output of the corresponding tuning subnetwork and the input of the decoding layer.

[0113] In other examples, the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, processed by the conditional weights of the conditional sub-network in the multiple conditional sub-networks, and output to the input of the decoding layer.

[0114] In other examples, the description data is a description text, the generated data includes multiple video frames, and the data processing device further includes: a video processing module that performs time-series encoding processing on the multiple video frames to obtain a generated video of the description text.

[0115] The following will be combined Figure 7 The intelligent interaction device according to some other embodiments of the present invention is described in detail. The intelligent interaction device corresponds to the intelligent interaction method, and specifically, the intelligent interaction device includes:

[0116] An acquisition module 710 acquires first description data input in the interactive interface;

[0117] The generation module 720 inputs at least the first description data into a pre-trained generative model to obtain first generated data corresponding to the first description data, wherein the first generated data includes at least one of an image, a video frame, and a video, and the generative model includes a residual structure, and the residual structure includes a jump connection between the encoding layer and the decoding layer, and at least one tuning subnetwork is formed between the encoding layer and the decoding layer.

[0118] The display module 730 displays the first generated data in the question-and-answer interface.

[0119] In some other examples, the acquisition module is further used to: acquire second description data modified for the first generated data from the interactive interface. The generation module is further used to: input at least the second description data into a pre-trained generative model to generate second generated data corresponding to the second description data as a modification result of the first generated data.

[0120] In some other examples, the first generated data is a picture or a video frame. The obtaining module is further configured to: obtain a modification mark for the picture or video from the interaction interface, identify the modification mark, and generate the second description data, or generate the second description data by combining the modification mark and the modification description text of the picture or video.

[0121] The following will be combined with Figure 8 A model training device according to some other embodiments of the present invention will be described in detail. The solution of this embodiment can be applied to any suitable electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.), and PC machines, etc. For example, in the model training (training) stage, a computing device (such as a data center) configured with a CPU (an example of a processing unit) + GPU (an example of an acceleration unit) architecture can be used to train an encoder-decoder model based on training samples. A computing device such as a data center can be deployed in a cloud server such as a proprietary cloud, a private cloud, or a hybrid cloud. Correspondingly, in the inference stage, a computing device configured with a CPU (an example of a processing unit) + GPU (an example of an acceleration unit) architecture can also be used for inference operations.

[0122] Specifically, Figure 8 The model training device corresponding to the model training method corresponds to the model training device including:

[0123] An obtaining module 810, which obtains a generative network for initial training. The generative network includes an encoder and a decoder, and a skip connection is formed between the encoding layer of the encoder and the decoding layer of the decoder.

[0124] A model adjustment module 820, which connects at least one tuning sub-network of the skip connection between the encoding layer and the decoding layer of the generative network for initial training to obtain a generative network to be fine-tuned, where the output of the encoding layer is connected to the input of the at least one tuning sub-network, and the output of the at least one tuning sub-network is connected to the input of the decoding layer.

[0125] A model training module 830, which performs fine-tuning training on the generative network to be fine-tuned based on training samples to obtain a generative model.

[0126] In the solution of the embodiment of the present invention, at least one tuning sub-network of the generative network to be fine-tuned is connected between the encoding layer and the decoding layer connected by the skip structure. The output of the encoding layer is connected to the input of at least one tuning sub-network, and the output of at least one tuning sub-network is connected to the input of the decoding layer. When fine-tuning and training the generative network, the skip structure is further optimized through at least one tuning sub-network, improving the efficiency of fine-tuning and training. In addition, at least one tuning sub-network is not in the encoder or the decoder, and there is no need to adjust parameters too much during fine-tuning and training, saving the computational overhead of fine-tuning and training.

[0127] In some other examples, fine-tuning and training the generative network to be fine-tuned based on training samples includes: performing multiple iterative trainings on the generative network to be fine-tuned based on training samples to obtain a generative model. Wherein, in each iterative training, the difference between the output of the input-side samples in the training samples passing through the forward propagation of the generative network to be fine-tuned and the input-side samples in the training samples is determined, and while keeping the parameters of the encoder unchanged, the parameters in the decoder and at least one tuning sub-network are adjusted through the gradient backpropagation of the difference in the decoder and at least one tuning sub-network.

[0128] In some other examples, fine-tuning and training the generative network to be fine-tuned based on training samples to obtain a generative model includes: using the input-side samples as the input of the encoder, using the output-side samples as the output of the decoder, and adjusting the network parameters in at least one tuning sub-network to obtain a generative model.

[0129] In some other examples, at least one tuning sub-network corresponds to at least one conditional sub-network respectively. The first output of each conditional sub-network is connected between the input of the corresponding tuning sub-network and the input of the encoding layer, and the second output of each conditional sub-network is connected between the output of the corresponding tuning sub-network and the input of the decoding layer; fine-tuning and training the generative network to be fine-tuned based on training samples to obtain a generative model includes: using the input-side samples as the input of the encoder, using the output-side samples as the output of the decoder, using the conditional samples as the input of at least one conditional sub-network, and adjusting the network parameters in at least one tuning sub-network and at least one conditional sub-network to obtain a generative model.

[0130] In some other examples, after the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, it is processed through the conditional weight of the conditional sub-network in the multiple conditional sub-networks and output to the input of the decoding layer.

[0131] In some other examples, N skip connections are respectively formed between multiple encoding layers of the encoder and multiple decoding layers of the decoder. Each skip connection corresponds to M tuning sub-networks and M conditional sub-networks, and the m-th conditional sub-network corresponds to the m-th tuning sub-network. The m-th conditional sub-network includes N network layers sequentially arranged from the input side to the output side. The n-th network layer corresponds to the n-th skip connection, and the output of the n-th network layer is connected to the input of the m-th tuning sub-network of the n-th skip connection.

[0132] The following will be combined with Figure 9 Describe an application development device according to some other embodiments of the present invention. The application development device corresponds to the application development method. Specifically, the application development device includes:

[0133] A creation module 910 that creates a user interface module of an application program. The user interface module is configured to generate description data at least based on user operation data and return presentation data based on the generated data of the description data;

[0134] An acquisition module 920 that acquires a call interface of a generative model. The generative model is obtained according to a model training method, and the call interface is configured to return the generated data when being called;

[0135] An embedding module 930 that embeds at least the call interface of the generative model into the user interface module.

[0136] In some other examples, the user interface module is further configured to call a service module to return the presentation data. The service module is configured to perform data processing on the generated data to obtain the presentation data, and the call interface is configured to return the generated data when being called by the service module.

[0137] For the specific implementation of each module in the above model training device or data processing device, reference can be made to the corresponding descriptions of the corresponding steps in the above method embodiments, and they have corresponding beneficial effects, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.

[0138] Referring to Figure 10 , a schematic structural diagram of an electronic device according to another embodiment of the present invention is shown. The specific implementation of the electronic device in the specific embodiments of the present invention is not limited.

[0139] As Figure 10As shown in the figure, the electronic device may include: a processor 1002 for executing program 1010, a communications interface 1004, a memory 1006, and a communication bus 1008.

[0140] The processor, the communications interface, and the memory communicate with each other through the communication bus.

[0141] The communications interface is used to communicate with other electronic devices or servers.

[0142] The processor is used to execute the program, and specifically may execute the relevant steps in the foregoing method embodiments.

[0143] Specifically, the program may include program code, and the program code includes computer operation instructions.

[0144] The processor may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0145] The memory is used to store the program. The memory may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0146] The program may include multiple computer instructions. Specifically, the program may cause the processor to execute the operations corresponding to the model training method or the data processing method described in any one of the foregoing multiple method embodiments through the multiple computer instructions.

[0147] For the specific implementation of each step in the program, reference may be made to the corresponding steps and the corresponding descriptions in the units in the foregoing method embodiments, and they have the corresponding beneficial effects, which will not be elaborated herein. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above may refer to the corresponding process descriptions in the foregoing method embodiments, and will not be elaborated herein.

[0148] An embodiment of the present invention also provides a computer storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method described in any one of the foregoing multiple method embodiments. The computer storage medium includes, but is not limited to: Compact Disc Read-Only Memory (CD-ROM), Random Access Memory (RAM), floppy disk, hard disk, magneto-optical disk, etc.

[0149] An embodiment of the present invention also provides a computer program product, including computer instructions, which direct a computing device to perform operations corresponding to the model training method or data processing method in the above-mentioned multiple method embodiments.

[0150] In addition, it should be noted that the information related to users (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data for training the model, data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the users or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with relevant regulations and standards, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0151] It should be pointed out that according to the needs of implementation, each component / step described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of the components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0152] The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and will be stored in a local recording medium. Thus, the method described herein can be stored on such a recording medium for software processing using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a Random Access Memory (RAM), a Read-Only Memory (ROM), a flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0153] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.

[0154] The above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention. The patent protection scope of the embodiments of the present invention shall be defined by the claims.

Claims

1. A data processing method, comprising: Obtaining a pre-trained generative model, the generative model including a residual structure, the residual structure including a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network being formed between the encoding layer and the decoding layer; Inputting at least description data into the generative model to obtain generated data corresponding to the description data.

2. The data processing method according to claim 1, wherein, The residual structure further includes at least one conditional sub-network, and the at least one conditional sub-network corresponds to at least one tuning sub-network; Inputting at least description data into the generative model to obtain generated data corresponding to the description data, including: Inputting the description data into the encoding layer of the generative model, inputting conditional data into at least one conditional sub-network of the generative model, and obtaining the generated data corresponding to the description data from the decoding layer of the generative model.

3. The data processing method according to claim 2, wherein, The at least one tuning sub-network respectively corresponds to at least one conditional sub-network, a first output of each conditional sub-network is connected between an input of the corresponding tuning sub-network and an input of the encoding layer, and a second output of each conditional sub-network is connected between an output of the corresponding tuning sub-network and an input of the decoding layer.

4. The data processing method according to claim 3, wherein, After the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, it is processed by the conditional weight of the conditional sub-network in the plurality of conditional sub-networks and output to the input of the decoding layer.

5. The data processing method according to claim 1, wherein, The description data is description text, and the generated data includes a plurality of video frames, The method further includes: Performing temporal encoding processing on the plurality of video frames to obtain a generated video of the description text.

6. An intelligent interaction method, comprising: Obtaining first description data input in an interaction interface; Inputting at least the first description data into a pre-trained generative model to obtain first generated data corresponding to the first description data, wherein the first generated data includes at least one of an image, a video frame, and a video, the generative model including a residual structure, the residual structure including a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network being formed between the encoding layer and the decoding layer; Displaying the first generated data in the question-and-answer interface.

7. The data processing method according to claim 6, wherein, The method further includes: Obtaining second description data for modifying the first generated data from the interaction interface; Inputting at least the second description data into a pre-trained generative model to generate second generated data corresponding to the second description data as a modification result of the first generated data.

8. The data processing method according to claim 6, wherein, The first generated data is a picture or a video frame, Obtaining second description data for modifying the first generated data from the interaction interface, including: Obtaining a modification mark for the picture or video from the interaction interface; Identifying the modification mark to generate the second description data, or combining the modification mark and the modification description text of the picture or video to generate the second description data.

9. A model training method, comprising: Obtain a generative network that has been initially trained, where the generative network includes an encoder and a decoder, and a skip connection is formed between the encoding layer of the encoder and the decoding layer of the decoder; Connect at least one tuning sub-network of the skip connection between the encoding layer and the decoding layer of the initially trained generative network to obtain a generative network to be fine-tuned, where the output of the encoding layer is connected to the input of the at least one tuning sub-network, and the output of the at least one tuning sub-network is connected to the input of the decoding layer; Based on training samples, perform fine-tuning training on the generative network to be fine-tuned to obtain a generative model.

10. The training method according to claim 9, wherein, Based on training samples, performing fine-tuning training on the generative network to be fine-tuned includes: Based on training samples, perform multiple iterative trainings on the generative network to be fine-tuned to obtain a generative model. Wherein, in each iterative training, determine the difference between the output of the forward propagation of the input-side samples in the training samples through the generative network to be fine-tuned and the input-side samples in the training samples, and while keeping the parameters of the encoder unchanged, adjust the parameters of the decoder and the at least one tuning sub-network through the gradient backpropagation of the difference in the decoder and the at least one tuning sub-network.

11. The training method according to claim 9, wherein, Based on training samples, performing fine-tuning training on the generative network to be fine-tuned to obtain a generative model includes: Use the input-side samples as the input of the encoder, use the output-side samples as the output of the decoder, and adjust the network parameters in the at least one tuning sub-network to obtain a generative model.

12. The training method according to claim 9, wherein, The at least one tuning sub-network respectively corresponds to at least one conditional sub-network. The first output of each conditional sub-network is connected between the input of the corresponding tuning sub-network and the input of the encoding layer, and the second output of each conditional sub-network is connected between the output of the corresponding tuning sub-network and the input of the decoding layer; Based on training samples, performing fine-tuning training on the generative network to be fine-tuned to obtain a generative model includes: Use the input-side samples as the input of the encoder, use the output-side samples as the output of the decoder, use the conditional samples as the input of the at least one conditional sub-network, and adjust the network parameters in the at least one tuning sub-network and the at least one conditional sub-network to obtain a generative model.

13. The training method according to claim 12, wherein, After the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, it is processed through the conditional weight of the conditional sub-network in the multiple conditional sub-networks and output to the input of the decoding layer.

14. An application development method, including: Create a user interface module for an application program, where the user interface module is configured to generate description data at least based on user operation data, and return presentation data based on the generated data of the description data; Obtain a call interface of a generative model, where the generative model is obtained according to the model training method described in any one of claims 9-13, and the call interface is configured to return the generated data when called; At least embed the call interface of the generative model into the user interface module.

15. The method according to claim 14, wherein The user interface module is further configured to call a service module to return the rendering data, the service module is configured to process the generated data to obtain the rendering data, and the call interface is configured to return the generated data when called by the service module.

16. An electronic device, comprising: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method according to any one of claims 1-15.

17. A computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method according to any one of claims 1-15.