Data processing method, intelligent interaction method, model training and development method, and device and medium

By introducing residual structure and tuning subnet in the generative model, the problem of insufficient training efficiency and data processing reliability of generative model is solved, and more efficient data processing and more reliable model prediction are achieved.

WO2025123765A1PCT designated stage expired Publication Date: 2025-06-19ALIBABA (CHINA) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/113737
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-08-21
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The existing generative models are costly and have poor reliability when training and data processing. Especially when the model scale is growing, there is room for improvement in training efficiency and data processing reliability.

Method used

A pre-trained generative model is adopted, which includes residual structure and tuning subnetwork. By combining jump connection and tuning subnetwork, the residual structure is efficiently trained to improve the prediction reliability of the generative model.

Benefits of technology

Through this method, the training efficiency of the generative model and the reliability of data processing are improved, the computational overhead of fine-tuning training is reduced, and more efficient data processing and more reliable model prediction are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024113737_19062025_PF_FP_ABST
    Figure CN2024113737_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing method, an intelligent interaction method, a model training and development method, and a device and a medium. The data processing method comprises: acquiring a pre-trained generative model, wherein the generative model comprises a residual structure, the residual structure comprises a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer; and at least inputting description data into the generative model, so as to obtain generated data corresponding to the description data. Since a residual structure comprises a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer, the at least one tuning sub-network can be connected between the encoding layer and the decoding layer, which form the skip connection, so as to efficiently train the residual structure to obtain a generative model having higher prediction reliability, and thus the generative model is used to execute more reliable data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing, intelligent interaction, model training and development methods, equipment and media

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 15, 2023, with application number 202311736419.9 and application name “Data processing, intelligent interaction, model training and development methods, equipment and media”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a method, device, and medium for data processing, intelligent interaction, model training, and development. Background Art

[0003] In the field of artificial intelligence, generative models such as large language models are constructed using generative networks. Diffusion generative networks, for example, have effective network structures and demonstrate excellent generalization capabilities after sample training. Basic generative models pre-trained on large-scale data can be fine-tuned for a range of downstream tasks and applications, such as image generation, image editing, and image transformation.

[0004] As the number of parameters in generative models continues to grow, adapting them to various tasks becomes prohibitively expensive. Low-cost training methods, however, result in poor training results, and consequently, data processing using trained generative models is less reliable. Therefore, there is room for improvement in both the training efficiency of generative models and the reliability of data processing using them.

[0005] Summary of the Invention

[0006] In view of this, embodiments of the present application provide a data processing, intelligent interaction, model training and development method, device and medium to at least partially solve the above problems.

[0007] According to a first aspect of an embodiment of the present application, a data processing method is provided, comprising: obtaining a pre-trained generative model, the generative model comprising a residual structure, the residual structure comprising a jump connection between an encoding layer and a decoding layer, and forming at least one tuning subnetwork between the encoding layer and the decoding layer; and inputting at least description data into the generative model to obtain generated data corresponding to the description data.

[0008] According to a second aspect of an embodiment of the present application, an intelligent interaction method is provided, comprising: obtaining first descriptive data input in an interactive interface; inputting at least the first descriptive data into a pre-trained generative model to obtain first generated data corresponding to the first descriptive data, wherein the first generated data includes at least one of an image, a video frame, and a video; the generative model includes a residual structure, the residual structure includes a jump connection between an encoding layer and a decoding layer, and at least one tuning subnetwork is formed between the encoding layer and the decoding layer. The first generated data is displayed in the question-and-answer interface.

[0009] According to a third aspect of an embodiment of the present application, a model training method is provided, comprising: obtaining an initially trained generative network, the generative network comprising an encoder and a decoder, wherein a jump connection is formed between the encoding layer of the encoder and the decoding layer of the decoder; connecting at least one tuning subnetwork of the jump connection between the encoding layer and the decoding layer of the initially trained generative network to obtain a generative network to be fine-tuned, wherein the output of the encoding layer is connected to the input of the at least one tuning subnetwork, and the output of the at least one tuning subnetwork is connected to the input of the decoding layer; and fine-tuning the generative network to be fine-tuned based on training samples to obtain a generative model.

[0010] According to the fourth aspect of the embodiments of the present application, an application development method is provided, comprising: creating a user interface module of an application, the user interface module being configured to generate description data based at least on user operation data, and returning presentation data based on the generated data of the description data; obtaining a calling interface of a generative model, the generative model being obtained according to the model training method described in the third aspect, the calling interface being configured to return the generated data when called; and embedding at least the calling interface of the generative model into the user interface module.

[0011] According to the fifth aspect of the embodiment of the present application, a data processing device is provided, including: an acquisition module, configured to acquire a pre-trained generative model, the generative model including a residual structure, the residual structure including a jump connection between the encoding layer and the decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer; a generation module, configured to input at least description data into the generative model to obtain generated data corresponding to the description data.

[0012] According to the sixth aspect of the embodiment of the present application, an intelligent interaction device is provided, including: an acquisition module, configured to acquire first descriptive data input in an interactive interface; a generation module, configured to at least input the first descriptive data into a pre-trained generative model to obtain first generated data corresponding to the first descriptive data, wherein the first generated data includes at least one of an image, a video frame, and a video, the generative model includes a residual structure, the residual structure includes a jump connection between the encoding layer and the decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer; a display module, configured to display the first generated data in the question-and-answer interface.

[0013] According to the seventh aspect of the embodiment of the present application, a model training device is provided, comprising: a module acquisition module, configured to acquire an initially trained generative network, the generative network comprising an encoder and a decoder, and a jump connection is formed between the encoding layer of the encoder and the decoding layer of the decoder; a model adjustment module, configured to connect at least one tuning subnetwork of the jump connection between the encoding layer and the decoding layer of the initially trained generative network, to obtain a generative network to be fine-tuned, wherein the output of the encoding layer is connected to the input of the at least one tuning subnetwork, and the output of the at least one tuning subnetwork is connected to the input of the decoding layer; a model training module, configured to perform fine-tuning training on the generative network to be fine-tuned based on training samples, to obtain a generative model.

[0014] According to the eighth aspect of the embodiment of the present application, an application development device is provided, including: a creation module, configured to create a user interface module of an application, the user interface module being configured to generate description data based on at least user operation data, and return presentation data based on the generated data of the description data; an acquisition module, configured to obtain a calling interface of a generative model, the generative model being obtained according to the model training method described in the third aspect, the calling interface being configured to return the generated data when called; and an embedding module, configured to embed at least the calling interface of the generative model into the user interface module.

[0015] According to the ninth aspect of the embodiments of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect.

[0016] According to a tenth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.

[0017] In this embodiment, the pre-trained generative model includes a residual structure, the residual structure includes a jump connection between the encoding layer and the decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer. Therefore, the at least one tuning sub-network can efficiently train the residual structure by connecting to the encoding layer and the decoding layer of the jump connection, thereby obtaining a generative model with higher prediction reliability, so that more reliable data processing can be performed using the generative model. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0019] FIG1 is a schematic diagram of a generative network structure according to some examples.

[0020] FIG2A is a flowchart of steps of a data processing method according to other embodiments of the present application.

[0021] FIG. 2B is a schematic diagram of a partial structure of a generative network of some examples of the embodiment of FIG. 2 .

[0022] FIG. 2C is a schematic diagram of a partial structure of a generative network of other examples of the embodiment of FIG. 2 .

[0023] FIG2D is a flowchart of steps of a data processing method according to other embodiments of the present application.

[0024] FIG2E is a schematic diagram of a partial structure of a generative network of other examples of the embodiment of FIG2 .

[0025] FIG3 is a flowchart of steps of an intelligent interaction method according to other embodiments of the present application.

[0026] FIG4A is a flowchart of steps of a model training method according to other embodiments of the present application.

[0027] FIG4B is a flowchart of the steps of a model training method according to other embodiments of the present application.

[0028] FIG5 is a flowchart of steps of an application development method according to other embodiments of the present application.

[0029] FIG6 is a schematic block diagram of a data processing device according to some other embodiments of the present application.

[0030] FIG7 is a schematic block diagram of an intelligent interaction device according to other embodiments of the present application.

[0031] FIG8 is a schematic block diagram of a model training device according to other embodiments of the present application.

[0032] FIG9 is a schematic block diagram of an application development device according to other embodiments of the present application.

[0033] FIG10 is a schematic structural diagram of an electronic device according to other embodiments of the present application. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.

[0035] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.

[0036] Generative models such as large language models are constructed through generative networks. For example, the generative network shown in Figure 1 includes an encoder 110 and a decoder 120. The output of the encoder 110 can be directly connected to the input of the decoder 120. The output of the encoder 110 can also be connected to the input of the decoder 120 through a context alignment layer (not shown) using an attention mechanism. The generative network of Figure 1 can be a generative diffusion network, which has a skip connection (SC) structure between the encoding layer and the decoding layer. The skip connection can be implemented as one or more, for example, skip connection #1, skip connection #2, ..., skip connection #N shown in Figure 1. Each skip connection forms a residual structure between the encoder 110 and the decoder 120, thereby enabling the generative network to obtain better generalization capabilities. However, as the number of parameters in the generative model scale continues to grow, the cost of adapting such a model to various tasks becomes quite expensive. The various embodiments of the present application provide a series of technical solutions that can improve the efficiency of fine-tuning training and reduce the fine-tuning training overhead of the generative network.

[0037] Figure 2A is a flowchart of the steps of the data processing method according to some other embodiments of the present application. The solution of this embodiment can be applied to any appropriate electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.) and PCs, etc. For example, in the model training (training) stage, a computing device (for example, a data center) configured with a CPU (an example of a processing unit) + GPU (an example of an acceleration unit) architecture can be used to train the encoder-decoder model based on training samples. Computing devices such as data centers can be deployed in cloud servers such as proprietary clouds, private clouds, or hybrid clouds. Accordingly, in the inference (inference) stage, a computing device configured with a CPU (an example of a processing unit) + GPU (an example of an acceleration unit) architecture can also be used to perform inference operations. Specifically, the data processing method of Figure 2A includes:

[0038] S210: Obtain a pre-trained generative model, where the generative model includes a residual structure, the residual structure includes a skip connection between an encoding layer and a decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer.

[0039] It should be understood that the generative model can be trained using a generative network, which can be a generative diffusion network. The generative model can be a generative model that generates images based on descriptive text, or a text processing model that generates text based on text, such as a knowledge question-answering model or a translation model.

[0040] It should also be understood that the generative model may include an encoder and a decoder, and the encoder and decoder may have a residual structure with one or more skip connections (SC), such as a U-Net structure. The network layer in the encoder is the encoding layer, and the network layer in the decoder is the decoding layer. The encoding layer and the decoding layer connected by the skip connection correspond to each other. The encoder and decoder may use an attention mechanism or the like to form a context-aligned structure including one or more network layers (e.g., a fully connected layer).

[0041] S220: Input at least the description data into the generative model to obtain generated data corresponding to the description data.

[0042] It should be understood that the description data can describe text or can be converted into description data that describes text. For example, the speech description data can be converted into corresponding description text and input into the generative model. As for the specific conversion method, the speech description data can be input into a pre-trained text conversion model to obtain the description text, or the text conversion model can be connected to the input side of the encoder of the generative model. Alternatively, the generative model can also be trained to receive speech description data as input. That is, when the generative model is trained end-to-end, the input side sample of the training sample is a speech sample.

[0043] It should also be understood that the generated data can be descriptive text (for example, in knowledge question-answering or translation scenarios); it can also be images or video frames, or videos based on video frame sequences. This depends on the sample type used in the end-to-end training phase of the generative model. If the output side samples are video frame sequences, the generated data in the inference phase is video; if the output side samples are images or video frames, the generated data in the inference phase is images or video frames.

[0044] In this embodiment, the pre-trained generative model includes a residual structure, the residual structure includes a jump connection between the encoding layer and the decoding layer, and at least one tuning sub-network is formed between the encoding layer and the decoding layer. Therefore, the at least one tuning sub-network can efficiently train the residual structure by connecting to the encoding layer and the decoding layer of the jump connection, thereby obtaining a generative model with higher prediction reliability, so that more reliable data processing can be performed using the generative model.

[0045] In other examples, various conditional image data are input into separate conditional sub-networks. The conditional image can be descriptive text and multiple additional conditions. Different image conditions are used to describe different types of image features. The various types of image feature data include, but are not limited to, image edge detection features, image depth features, image segmentation features, image color features, image expansion features, image restoration features, and other different image features. Specifically, the descriptive text is input into the encoder of the generative model, and each additional condition is input into the corresponding tuning sub-network of multiple jump connections to obtain the output generated image.

[0046] Specifically, in the residual structure (U-Net structure) shown in FIG2B , the output of the encoding layer is connected to the input of each tuning sub-network 30, and the output of each tuning sub-network is connected to the input of the decoding layer. In other words, the output of each tuning sub-network is fused with the input of the encoding layer via a skip connection as the input of the skip-connected decoding layer. That is:

[0047] Among them, Oj SC is the output of the jth layer in the residual structure, XN-j is the input of the skip connection, Tj is the output of the tuning sub-network, and N is the total number of encoder layers in the residual structure.

[0048] In other examples, the residual structure further includes at least one conditional sub-network, and the at least one conditional sub-network corresponds to the at least one tuning sub-network. As an example of inputting at least the description data into the generative model to obtain generated data corresponding to the description data, the description data can be input into the encoding layer of the generative model, the conditional data can be input into at least one conditional sub-network of the generative model, and the generated data corresponding to the description data can be obtained from the decoding layer of the generative model.

[0049] Specifically, at least one tuning subnetwork corresponds to at least one conditional subnetwork, the first output of each conditional subnetwork is connected between the input of the corresponding tuning subnetwork and the input of the encoding layer, and the second output of each conditional subnetwork is connected between the output of the corresponding tuning subnetwork and the input of the decoding layer.

[0050] More specifically, the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, processed by the conditional weights of the conditional sub-network in multiple conditional sub-networks, and output to the input of the decoding layer.

[0051] In other examples, the description data is a description text, and the generated data includes multiple video frames. The data processing method further includes: performing time-series coding processing on the multiple video frames to obtain a generated video of the description text.

[0052] FIG2D shows a data processing method according to other embodiments of the present application.

[0053] S280: Obtain a pre-trained generative model.

[0054] S290: Input the description data into the encoding layer of the generative model, input the conditional data into at least one conditional sub-network of the generative model, and obtain generated data corresponding to the description data from the decoding layer of the generative model.

[0055] It should be understood that the same or similar parts of the embodiment of FIG2D and the embodiment of FIG2A are not described in detail here. For example, step S280 is similar to step S210.

[0056] Furthermore, the generative model of this embodiment may include a tuning subnetwork with a conditional subnetwork. For example, as shown in FIG2C , at least one tuning subnetwork 30 corresponds to at least one conditional subnetwork 40, with the first output of each conditional subnetwork 40 connected between the input of the corresponding tuning subnetwork 30 and the input of the encoding layer, and the second output of each conditional subnetwork 40 connected between the output of the corresponding tuning subnetwork 30 and the input of the decoding layer.

[0057] Right now: Among them, Oj CSC is the output of the j-th layer of the residual structure after the tuning sub-network is processed; m is the m-th condition, traversing values ​​between 1 and M; X N-j is the input; Tj m is the tuning subnetwork of the jth layer and the mth condition, α m is the conditional weight coefficient of the mth condition; Cj is the input after the mth conditional encoding of the jth layer, and Cj is the condition set of {Cj0…Cjm}.

[0058] The M conditions are conditional data that supplement the encoder input data during fine-tuning training. For example, when the encoder input is a description text, the M conditions can be conditional images that supplement the description text.

[0059] Furthermore, as shown in FIG2E , for each conditional subnetwork (eg, M conditional subnetworks), the number of conditional subnetworks may be consistent with the number of tuning subnetworks, that is, the mth conditional subnetwork corresponds to the mth tuning subnetwork.

[0060] More specifically, N skip connections are formed between the multiple encoding layers of the encoder and the multiple decoding layers of the decoder. Each skip connection corresponds to M tuning subnetworks and M conditional subnetworks, and the mth conditional subnetwork corresponds to the mth tuning subnetwork. The mth conditional subnetwork includes N network layers arranged sequentially from the input side to the output side. The nth network layer corresponds to the nth skip connection, and the output of the nth network layer is connected to the input of the mth tuning subnetwork of the nth skip connection. Where N is an integer greater than 1, M is an integer greater than 1, 1≤n≤N, and 1≤m≤M.

[0061] Furthermore, the output of the nth network layer is connected to the encoding layer of the nth skip connection. It should be understood that each network layer can be implemented by a linear layer such as a downsampling layer and an activation layer, and the data dimension of the output of the nth network layer is aligned or consistent with the data dimension of the encoding layer of the nth skip connection, thereby improving the reliability of the generative model training process and the reliability of data processing using the generative model.

[0062] In some variations of the data processing method, a generated image of a generative model including a tuning subnetwork without a conditional subnetwork can be used as a conditional image of a generative model including a tuning subnetwork with a conditional subnetwork, thereby generating a conditional image through a descriptive text, and further using the generative model of the tuning subnetwork with a conditional subnetwork to obtain a generated image by combining both the conditional image and the further descriptive text.

[0063] In some other variations of the data processing method, the generated image of the generative model of the tuning subnetwork without the conditional subnetwork can be used as input data of the encoder of the generative model of the tuning subnetwork including the conditional subnetwork, so as to further obtain the generated image that better matches the description text by using the attached conditional subnetwork.

[0064] The following describes in detail the flow chart of the steps of the intelligent interaction method based on the generative model in some other embodiments of the present application in conjunction with Figure 3. The intelligent interaction method in Figure 3 includes:

[0065] S310: Acquire first description data input in the interactive interface.

[0066] S320: Input at least the first description data into a pre-trained generative model to obtain first generated data corresponding to the first description data, wherein the first generated data includes at least one of an image, a video frame, and a video, and the generative model includes a residual structure, the residual structure includes a jump connection between the encoding layer and the decoding layer, and at least one tuning subnetwork is formed between the encoding layer and the decoding layer.

[0067] S330: Display the first generated data in the question-and-answer interface.

[0068] In this embodiment, the user interface is used to receive the descriptive data input by the user and the data processing results are displayed to the user, thereby achieving quick interaction. In addition, this embodiment does not limit the type of descriptive data input by the user, making it convenient for the user to input various types of descriptive data such as text descriptive data or voice descriptive data. In addition, the generated data presented by the user interface includes but is not limited to text, pictures or videos. Therefore, this embodiment provides users with intelligent information services by combining the convenient interaction method of the user interface and the data processing capabilities of the generative model. For example, it can be applied to intelligent products such as intelligent assistants and virtual experts.

[0069] In some examples, the intelligent interaction method also includes: obtaining second descriptive data modified from the interactive interface for the first generated data; at least inputting the second descriptive data into a pre-trained generative model to generate second generated data corresponding to the second descriptive data as a modification result of the first generated data.

[0070] Through the above data modification process, after the initial generation process of the first generated data, the first description data is used as context data and the second description data is used as modified description data to adjust the first generated data to obtain the second generated description.

[0071] It should be understood that the types of the first generated data and the second generated data here can be consistent. For example, the voice data is modified to obtain modified voice data, the text data is modified to obtain modified text data, and the video data is modified to obtain modified video data.

[0072] Furthermore, the first description data and the second description data may be of the same or different types. For example, in some other examples, the first generated data is a picture or video frame. Obtaining second description data modified from the interactive interface for the first generated data includes: obtaining a modification mark for the picture or video from the interactive interface; identifying the modification mark and generating the second description data; or generating the second description data by combining the modification mark and modified description text of the picture or video.

[0073] That is to say, the second description data of the same type as the first description data can be used as a supplement and modification prompt for the first description data, or the second description data of a different type from the first description data can be used as a supplement and modification prompt for the first description data. For example, when the first description data is text data, the text data can be further supplemented to modify the first generated data, or voice data / operation data different from the text data can be used as a supplement or modification prompt for the text data. The operation data can be the target area marked in the above-mentioned picture or video frame, that is, the target area is the area to be modified or the area with a larger modification weight. Through such an operation, the second description data combined with the operation data can further improve the accuracy and reliability of the modification because it provides more information on the modification weight.

[0074] Specifically, FIG4A shows a model training method according to some embodiments of the present application. The model training method of FIG4A includes:

[0075] S410: Obtain an initially trained generative network, where the generative network includes an encoder and a decoder, and a skip connection is formed between an encoding layer of the encoder and a decoding layer of the decoder.

[0076] It should be understood that the generative network initially trained may be a generative network obtained through pre-training, for example, a generative diffusion network.

[0077] S420: Connect at least one tuning subnetwork with a jump connection between the encoding layer and the decoding layer of the initially trained generative network to obtain a generative network to be fine-tuned, wherein the output of the encoding layer is connected to the input of the at least one tuning subnetwork, and the output of the at least one tuning subnetwork is connected to the input of the decoding layer.

[0078] It should be understood that at least one tuning subnetwork can be connected between some or all of the skip-connected encoding layers and decoding layers, and the structure formed by the tuning subnetwork can be called an SC-Tuner. The tuning subnetwork includes at least a linear layer and an activation layer. For example, the tuning subnetwork can be formed by a first linear layer, an activation layer, and a second linear layer, and the activation layer is arranged between the first linear layer and the second linear layer. In addition, other activation layers can be added to the input side of the first linear layer and / or the output side of the second linear layer, and other linear layers can be added to any position of the tuning subnetwork. The above-mentioned linear layer can be a fully connected layer or a convolutional layer. In the case where the output of the generative network to be fine-tuned is image data, the linear layer in the tuning subnetwork can be two-dimensional data. In the case where the output of the generative network to be fine-tuned is text data, the linear layer in the tuning subnetwork can be one-dimensional data. This is because the tuning subnetwork is connected between the skip-connected encoding layer and the decoding layer and only needs to have the same dimension as the output data of the decoder.

[0079] S430: Based on the training samples, fine-tune the generative network to be fine-tuned to obtain a generative model.

[0080] It should be understood that the fine-tuning training of the fine-tuning generative network can be supervised training, and the training samples can include input side samples and output side samples. The input side samples can be text data injected into a specific coding layer of the encoder, and the input side samples can also include image samples input to the encoder. The output side samples can be text data or image data such as pictures or video frames.

[0081] In the embodiment of the present application, at least one tuning subnetwork of the generative network to be fine-tuned is connected between the encoding layer and the decoding layer connected to the skip structure. The output of the encoding layer is connected to the input of at least one tuning subnetwork, and the output of at least one tuning subnetwork is connected to the input of the decoding layer. This allows the skip structure to be further optimized by the at least one tuning subnetwork during fine-tuning training of the generative network, thereby improving the efficiency of fine-tuning training. In addition, since at least one tuning subnetwork is not in the encoder or decoder, there is no need for excessive parameter adjustments during fine-tuning training, which saves the computational overhead of fine-tuning training.

[0082] Furthermore, as an example of fine-tuning training of a generative network to be fine-tuned based on training samples, the generative network to be fine-tuned can be subjected to multiple iterative training based on the training samples to obtain a generative model, wherein, in each iterative training, the difference between the input side sample in the training sample after the forward propagation output of the generative network to be fine-tuned and the input side sample in the training sample is determined, and while maintaining the various parameters of the encoder, the various parameters in the decoder and the at least one tuning sub-network are adjusted by back-propagating the gradient of the difference in the decoder and the at least one tuning sub-network. That is, in each iterative training, there is no need to train the parameters in the encoder, and the training process of the encoder and the at least one tuning sub-network is decoupled, thereby significantly reducing the computational overhead in fine-tuning training.

[0083] Specifically, the input side sample can be a descriptive text, and the output side sample can be a generated image. For example, the descriptive text is used to describe the generated image, so that the trained generative model can more effectively adapt to the downstream tasks of the initially trained generative network.

[0084] For example, taking the stable diffusion generative network as an example, residual structures such as U-Net often contain 12 layers of jump connections. By adding a tuning subnetwork to the jump connection of each layer, efficient training can be performed for different tasks. Compared with other fine-tuning training methods such as LoRA, training can be performed with smaller tuning parameters and memory consumption. In other words, the application of each tuning subnetwork in efficient generative tuning or few-sample task tuning can be applied to fast and customized transfer training under the conditions of a small number of samples, such as learning specific character features, specific image styles, and other image features.

[0085] In the case of a tuning subnetwork without a conditional subnetwork (i.e., SC-Tuner), as an example of fine-tuning a generative network to be fine-tuned, the input side samples can be used as the input of the encoder, and the output side samples can be used as the output of the decoder. The network parameters in at least one tuning subnetwork can be adjusted to obtain a generative model, thereby reliably training the tuning subnetwork without the conditional subnetwork. In other words, the input side constraints and output side constraints of the tuning subnetwork are derived from the encoding layer and decoding layer of the skip connection corresponding to the tuning subnetwork. Accordingly, the input side samples include at least descriptive text, and the output side samples include at least image samples, which include but are not limited to pictures and video frames.

[0086] Furthermore, after the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, it is processed by the conditional weights of the conditional sub-network in multiple conditional sub-networks and output to the input of the decoding layer. It should be understood that since the importance of the conditions corresponding to the multiple conditional sub-networks can be different, the conditional weights of a conditional sub-network in multiple conditional sub-networks indicate the importance of the conditional sub-network. In addition, the conditional weights of the conditional sub-network can also be implemented as a linear layer such as a fully connected layer, connected between each tuning sub-network and the corresponding encoding layer. At least one tuning sub-network and its conditional sub-network correspond to a skip connection, and the tuning sub-network with the conditional sub-network can also be referred to as a CSC-Tuner.

[0087] Furthermore, different conditional sub-networks have different functions, that is, multiple conditional sub-networks are used to output multiple image feature data respectively. For example, the multiple image feature data include but are not limited to different image features such as image edge detection features, image depth features, image segmentation features, image color features, image expansion features, and image restoration features.

[0088] Without loss of generality, in the model training framework, components for fine-tuning the training residual structure can be configured by inputting skip connections and related conditions into SC-Tuner (i.e., the tuning subnetwork of the subnetwork without conditions) and CSC-Tuner (i.e., the tuning subnetwork of the subnetwork with conditions), and concatenating them with the input of the original decoding layer of the corresponding residual structure and sending them to the next decoding layer.

[0089] That is, g j+1 =G j [O j (X N-j-1 +C j );g j ], where gj is the input of the decoding layer of the residual structure of the j-th layer; Gj is the output of the decoding layer of the residual structure of the j-th layer after operation; Oj is the output of the tuning sub-network of the j-th layer in the residual structure after processing.

[0090] Multiple skip connections 130 are respectively formed between the multiple encoding layers of the encoder 120 and the multiple decoding layers of the decoder 120. The conditional sub-network 40 includes multiple network layers arranged in sequence from the input side to the output side, and the multiple network layers correspond to the multiple skip connections respectively.

[0091] Specifically, as shown in FIG2E , for each conditional subnetwork (eg, M conditional subnetworks), the number of conditional subnetworks may be consistent with the number of tuning subnetworks, that is, the mth conditional subnetwork corresponds to the mth tuning subnetwork.

[0092] More specifically, N skip connections are formed between the multiple encoding layers of the encoder and the multiple decoding layers of the decoder. Each skip connection corresponds to M tuning subnetworks and M conditional subnetworks, and the mth conditional subnetwork corresponds to the mth tuning subnetwork. The mth conditional subnetwork includes N network layers arranged sequentially from the input side to the output side. The nth network layer corresponds to the nth skip connection, and the output of the nth network layer is connected to the input of the mth tuning subnetwork of the nth skip connection. Where N is an integer greater than 1, M is an integer greater than 1, 1≤n≤N, and 1≤m≤M.

[0093] Furthermore, the output of the nth network layer is connected to the encoding layer of the nth skip connection. It should be understood that each network layer can be implemented by a linear layer such as a downsampling layer and an activation layer, and the data dimension of the output of the nth network layer is aligned or consistent with the data dimension of the encoding layer of the nth skip connection, thereby improving the reliability of the training process of the generative model.

[0094] That is, the output of the mth conditional subnetwork of the n+1th jump connection is connected to the input of the n+1th network layer, resulting in the mth conditional subnetwork of the nth jump connection. That is, the output of the n+1th network layer is used as the output of the mth conditional subnetwork of the nth jump connection.

[0095] FIG4B shows a model training method according to some other embodiments of the present application, including:

[0096] S470: Obtain an initially trained generative network, where the generative network includes an encoder and a decoder, and a skip connection is formed between an encoding layer of the encoder and a decoding layer of the decoder.

[0097] S480: Connect at least one tuning subnetwork and at least one conditional subnetwork with skip connections between the encoding layer and the decoding layer of the initially trained generative network to obtain a generative network to be fine-tuned, wherein the output of the encoding layer is connected to the input of the at least one tuning subnetwork, and the output of the at least one tuning subnetwork is connected to the input of the decoding layer. At least one tuning subnetwork corresponds to at least one conditional subnetwork, and the first output of each conditional subnetwork is connected between the input of the corresponding tuning subnetwork and the input of the encoding layer, and the second output of each conditional subnetwork is connected between the output of the corresponding tuning subnetwork and the input of the decoding layer.

[0098] S490: Using the input side sample as the input of the encoder, using the output side sample as the output of the decoder, using the conditional sample as the input of at least one conditional sub-network, adjusting the network parameters in at least one tuning sub-network and at least one conditional sub-network, and obtaining a generative model.

[0099] [Corrected 11.09.2024 according to Rule 91] In this embodiment, the network structure of the generative network to be fine-tuned can refer to Figures 2C and 2E, and will not be repeated here. For example, step S470 is similar to step S410.

[0100] That is, in this embodiment, when a tuning subnetwork is accompanied by a conditional subnetwork (i.e., CSC-Tuner), as an example of fine-tuning training for a generative network to be fine-tuned, at least one tuning subnetwork corresponds to at least one conditional subnetwork. Input samples can be used as input to the encoder, output samples as output to the decoder, and conditional samples as input to at least one conditional subnetwork. The network parameters in at least one tuning subnetwork and at least one conditional subnetwork are adjusted to obtain a generative model. In other words, in a conditionally controlled generative image processing task, the condition is first encoded (i.e., Cj) through a cascade of network layers and then applied to the tuning subnetwork as an input. This is then independently applied to the tuning subnetwork along with another input from the encoding layer, thereby improving the training effect of the tuning subnetwork.

[0101] The application development method according to other embodiments of the present application will be described in detail below in conjunction with Figure 5. The application development method of Figure 5 can be applied to PAAS cloud services or IAAS cloud services, that is, the generative model performs the reasoning process on the server side of the PAAS cloud service or IAAS cloud service (its training process can be executed on the server side or on different server sides before being deployed to the server side), and opens the calling interface of the generative model to the tenants of the PAAS cloud service or IAAS cloud service in the form of a virtual machine, that is, the calling interface of the generative model is provided to one or more virtual machines, and the tenant can perform application development in the virtual based on the calling interface.

[0102] Specifically, application development methods include:

[0103] S510: Create a user interface module of the application, where the user interface module is configured to generate description data based at least on user operation data, and return presentation data based on the generated data of the description data.

[0104] S520: Obtain a calling interface for a generative model, where the generative model is obtained according to a model training method, and the calling interface is configured to return generated data when called;

[0105] S530: Embed at least the calling interface of the generative model into the user interface module.

[0106] In other examples, the user interface module is further configured to call the service module to return presentation data, the service module is configured to process the generated data to obtain presentation data, and the calling interface is configured to return the generated data when called by the service module.

[0107] Specifically, the user interface module is used to implement the interface for the application to interact with the user, and can be used to receive user input operations and provide output results to the user. The application also includes at least one service module, which corresponds to the user interface module and can perform data processing based on the user input and obtain processing results. The processing process of the service module can be completely executed by the generative model, that is, the input of the service module is directly or through simple preprocessing as the input of the generative model, and the output of the generative module is directly or through simple subsequent processing as the output of the service module. Alternatively, the processing process of the service module can be partially executed by the generative model, that is, the data processing capability of the service module is enhanced based on the reasoning capability of the generative model. In this case, the functional function of the service module is configured to call the calling interface of the generative model, or the output data returned by the generative model through the calling interface is used as the input of the functional function of the service module.

[0108] For example, the application is a content editing application, which includes service modules such as subtitle generation, copywriting generation, and dubbing generation. Accordingly, a content editing interface can be provided in the user interface module. Furthermore, by embedding the calling interface of the production model, the generated subtitle data, copywriting data, or dubbing data, etc. can be returned from the calling interface, and the functional functions in the service module can further perform timestamp-based alignment processing on the subtitle data, copywriting data, or dubbing data, etc., and return the alignment processing results as a whole as the output of the service module to the user interface module to provide the alignment processing results in a visual manner to the content editing interface, so that users can achieve efficient development of applications. In addition, the virtual machine and the calling interface of the generative model are packaged and provided to tenants, which not only improves the development efficiency of the application but also reduces the resource requirements of local development and improves the flexibility of development.

[0109] The following describes a data processing device according to some other embodiments of the present application in conjunction with FIG6 . The data processing device in FIG6 corresponds to the data processing method, and the data processing device includes:

[0110] An acquisition module 610 is configured to acquire a pre-trained generative model, wherein the generative model includes a residual structure, the residual structure includes a skip connection between an encoding layer and a decoding layer, and at least one tuning subnetwork is formed between the encoding layer and the decoding layer;

[0111] The generation module 620 is configured to at least input the description data into the generative model to obtain generated data corresponding to the description data.

[0112] In other examples, the residual structure also includes at least one conditional sub-network, which corresponds to at least one tuning sub-network; the generation module is configured to: input the description data into the encoding layer of the generative model, input the conditional data into at least one conditional sub-network of the generative model, and obtain the generated data corresponding to the description data from the decoding layer of the generative model.

[0113] In other examples, the at least one tuning subnetwork corresponds to at least one conditional subnetwork, the first output of each conditional subnetwork is connected between the input of the corresponding tuning subnetwork and the input of the encoding layer, and the second output of each conditional subnetwork is connected between the output of the corresponding tuning subnetwork and the input of the decoding layer.

[0114] In other examples, the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, processed by the conditional weights of the conditional sub-network in the multiple conditional sub-networks, and output to the input of the decoding layer.

[0115] In other examples, the description data is description text, the generated data includes multiple video frames, and the data processing device also includes: a video processing module, configured to perform time-series encoding processing on the multiple video frames to obtain a generated video of the description text.

[0116] The following will describe in detail the intelligent interaction device according to other embodiments of the present application in conjunction with FIG7 . The intelligent interaction device corresponds to the intelligent interaction method, and specifically, the intelligent interaction device includes:

[0117] An acquisition module 710 is configured to acquire first description data input in the interactive interface;

[0118] The generation module 720 is configured to input at least the first description data into a pre-trained generative model to obtain first generated data corresponding to the first description data, wherein the first generated data includes at least one of an image, a video frame, and a video, and the generative model includes a residual structure, and the residual structure includes a jump connection between the encoding layer and the decoding layer, and at least one tuning subnetwork is formed between the encoding layer and the decoding layer.

[0119] The display module 730 is configured to display the first generated data in the question-and-answer interface.

[0120] In some other examples, the acquisition module is further configured to: acquire, from the interactive interface, second description data that is modified for the first generated data. The generation module is further configured to: input at least the second description data into a pre-trained generative model to generate second generated data corresponding to the second description data as a modification result of the first generated data.

[0121] In some other examples, the first generated data is a picture or video frame. The acquisition module is further configured to: obtain a modification mark for the picture or video from the interactive interface, identify the modification mark, and generate the second description data, or generate the second description data by combining the modification mark and a modification description text of the picture or video.

[0122] The model training device according to other embodiments of the present application will be described in detail below in conjunction with Figure 8. The solution of this embodiment can be applied to any appropriate electronic device with data processing capabilities, including but not limited to: servers, mobile terminals (such as mobile phones, PADs, etc.) and PCs, etc. For example, in the model training stage, a computing device (for example, a data center) configured with a CPU (an example of a processing unit) + GPU (an example of an acceleration unit) architecture can be used to train the encoder-decoder model based on training samples. Computing devices such as data centers can be deployed in cloud servers such as proprietary clouds, private clouds, or hybrid clouds. Accordingly, in the inference stage, a computing device configured with a CPU (an example of a processing unit) + GPU (an example of an acceleration unit) architecture can also be used to perform inference operations.

[0123] Specifically, the model training device of FIG8 corresponds to the model training method, and the model training device includes:

[0124] The module acquisition module 810 is configured to acquire an initially trained generative network, where the generative network includes an encoder and a decoder, and a skip connection is formed between the encoding layer of the encoder and the decoding layer of the decoder.

[0125] The model adjustment module 820 is configured to connect at least one tuning subnetwork of the jump connection between the encoding layer and the decoding layer of the initially trained generative network to obtain a generative network to be fine-tuned, wherein the output of the encoding layer is connected to the input of the at least one tuning subnetwork, and the output of the at least one tuning subnetwork is connected to the input of the decoding layer.

[0126] The model training module 830 is configured to perform fine-tuning training on the generative network to be fine-tuned based on the training samples to obtain a generative model.

[0127] In the solution of the embodiment of the present application, at least one tuning sub-network of the generative network to be fine-tuned is connected between the encoding layer and the decoding layer connected to the jump structure, the output of the encoding layer is connected to the input of at least one tuning sub-network, and the output of at least one tuning sub-network is connected to the input of the decoding layer, so that when the generative network is trained by fine-tuning, the jump structure is further optimized by at least one tuning sub-network, thereby improving the efficiency of fine-tuning training. In addition, at least one tuning sub-network is not in the encoder or the decoder, and there is no need to perform excessive parameter adjustments during fine-tuning training, thereby saving the computational overhead of fine-tuning training.

[0128] In other examples, fine-tuning training is performed on the generative network to be fine-tuned based on the training samples, including: performing multiple iterative training on the generative network to be fine-tuned based on the training samples to obtain a generative model, wherein in each iterative training, the difference between the input side sample in the training sample after the forward propagation output of the generative network to be fine-tuned and the input side sample in the training sample is determined, and while maintaining the various parameters of the encoder, the gradient of the difference in the decoder and the at least one tuning sub-network is back-propagated to adjust the various parameters in the decoder and the at least one tuning sub-network.

[0129] In other examples, based on the training samples, the generative network to be fine-tuned is fine-tuned to obtain a generative model, including: using the input side samples as the input of the encoder, using the output side samples as the output of the decoder, and adjusting the network parameters in at least one tuning sub-network to obtain a generative model.

[0130] In other examples, at least one tuning subnetwork corresponds to at least one conditional subnetwork, the first output of each conditional subnetwork is connected between the input of the corresponding tuning subnetwork and the input of the encoding layer, and the second output of each conditional subnetwork is connected between the output of the corresponding tuning subnetwork and the input of the decoding layer; based on the training samples, the generative network to be fine-tuned is fine-tuned to obtain a generative model, including: using the input side sample as the input of the encoder, using the output side sample as the output of the decoder, using the conditional sample as the input of at least one conditional subnetwork, adjusting the network parameters in at least one tuning subnetwork and the at least one conditional subnetwork to obtain a generative model.

[0131] In other examples, the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, processed by the conditional weights of the conditional sub-network in the multiple conditional sub-networks, and output to the input of the decoding layer.

[0132] In other examples, N skip connections are formed between the multiple encoding layers of the encoder and the multiple decoding layers of the decoder, each skip connection corresponding to M tuning sub-networks and M conditional sub-networks, and the mth conditional sub-network corresponds to the mth tuning sub-network. The mth conditional sub-network includes N network layers arranged sequentially from the input side to the output side, the nth network layer corresponds to the nth skip connection, and the output of the nth network layer is connected to the input of the mth tuning sub-network of the nth skip connection.

[0133] The following describes application development devices according to other embodiments of the present application in conjunction with FIG9 . The application development device corresponds to the application development method. Specifically, the application development device includes:

[0134] A creation module 910 is configured to create a user interface module of an application, wherein the user interface module is configured to generate description data based on at least user operation data, and return presentation data based on the generated data of the description data;

[0135] An acquisition module 920 is configured to obtain a calling interface for a generative model, wherein the generative model is obtained according to a model training method, and the calling interface is configured to return the generated data when called;

[0136] The embedding module 930 is configured to embed at least the calling interface of the generative model into the user interface module.

[0137] In other examples, the user interface module is also configured to call a service module to return the presentation data, the service module is configured to perform data processing on the generated data to obtain the presentation data, and the calling interface is configured to return the generated data when called by the service module.

[0138] The specific implementation of each module in the above-mentioned model training device or data processing device can refer to the corresponding description of the corresponding steps in the above-mentioned method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-mentioned device and module can refer to the corresponding process description in the above-mentioned method embodiment, and will not be repeated here.

[0139] 10 , a schematic structural diagram of an electronic device according to another embodiment of the present application is shown. The specific embodiments of the present application do not limit the specific implementation of the electronic device.

[0140] As shown in FIG. 10 , the electronic device may include: a processor 1002 for executing a program 1010 , a communications interface 1004 , a memory 1006 , and a communication bus 1008 .

[0141] The processor, the communication interface, and the memory communicate with each other via a communication bus.

[0142] Communication interface, used to communicate with other electronic devices or servers.

[0143] The processor is used to execute the program, and specifically can execute the relevant steps in the above method embodiment.

[0144] Specifically, the program may include program codes including computer operation instructions.

[0145] The processor may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or processors of different types, such as one or more CPUs and one or more ASICs.

[0146] Memory is used to store programs. The memory may include high-speed RAM memory and may also include non-volatile memory (non-volatile memory), such as at least one disk storage.

[0147] The program may include multiple computer instructions. Specifically, the program may enable the processor to execute operations corresponding to the model training method or data processing method described in any of the aforementioned method embodiments through multiple computer instructions.

[0148] The specific implementation of each step in the program can refer to the corresponding description of the corresponding steps and units in the above method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiment, and will not be repeated here.

[0149] The present application also provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the aforementioned method embodiments. The computer storage medium includes, but is not limited to, a compact disc read-only memory (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk.

[0150] An embodiment of the present application also provides a computer program product, including computer instructions, which instruct a computing device to perform operations corresponding to the model training method or data processing method in the above-mentioned multiple method embodiments.

[0151] In addition, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used to train the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0152] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0153] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.

[0154] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.

[0155] The above implementation methods are only used to illustrate the embodiments of the present application, and are not intended to limit the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present application, and the scope of patent protection of the embodiments of the present application should be defined by the claims. Industrial Applicability

[0156] The pre-trained generative model provided in the embodiment of the present application includes a residual structure, which includes a jump connection between the encoding layer and the decoding layer. At least one tuning sub-network is formed between the encoding layer and the decoding layer. Therefore, at least one tuning sub-network can efficiently train the residual structure by connecting to the encoding layer and the decoding layer of the jump connection, thereby obtaining a generative model with higher prediction reliability, so that more reliable data processing can be performed using the generative model.

Claims

1. A data processing method, comprising: Acquire a pre-trained generative model, wherein the generative model includes a residual structure, wherein the residual structure includes a skip connection between a coding layer and a decoding layer, and at least one tuning subnetwork is formed between the coding layer and the decoding layer; At least the description data is input into the generative model to obtain the generated data corresponding to the description data.

2. The data processing method according to claim 1, wherein: The residual structure further includes at least one conditional sub-network, and the at least one conditional sub-network corresponds to at least one tuning sub-network; At least inputting the description data into the generative model to obtain the generated data corresponding to the description data includes: The description data is input into the encoding layer of the generative model, the conditional data is input into at least one conditional sub-network of the generative model, and the generated data corresponding to the description data is obtained from the decoding layer of the generative model.

3. The data processing method according to claim 2, wherein: The at least one tuning subnetwork corresponds to at least one conditional subnetwork respectively, the first output of each conditional subnetwork is connected between the input of the corresponding tuning subnetwork and the input of the encoding layer, and the second output of each conditional subnetwork is connected between the output of the corresponding tuning subnetwork and the input of the decoding layer.

4. The data processing method according to claim 3, wherein: After the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, the conditional weights in the multiple conditional sub-networks are processed by the conditional sub-network and output to the input of the decoding layer.

5. The data processing method according to claim 1, wherein: The description data is a description text, and the generated data includes a plurality of video frames. The method further comprises: The multiple video frames are subjected to time-sequential coding processing to obtain a generated video of the description text.

6. An intelligent interaction method, comprising: Acquire the first description data input in the interactive interface; Inputting at least the first description data into a pre-trained generative model to obtain first generated data corresponding to the first description data, wherein the first generated data includes at least one of an image, a video frame, and a video, and the generative model includes a residual structure, the residual structure includes a jump connection between a coding layer and a decoding layer, and at least one tuning subnetwork is formed between the coding layer and the decoding layer; The first generated data is displayed in the question-and-answer interface.

7. The intelligent interaction method according to claim 6, wherein: The method further comprises: Acquire, from the interactive interface, second description data modified for the first generated data; At least the second description data is input into a pre-trained generative model to generate second generated data corresponding to the second description data as a modification result of the first generated data.

8. The intelligent interaction method according to claim 6, wherein: The first generated data is a picture or a video frame, Acquiring second description data modified for the first generated data from the interactive interface includes: Acquire a modification mark for the picture or video from the interactive interface; The modification mark is identified to generate the second description data, or the second description data is generated by combining the modification mark and the modification description text of the picture or video.

9. A model training method, comprising: Acquire an initially trained generative network, wherein the generative network includes an encoder and a decoder, wherein a skip connection is formed between an encoding layer of the encoder and a decoding layer of the decoder; Connecting at least one tuning subnetwork of the jump connection between the encoding layer and the decoding layer of the initially trained generative network to obtain a generative network to be fine-tuned, wherein the output of the encoding layer is connected to the input of the at least one tuning subnetwork, and the output of the at least one tuning subnetwork is connected to the input of the decoding layer; Based on the training samples, fine-tuning training is performed on the generative network to be fine-tuned to obtain a generative model.

10. The training method according to claim 9, wherein: Based on the training samples, fine-tuning training is performed on the generative network to be fine-tuned, including: Based on the training samples, the generative network to be fine-tuned is iteratively trained for multiple times to obtain a generative model, wherein in each iterative training, the difference between the input side sample in the training sample after the forward propagation output of the generative network to be fine-tuned and the input side sample in the training sample is determined, and while keeping the various parameters of the encoder, the various parameters in the decoder and the at least one tuning subnetwork are adjusted by back-propagating the gradient of the difference in the decoder and the at least one tuning subnetwork.

11. The training method according to claim 9, wherein: Based on the training samples, fine-tuning training is performed on the generative network to be fine-tuned to obtain a generative model, including: The input side samples are used as the input of the encoder, the output side samples are used as the output of the decoder, and the network parameters in the at least one tuning subnetwork are adjusted to obtain a generative model.

12. The training method according to claim 9, wherein: The at least one tuning subnetwork corresponds to at least one conditional subnetwork respectively, a first output of each conditional subnetwork is connected between an input of the corresponding tuning subnetwork and an input of the encoding layer, and a second output of each conditional subnetwork is connected between an output of the corresponding tuning subnetwork and an input of the decoding layer; Based on the training samples, fine-tuning training is performed on the generative network to be fine-tuned to obtain a generative model, including: The input side samples are used as the input of the encoder, the output side samples are used as the output of the decoder, the conditional samples are used as the input of at least one conditional sub-network, and the network parameters in the at least one tuning sub-network and the at least one conditional sub-network are adjusted to obtain a generative model.

13. The training method according to claim 12, wherein: After the second output of each conditional sub-network is fused with the output of the corresponding tuning sub-network, the conditional weights in the multiple conditional sub-networks are processed by the conditional sub-network and output to the input of the decoding layer.

14. An application development method, comprising: Creating a user interface module of the application, the user interface module being configured to generate description data based at least on user operation data, and return presentation data based on the generated data of the description data; Obtaining a calling interface of a generative model, wherein the generative model is obtained according to the model training method according to any one of claims 9 to 13, and the calling interface is configured to return the generated data when called; At least the calling interface of the generative model is embedded in the user interface module.

15. The method according to claim 14, wherein: The user interface module is also configured to call the service module to return the presentation data, the service module is configured to perform data processing on the generated data to obtain the presentation data, and the calling interface is configured to return the generated data when called by the service module.

16. An electronic device, comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1-15.

17. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 15 is implemented.

Citation Information

Patent Citations

  • Medical image segmentation method based on T-shaped attention structure

    CN111612790A

  • Liver segmentation method based on improved U-Net network

    CN115482242A

  • Methods for training generative large language models and for processing image tasks

    CN117114063A