Content generation method, apparatus, device, medium, and product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2026-05-08
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]在此提供一种内容生成方法、装置、设备、介质以及产品,解决了对特征向量处理时,内容输出效率较低的问题,达到了提高内容生成效率的技术效果
[0009] The beneficial effects of the above content generation method are as follows: First input information is obtained, which includes at least text information; a first feature vector is obtained based on a first encoding module; a third feature vector is obtained by denoising a second feature vector based on a first diffusion module; and the third feature vector is parsed based on a first decoding module. The first feature vector indicates the feature vector of the first input information in a first space, and the second feature vector indicates the vector obtained by concatenating the first feature vector with first noise. The first encoding module, the first diffusion module, and the first decoding module are modules in a first machine learning model. First content is output, which reflects the parsing result of the third feature vector. Therefore, the first input information can be processed accurately and quickly based on the first machine learning model, solving the problem that existing machine learning models can only process individual data units one by one, i.e., only one data unit's content can be generated each time, resulting in low content generation efficiency. This method achieves the technical effect of improving content generation efficiency.
Smart Images

Figure CN122528964A_ABST
Abstract
Description
Technical Field
[0001] This relates to the field of computer processing technology, and in particular to a content generation method, apparatus, device, medium, and product. Background Technology
[0002] Content generation tasks commonly employ machine learning models. However, when using machine learning models to process text information, only one data unit can be generated at a time, resulting in low output efficiency. Summary of the Invention
[0003] This invention provides a content generation method, apparatus, device, medium, and product that solves the problem of low content output efficiency when processing feature vectors, thereby achieving the technical effect of improving content generation efficiency.
[0004] In one scenario, this paper provides a content generation method that includes: Obtain first input information, wherein the first input information includes at least text information; The first feature vector is obtained based on the first encoding module, the second feature vector is denoised based on the first diffusion module to obtain the third feature vector, and the third feature vector is parsed based on the first decoding module. The first feature vector is used to indicate the feature vector of the first input information in the first space, and the second feature vector is used to indicate the vector after the first feature vector is concatenated with the first noise. The first encoding module, the first diffusion module and the first decoding module are modules in the first machine learning model. Output the first content, which reflects the parsing result of the third feature vector.
[0005] In one instance, this document also provides a content generation apparatus, which includes: The information receiving module is used to acquire first input information, wherein the first input information includes at least text information; The information processing module is used to obtain a first feature vector based on the first encoding module, to perform noise reduction processing on the second feature vector based on the first diffusion module to obtain a third feature vector, and to parse the third feature vector based on the first decoding module. The first feature vector is used to indicate the feature vector of the first input information in the first space, and the second feature vector is used to indicate the vector after concatenating the first feature vector with the first noise. The first encoding module, the first diffusion module, and the first decoding module are modules in the first machine learning model. The content output module is used to output first content, which reflects the parsing result of the third feature vector.
[0006] In one instance, this document also provides an electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the content generation method as described herein.
[0007] In one instance, this document also provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform the content generation methods described herein.
[0008] In another scenario, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the content generation method as described herein.
[0009] The beneficial effects of the above content generation method are as follows: First input information is obtained, which includes at least text information; a first feature vector is obtained based on a first encoding module; a third feature vector is obtained by denoising a second feature vector based on a first diffusion module; and the third feature vector is parsed based on a first decoding module. The first feature vector indicates the feature vector of the first input information in a first space, and the second feature vector indicates the vector obtained by concatenating the first feature vector with first noise. The first encoding module, the first diffusion module, and the first decoding module are modules in a first machine learning model. First content is output, which reflects the parsing result of the third feature vector. Therefore, the first input information can be processed accurately and quickly based on the first machine learning model, solving the problem that existing machine learning models can only process individual data units one by one, i.e., only one data unit's content can be generated each time, resulting in low content generation efficiency. This method achieves the technical effect of improving content generation efficiency. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments described herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 This is a schematic diagram of the architecture of an exemplary system in one scenario. Figure 2 This is an example diagram of the model architecture for a second machine learning model in one scenario. Figure 3 This is an example diagram of the model architecture for a third machine learning model in one scenario. Figure 4 This is a schematic diagram of the model architecture of the first machine learning model in one scenario; Figure 5 This is a flowchart illustrating a content generation method for one scenario. Figure 6 This is a schematic diagram of the noise reduction process of the first diffusion module under one scenario. Figure 7 This is a schematic diagram of the model architecture of the first machine learning model in another scenario; Figure 8 This is a schematic diagram of the model architecture of the first machine learning model in another scenario; Figure 9 This is a flowchart illustrating the training process of a second machine learning model in one scenario. Figure 10 A schematic diagram illustrating the training and adjustment process of a third machine learning model in one scenario; Figure 11 This is a schematic diagram of the noise reduction process of the second diffusion module in one scenario. Figure 12 Example diagram of the model architecture for a third machine learning model in another scenario; Figure 13 A flowchart illustrating the process of adjusting model parameters for a third machine learning model in one scenario; Figure 14 A flowchart illustrating the process of adjusting model parameters for a third machine learning model in another scenario; Figure 15 This is a schematic diagram of a content generation device in one scenario. Figure 16 This is a schematic diagram of the structure of an electronic device under one specific scenario. Detailed Implementation
[0012] The embodiments will now be described in more detail with reference to the accompanying drawings. While some embodiments are shown in the drawings, it should be understood that the technical solutions can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the technical solutions herein. It should be understood that the illustrated drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the technical solutions.
[0013] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.
[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.
[0015] It should be noted that the concepts of "first" and "second" mentioned are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0016] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this document, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this document in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described herein, based on the prompt message.
[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0021] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.
[0022] It is understood that the data involved in the technical solutions in this article (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] In this paper, the application scenarios of this technical solution can be illustrated by example: This solution can be applied to processing input information in a text modal, or processing input information in both text and non-text modalities, and outputting content adapted to the input information. The input information in the text modal can be text information. The non-text modal can be a video modal or an image modal. Correspondingly, the input information in the non-text modal can be video, images, etc.
[0024] Typically, text-based input information is discrete. To enable machine learning models to output content corresponding to this discrete information, this paper maps the text-based input information into a latent space, transforming the discrete information into vectors in a continuous feature space. These vectors are then processed to output content that matches the input information. It should be noted that even when processing multimodal input information, each input is processed separately in a single-modal format.
[0025] In some cases, the provided content generation method can be applied to Figure 1 The content generation system shown may include a client 101 and a server 102. The client 101 may include, but is not limited to, browsers, applications (Apps), HyperText Markup Language (HTML) applications, lightweight applications (also known as mini-programs), or cloud applications. The client 101 may be deployed on an electronic device and relies on the operation of that device or certain applications on the device to implement its functions. The electronic device may be, for example, a device with a display screen that supports information browsing, such as a smartphone, tablet, personal computer, or other client terminal. For ease of understanding, Figure 1 The client is primarily represented in the form of a device. Other types of applications can also be configured on the electronic device, such as media content publishing applications, session applications, etc. Server 102 can be one or more servers providing various services. That is, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; furthermore, it can be a server for a distributed system, a server integrating blockchain technology, a cloud server, or an intelligent cloud computing server or intelligent cloud host deployed with machine learning models, etc.
[0026] The content generation method described herein allows interaction between client 101 and server 102, such as receiving or sending messages. For example, in this paper, server 102 can receive first input information sent by client 101 based on an information carrier, and send the first content corresponding to the first input information to client 101 for display on the display interface.
[0027] It should be noted that the content generation method can be executed by client 101. In this case, the first machine learning model is deployed on client 101, and the first machine learning model deployed on client 101 processes the first input information. Alternatively, the content generation method can be executed by client 101 and server 102, with different functional parts of the corresponding content generation device deployed on client 101 and server 102 respectively; wherein, the information receiving module and content output module of the device are deployed on client 101, and the information processing module is deployed on server 102. In this case, the first input information sent by client 101 is processed based on the first machine learning model in the information processing module to provide feedback on the first content corresponding to the first input information. Client 101 and server 102 achieve data interaction and functional collaboration through network communication. It should be understood that... Figure 1 The number of clients and servers shown is for illustrative purposes only. Any number of clients and servers can be configured to meet specific implementation requirements.
[0028] Before introducing this scenario, we can first give a brief introduction to the machine learning model used in this scenario.
[0029] In one scenario, the model architecture of the first machine learning model is obtained based on the model architectures of the second and third machine learning models.
[0030] For ease of explanation, the model architecture of each machine learning model will be introduced in the following order: second machine learning model, third machine learning model, and first machine learning model.
[0031] The second machine learning model is a generative model based on a variational autoencoder (VAE). In one case, the second machine learning model includes a fourth encoding module and a fourth decoding module. See also Figure 2 The architecture of the second machine learning model can be as follows: Figure 2 As shown. The second machine learning model includes a fourth encoding module 201 and a fourth decoding module 202. The fourth encoding module 201 is used to perform latent space mapping on the input information to obtain the feature vector of the input information in the first space. The fourth decoding module 202 is used to parse the feature vector in the first space to obtain the output content corresponding to the input information. The second machine learning model can ensure the consistency of input and output.
[0032] The third machine learning model is a latent diffusion model (DIT) obtained based on the model architecture of the second machine learning model. In one scenario, the third machine learning model includes a second diffusion module, a first processing module, a fifth encoding module, and a fifth decoding module. The second diffusion module is located between the fifth encoding and fifth decoding modules. The fifth encoding module is identical to the fourth encoding module, and the fifth decoding module is identical to the fourth decoding module. The fourth encoding and fourth decoding modules are modules within the trained second machine learning model. The first processing module includes the fourth encoding module, used to instruct the adjustment of the model parameters of the second diffusion module based on its output.
[0033] The architecture of the third machine learning model can be combined with Figure 3 Let's find out. Figure 3 Example diagram of the model architecture for the third machine learning model; Figure 3 The third machine learning model in the model includes: a first processing module 301, a fifth encoding module 302, a fifth decoding module 303, and a second diffusion module 304.
[0034] The first processing module 301 performs latent space mapping on the input information to obtain feature vectors mapped into the first space. The fifth encoding module 302 has the same model parameters as the fourth encoding module of the trained second machine learning model. The fifth encoding module performs latent space mapping on the input information to obtain feature vectors mapped into the first space. The fifth decoding module 303 has the same model parameters as the fourth decoding module of the trained second machine learning model. The fifth decoding module parses the feature vectors in the first space.
[0035] The second diffusion module 304 is used for denoising the feature vector. The input to the second diffusion module is the feature vector obtained by concatenating noise from the feature vector output by the fifth encoding module. The purpose of introducing the second diffusion module is to ensure the accuracy and speed of the model's information processing through multi-step iterative denoising. The purpose of introducing the first processing module is to constrain the adjustment of the model parameters of the third machine learning model during the training phase, so as to avoid the model output deviating too much from the actual expectation.
[0036] The first machine learning model is obtained by removing the first processing module from the pre-trained third machine learning model architecture, and using the second diffusion module of the third machine learning model as the first diffusion module, the fifth encoding module as the first encoding module, and the fifth decoding module as the first decoding module. See also Figure 4 , Figure 4This is an example diagram of the model structure of the first machine learning model. Figure 4 The first machine learning model includes: a first encoding module 401, a first diffusion module 402, and a first decoding module 403.
[0037] The first encoding module 401 performs latent space mapping on the input information to obtain a feature vector corresponding to the first space. The first diffusion module 402 performs noise reduction on the feature vector after concatenation of noise in the first space. The first decoding module 403 parses the noise-reduced feature vector to obtain content with the same information dimension as the input information, thereby improving the content output efficiency while ensuring that the output content matches the input information.
[0038] The application process of the first machine learning model will be explained from an application perspective below. Figure 5 This is a flowchart illustrating a content generation method in one scenario. First content corresponding to the first input information can be obtained based on a first machine learning model. For a detailed description of the specific implementation method, please refer to the detailed explanation of this scenario. Technical terms that are the same as or corresponding to those in the above scenario will not be repeated here. Figure 5 As shown, the content generation method may specifically include: S510. Obtain first input information, the first input information including at least text information.
[0039] The first input information can be instruction description information obtained according to actual needs. The first input information can include at least text-based input information, i.e., text information. Optionally, the first input information can include text-based input information, or it can include both text-based and non-text-based input information. The non-text-based input information can include image-based and video-based input information. By inputting input information in one or more modalities, the content generation requirements in different scenarios can be met, thus expanding the scope of application for content generation.
[0040] For example, the first input information can be text-based input information, such as a question text that needs to be answered, like "What color is an apple?". The first input information can also be both text-based and image-based input information, such as an image that needs to be styled and the style description information that is expected to be converted.
[0041] S520: Obtain a first feature vector based on the first encoding module, perform noise reduction processing on the second feature vector based on the first diffusion module to obtain a third feature vector, and parse the third feature vector based on the first decoding module.
[0042] The first machine learning model includes a first encoding module, a first decoding module, and a first diffusion module. The first encoding module is used to obtain the first feature vector of the first input information in a first space. The first space can be understood as a low-dimensional, continuous abstract feature space, which can also be called the latent space. Since the first input information is usually high-dimensional, discrete data, in order to improve the efficiency of subsequent content generation, reduce computational complexity, and achieve the extraction of semantic features of the first input information, the first input information can be mapped based on the first space to obtain the feature vector in the latent space, i.e., the first feature vector. It can be understood that the first feature vector is obtained after feature extraction and redundancy filtering of the first input information, and is used to characterize the semantic features of the first input information.
[0043] The first diffusion module can be a module obtained by fusing a diffusion model with a Transformer architecture. Optionally, the first diffusion module can be a Diffusion Transformer Block (DiT) module. The first diffusion module is used to instruct the denoising process of the second feature vector to obtain the third feature vector. It can be understood that the first diffusion module is used to denoise the second feature vector. The output of the first diffusion module is the denoised second feature vector, i.e., the third feature vector.
[0044] The second feature vector indicates the vector obtained by concatenating the first feature vector with the first noise. This can be understood as the input to the first diffusion module being the second feature vector obtained by concatenating the first feature vector output by the first encoding module with the first noise. The first noise is Gaussian noise randomly sampled from a Gaussian distribution. The second feature vector can be obtained by concatenating the first noise and the first feature vector.
[0045] The noise concatenation process can be as follows: Random sampling is performed from a Gaussian distribution according to a preset noise concatenation ratio to obtain first noise. The first noise is then concatenated with a first feature vector to obtain a second feature vector. The preset noise concatenation ratio is the ratio of the length of the added first noise to the length of the first feature vector. Different preset noise concatenation ratios will result in different denoising outcomes for the second feature vector.
[0046] The first decoding module is used to instruct the parsing of the third feature vector. The parsing result of the first decoding module is the first content corresponding to the first input information. For example, if the first input information is the text to be completed, "Attention is", the corresponding first content could be "Attention is all you need!".
[0047] Specifically, the first encoding module based on the first machine learning model maps the first input information to convert discrete information into feature vectors in a continuous feature space, obtaining a first feature vector in the first space. First noise is obtained according to a preset noise concatenation ratio. The first noise is concatenated with the first feature vector to obtain a second feature vector with added noise. The second feature vector is then denoised using the first diffusion module to obtain a third feature vector. The third feature vector is then parsed using the first decoding module to obtain first content related to the first input information.
[0048] For example, in combination Figure 4 This describes how the first machine learning model processes the first input information. Figure 4 This is an example diagram of the model structure of the first machine learning model, which includes: a first encoding module 401, a first diffusion module 402, and a first decoding module 403.
[0049] The first encoding module 401 performs latent space mapping processing on the first input information to obtain a first feature vector corresponding to the first space, thereby converting high-dimensional discrete information into semantic feature vectors in a continuous feature space. The first feature vector is concatenated with first noise to obtain a second feature vector. The second feature vector is then denoised using the first diffusion module 402 to obtain a third feature vector, improving the efficiency of the denoising process while ensuring denoising accuracy. Finally, the first decoding module 403 analyzes the third feature vector in the first space to obtain first content with the same information dimension as the first input information, improving the quality and efficiency of the generated content.
[0050] S530. Output the first content, which reflects the analysis result of the third feature vector.
[0051] The first content can also be multimodal output information. For example, if the first input information is the question text to be answered, the first content can be the answer text corresponding to the question text. If the first input information is an image that needs to be style-transformed and the style description information to be transformed, the first content can be the style-transformed image.
[0052] Specifically, it outputs the first content determined by the first machine learning model. In one scenario, this first content can be displayed.
[0053] The beneficial effects of the above content generation method are as follows: First input information is obtained, which includes at least text information; a first feature vector is obtained based on a first encoding module; a third feature vector is obtained by denoising a second feature vector based on a first diffusion module; and the third feature vector is parsed based on a first decoding module. The first feature vector indicates the feature vector of the first input information in a first space, and the second feature vector indicates the vector obtained by concatenating the first feature vector with first noise. The first encoding module, the first diffusion module, and the first decoding module are modules in a first machine learning model. First content is output, which reflects the parsing result of the third feature vector. Therefore, the first input information can be processed accurately and quickly based on the first machine learning model, solving the problem that existing machine learning models can only process individual data units one by one, i.e., only one data unit's content can be generated each time, resulting in low content generation efficiency. This method achieves the technical effect of improving content generation efficiency.
[0054] Figure 6 This is a schematic diagram of the noise reduction process of the first diffusion module provided in this paper. Based on the above scheme, the noise reduction process of the first diffusion module on the second feature vector can also be described. The specific noise reduction process for the second feature vector can be found in the detailed description in this paper. Technical features that are the same as or similar to those described above will not be repeated here.
[0055] like Figure 6 As shown, the noise reduction process of the first diffusion module may specifically include: S610. Based on the first diffusion module, the second feature vector is denoised to obtain the fourth feature vector.
[0056] The fourth feature vector is the feature vector obtained by denoising the second feature vector.
[0057] Specifically, to ensure the accuracy of the first content, the second feature vector can be iteratively denoised based on the first diffusion module. That is, the second feature vector can be denoised first to remove the first noise and obtain the fourth feature vector.
[0058] S620: The fourth feature vector is concatenated with the second noise to obtain the fifth feature vector.
[0059] The second noise can be Gaussian noise randomly sampled from a Gaussian distribution. The fifth eigenvector can be a eigenvector obtained by concatenating the fourth eigenvector with the second noise.
[0060] Specifically, when concatenating the fourth feature vector with the second noise, the noise length of the second noise can be obtained according to a preset noise concatenation ratio. This second noise length is then concatenated with the fourth feature vector to obtain the fifth feature vector.
[0061] S630. The fifth feature vector is processed based on the first diffusion module to obtain the third feature vector.
[0062] Specifically, the fifth feature vector is denoised using the first diffusion module to obtain a denoised fifth feature vector. The denoised fifth feature vector is then used as the fourth feature vector, and the second noise concatenation process and the denoising process based on the first diffusion module are repeated until the denoised fifth feature vector is free of processable noise. The denoised fifth feature vector at this point is then used as the third feature vector.
[0063] For example, see Figure 7 , Figure 7 This is an example diagram of the model architecture for the first machine learning model. Figure 7 The first machine learning model includes: a first encoding module 701, a first diffusion module 702, and a first decoding module 703. Combined with... Figure 7 Taking the first input information as "Attention is" in text mode as an example, the first input information can be mapped into a first space to obtain a first feature vector. According to a preset noise concatenation ratio, first noise is concatenated to the first feature vector to obtain a second feature vector. The second feature vector is then denoised using a first diffusion module to obtain a fourth feature vector. For example, the fourth feature vector can be used to represent the feature corresponding to "Attention is all you". The fourth feature vector is then concatenated with second noise to obtain a fifth feature vector. The fifth feature vector is then denoised using the first diffusion module to obtain a denoised fifth feature vector. This denoised fifth feature vector can then be used to represent the feature corresponding to "Attention is all you need". The denoised fifth feature vector is used as the fourth feature vector, and the above second noise concatenation and denoising processing based on the first diffusion module are repeated until the denoised fifth feature vector has no further noise. This denoised fifth feature vector is then used as the third feature vector. This third feature vector can then be used to represent the feature corresponding to "Attention is all you need!". The first decoding module analyzes and processes the third feature vector to obtain the first content, which can be "Attention is all you need!" The beneficial effects of the above method are as follows: Based on the first diffusion module, the second feature vector is denoised to obtain a fourth feature vector; second noise is concatenated to the fourth feature vector to obtain a fifth feature vector. The fifth feature vector is then processed based on the first diffusion module to obtain a third feature vector. Through this layer-by-layer noise addition and denoising process, feature bias can be gradually corrected. Simultaneously, by adding noise to the denoised feature vector, prior knowledge and denoising constraints can be provided for subsequent denoising processes, gradually correcting denoising bias and ensuring denoising accuracy. This also solves the problem in existing technologies where only one data unit's content can be output at a time, facilitating the acquisition of the first content matching the input information in one go.
[0064] In one scenario, a first machine learning model can process multimodal first input information. This first machine learning model includes a first encoding module, a first decoding module, and a first diffusion module. The first encoding module includes a second encoding module and a third encoding module, and the first decoding module includes a second decoding module and a third decoding module. The output of the first encoding module is the input of the first diffusion module, and the output of the first diffusion module is the input of the first decoding module.
[0065] The second encoding module is related to the second decoding module and is used to process the first input information in the text modality; the third encoding module is related to the third decoding module and is used to process the first input information in the non-text modality.
[0066] See Figure 8 , Figure 8 This is the model architecture diagram of the first machine learning model provided in this article. Figure 8 The first encoding module 801 includes a second encoding module and a third encoding module, used to process input information of one or more modalities. The first decoding module 803 includes a second decoding module and a third decoding module. Specifically, after noise concatenation is applied to the output feature vector of the first encoding module 801, the input feature vector of the first diffusion module 802 is obtained, and the output feature vector of the first diffusion module 802 becomes the input feature vector of the first decoding module 803.
[0067] Combination Figure 8 The following describes the processing of first input information of one or more modalities by the first machine learning model, based on the model architecture shown. The second encoding module is used to process the first input information in the text modality, and the third encoding module is used to process the first input information in the non-text modality (image modality or video modality), respectively, to describe the processing of first input information of different modalities by the first machine learning model.
[0068] When the first input information is only in text mode, the second encoding module converts the text mode input information into a first space to obtain a first feature vector corresponding to the text mode. First noise is concatenated to the first feature vector corresponding to the text mode to obtain a second feature vector corresponding to the text mode. The second feature vector corresponding to the text mode is then denoised using a first diffusion module to obtain a third feature vector corresponding to the text mode. Finally, the third feature vector corresponding to the text mode is parsed using a second decoding module to obtain the first content corresponding to the text mode input information.
[0069] It should be noted that when the third encoding module is used to process the first input information in the image mode or video mode, since the video is composed of multiple images, the third encoding module and the corresponding third decoding module process the first input information in the image mode and video mode in a similar way. The following explanation will use the first input information as the input information in the text mode and the input information in the image mode.
[0070] When the first input information consists of text-based and image-based input information, the second encoding module converts the text-based input information into a first space to obtain a first sub-feature vector corresponding to the text modality, and the third encoding module converts the image-based input information into the first space to obtain a first sub-feature vector corresponding to the image modality. Noise concatenation is performed on the first sub-feature vectors corresponding to the text and image modalities to obtain a second feature vector. A first diffusion module performs noise reduction on the second feature vector to obtain second sub-feature vectors corresponding to the text and image modalities. The second sub-feature vectors corresponding to the text and image modalities are then parsed using a second decoding module and a third decoding module, respectively, to obtain the first content corresponding to the first input information.
[0071] The beneficial effects of the above method are as follows: When the first machine learning model includes a first encoding module, a first decoding module, and a first diffusion module, the first encoding module is configured to include a second encoding module and a third encoding module, and the first decoding module includes a second decoding module and a third decoding module; the output of the first encoding module is the input of the first diffusion module, and the output of the first diffusion module is the input of the first decoding module. Based on this, it is possible to perform targeted processing of input information under corresponding modalities based on different encoding modules, avoid information interference caused by the joint processing of multimodal input information, ensure the accuracy of processing multimodal input information, and facilitate the individual module parameter adjustment for the encoding module or decoding module corresponding to a certain modality, ensuring the efficiency of model training and the accuracy of model output results.
[0072] Based on the above, the first machine learning model includes a first encoding module and a first decoding module. Therefore, the second machine learning model can be trained first to obtain the encoding and decoding modules related to the first machine learning model.
[0073] Figure 9 This is a flowchart illustrating the training process of a second machine learning model under one scenario. Based on the above scenario, the training process of the second machine learning model is explained; for details, please refer to the detailed explanation in this article. Technical features that are the same as or similar to those described above will not be repeated here.
[0074] S910. Obtain second input information, the second input information including first information and second information, the second information being used to indicate information after at least a portion of the content in the first information is masked.
[0075] To achieve content reconstruction and prediction, the second input information may include both the first and second information. The first information may include input information from multiple modalities. This can be understood as including at least text-modal input information. Optionally, the first information may be text-modal input information, or it may be both text-modal and non-text-modal input information. The first information enables the second machine learning model to perform content reconstruction. The second information is the input information obtained after randomly masking the first information. Taking text information as an example, the second information may be incomplete text information after masking a portion of the text content. The second information enables the second machine learning model to perform contextual semantic learning based on the unmasked portion of the content to achieve content prediction.
[0076] Since the first information can include input information from multiple modalities, and the second information is the first information after masking, and the training method is the same for the second input information from different modalities, we will elaborate on the second input information from different modalities here.
[0077] When the second input information is in text mode, the first information obtained is text information, and the second information is the masked text information after the masked portion of the text content. For example, with Figure 2 For example, the first message could be “Attention is all you need!”, and the second message could be the text information obtained after masking “Attention”, “you”, and “need”.
[0078] S920. Based on the fourth encoding module in the second machine learning model, the second input information is processed to obtain the sixth and seventh feature vectors of the second input information in the first space.
[0079] The sixth feature vector is related to the first information. This can be understood as the feature vector obtained by the fourth encoding module of the second machine learning model through latent space mapping of the first information. The seventh feature vector is related to the second information. This can also be understood as the feature vector obtained by the fourth encoding module of the second machine learning model through latent space mapping of the second information.
[0080] Specifically, the first and second information are processed by the fourth encoding module in the second machine learning model to map the first information into a first space, obtaining a sixth feature vector corresponding to the first information. Similarly, the second information is mapped into the first space to obtain a seventh feature vector corresponding to the second information. Based on the above, the process of converting discrete information into feature vectors in a continuous space is achieved, enabling the extraction of semantic features.
[0081] S930. Based on the fourth decoding module in the second machine learning model, the sixth and seventh feature vectors are processed to obtain the second content.
[0082] The second part includes the third and fourth information. The third information is used to characterize the reconstructed content of the sixth feature vector. This can be understood as the reconstructed content obtained by the fourth decoding module parsing and processing the sixth feature vector from the first space. The fourth information is used to characterize the predicted content of the seventh feature vector. Correspondingly, the fourth information is the predicted content obtained by the fourth decoding module parsing and processing the seventh feature vector from the first space.
[0083] Specifically, the fourth decoding module based on the second machine learning model analyzes the sixth feature vector to obtain the reconstructed content corresponding to the first information, i.e., the third information. And, based on the fourth parsing module, the seventh feature vector is analyzed to obtain the predicted content corresponding to the second information, i.e., the fourth information.
[0084] The following section uses text modality as an example to introduce the specific processing of second input information by the second machine learning.
[0085] When the first information is text information and the second information is masked text information, the fourth encoding module based on the second machine learning model performs latent space mapping on the text information to obtain the sixth feature vector in the first space, and performs latent space mapping on the masked text information based on the fourth encoding module to obtain the seventh feature vector in the first space.
[0086] The fourth decoding module analyzes the sixth feature vector in the first space to obtain the reconstructed text corresponding to the text information, i.e., the third information. The fourth decoding module also analyzes the seventh feature vector in the first space to obtain the predicted text corresponding to the masked text information, i.e., the fourth information.
[0087] Optionally, by training the second machine learning model, it can learn the correspondence between the latent representation (feature vectors in the first space) and the input information and output content in the first space. That is, the second input information is latently mapped using the fourth encoding module, and parsed based on the feature vectors in the first space using the fourth decoding module to reconstruct the third information and predict the fourth information. The encoding process of the fourth encoding module is described in the following function: ; in, It is used to characterize the latent space mapping process of the second input information. This indicates the second input information. Let represent the sixth and seventh eigenvectors in the first space.
[0088] The decoding process of the fourth decoding module is described in the following functions: ; in, This is used to characterize the analysis of the sixth and seventh eigenvectors of the first space. This indicates the second input information. Let the sixth and seventh eigenvectors in the first space be represented. This indicates the second content.
[0089] For example, with Figure 2 For example, if the first information can be the text "Attention is all you need!", and the second information can be the masked text obtained after masking "Attention", "you", and "need", then the first and second information are input into the second machine learning model. The fourth encoding module of the second machine learning model performs latent space mapping processing on the first and second information respectively, obtaining the sixth feature vector corresponding to the first information and the seventh feature vector corresponding to the second information. The fourth decoding module then analyzes the sixth feature vector to reconstruct the third information corresponding to the first information, i.e., Figure 2 The text shows "Attention is all you need." and how the fourth decoding module analyzes the seventh feature vector to predict the fourth information corresponding to the second information. Figure 2 The image shows "Cola is all you need!" S940. Based on the first information and the second content, adjust the parameters of the fourth encoding module and the fourth decoding module, and use the model parameters of the fourth encoding module and the fourth decoding module when the convergence condition is met as the first model parameters.
[0090] The first model parameter is used to characterize the model parameters of the trained second machine learning model.
[0091] Specifically, based on the first information and the second content, the model parameters of the fourth encoding module and the fourth decoding image in the second machine learning model are adjusted so that when the training error of the loss function associated with the first information and the second content converges, the trained second machine learning model is obtained, and the model parameters of the fourth encoding module and the fourth decoding module of the trained second machine learning model are used as the first model parameters.
[0092] Under a request, the parameters of the fourth encoding module and the fourth decoding module are adjusted based on the first information and the second content as follows: a first loss is obtained based on the first information and the third information, the first loss being used to reflect the reconstruction loss of the second machine learning model; a second loss is obtained based on the first information and the fourth information, the second loss being used to reflect the prediction loss of the second machine learning model; a third loss is obtained based on the first information, the second information, and the third information; and the parameters of the fourth encoding module and the fourth decoding module are adjusted based on the first loss, the second loss, and the third loss to obtain the first model parameters.
[0093] The first loss characterizes the difference between the first and third pieces of information, and is used to characterize the reconstruction loss. The reconstruction loss can be used to ensure that the third information obtained after encoding and decoding the first information remains consistent with the original first information. The second loss characterizes the difference between the first and fourth pieces of information, and is used to characterize the prediction loss. The prediction loss can be used to supervise the second machine learning model to accurately predict the semantics of the masked missing content in the second information. The third loss can be the KL divergence loss obtained based on the first, second, and third pieces of information, used to regularize the first spatial distribution.
[0094] Specifically, reconstruction loss processing is applied to the first and third information to obtain the first loss. Optionally, the reconstruction loss processing can be characterized by the following function: ; in, Indicates the first loss. It is used to characterize the latent space mapping process of the second input information. This indicates the second input information. Let represent the sixth and seventh eigenvectors in the first space. This represents the analytical processing of the sixth and seventh eigenvectors in the first space. Let represent the sixth and seventh eigenvectors in the first space.
[0095] Applying a prediction loss to the first and fourth pieces of information yields a second loss. Optionally, the prediction loss processing can be characterized by the following function: ; in, Indicates the second loss. Indicates BERT-style mask loss, This represents the weighting coefficient.
[0096] Loss processing is applied to the first, second, and third information to obtain the third loss. Optionally, the third loss can be characterized by the following function: ; in, Indicates the third loss. It is used to characterize the latent space mapping process of the second input information. This indicates the second input information. Let the sixth and seventh eigenvectors in the first space be represented. This indicates the handling of KL divergence loss. This represents the weighting coefficient.
[0097] The first loss, second loss, and third loss are summed to obtain the total loss. Optionally, the total loss can be expressed as: ; in, Indicates the total loss. Indicates the first loss. Indicates the third loss. This indicates the second loss.
[0098] The parameters of the second machine learning model are adjusted using the total loss. Specifically, the convergence of the total loss function can be used as a training objective, such as whether the training error is less than a preset error, whether the error change tends to stabilize, or whether the current number of iterations equals a preset number. If convergence is detected (e.g., the training error of the loss function is less than the preset error, or the error change trend tends to stabilize), it indicates that the second machine learning model has completed training, and iterative training can be stopped. If convergence is not detected, other second input information can be obtained to continue training the second machine learning model until the training error of the loss function is within a preset range. When the training error of the loss function converges, the trained second machine learning model is obtained, and the model parameters of the fourth encoding module and the fourth decoding module of the trained second machine learning model are used as the first model parameters. This loss processing prevents the fourth encoding module from collapsing at the semantic level, while the fourth decoding module only needs to memorize the surface text, ensuring consistency between input and output. During training, the second machine learning model does not compress the feature sequence length. Furthermore, to prevent information leakage and facilitate subsequent second content generation, the fourth encoding module and the fourth decoding module must strictly adhere to causal relationships.
[0099] For example, referring to the examples above, see [link to previous section]. Figure 2 Given the following scenarios: the first information is "Attention is all you need!", the second information is the masked text information obtained after the mask "Attention", "you", and "need", the third information is "Attention is all you need.", and the fourth information is "Cola is all you need!", a first loss is obtained based on the first information "Attention is all you need!" and the third information "Attention is all you need.", a second loss is obtained based on the first information "Attention is all you need!" and the fourth information "Cola is all you need!", and a third loss is obtained based on the first information, the second information, and the third information. The parameters of the fourth encoding module and the fourth decoding module are adjusted using the first loss, the second loss, and the third loss to obtain a trained second machine learning model. The model parameters of the fourth encoding module of the trained second machine learning model are used as the first model parameters.
[0100] The beneficial effects of the above method are as follows: First, the second input information is obtained. Then, the second input information is processed by the fourth encoding module in the second machine learning model to obtain the sixth and seventh feature vectors of the second input information in the first space. Next, the sixth and seventh feature vectors are processed by the fourth decoding module in the second machine learning model to obtain the second content. Based on the first information and the second content, the parameters of the fourth encoding and fourth decoding modules are adjusted, and the model parameters of the fourth encoding and fourth decoding modules when the convergence condition is met are used as the first model parameters. By training the second machine learning model, the fourth decoding and fourth encoding modules can achieve content reconstruction and content prediction processing, facilitating module transfer of the trained fourth encoding and fourth decoding modules and providing data support for the subsequent application of the first machine learning model.
[0101] Based on the above, it can be seen that, after obtaining a trained second machine learning model, the fourth encoding module and the fourth decoding module of the second machine learning model can be reused to construct a third machine learning model. The third machine learning model is then trained to obtain the encoding module, decoding module, and diffusion module related to the first machine learning model.
[0102] It should be noted that during the training of the third machine learning model, the model parameters of the first processing module do not need to be adjusted. This can be understood as using the output of the first processing module as a benchmark to adjust the model parameters of other modules.
[0103] Figure 10 This is a flowchart illustrating the training and tuning process of a third-party machine learning model under one specific scenario. Based on the above scenario, the training and tuning process of the third-party machine learning model is explained in detail in this paper. Technical features that are the same as or similar to those described above will not be repeated here.
[0104] S1010. Based on the third machine learning model and the third input information, adjust the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module.
[0105] The third input information may include at least the input information in the text modality. Optionally, the third input information may include the input information in the text modality and the mask information corresponding to the text modality. Alternatively, the third input information may include the input information in the text modality and the mask information corresponding to the text modality, as well as the input information in the non-text modality and the mask information corresponding to the non-text modality. The second model parameters may be the model parameters of the third machine learning model.
[0106] Specifically, the third input information is processed based on the third machine learning model. Based on the feature vectors obtained during the processing, the model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module of the third machine learning model are adjusted to obtain a trained third machine learning model.
[0107] In one scenario, based on the third machine learning model and the third input information, the second model parameters of the fifth encoding module, the fifth decoding module, and the first diffusion module are adjusted as follows: the third input information is mapped into the first space based on the fifth encoding module to obtain the eighth feature vector; the eighth feature vector is parsed based on the fifth decoding module to obtain the ninth feature vector; the third input information is mapped into the first space based on the first processing module to obtain the tenth feature vector; the eighth feature vector after splicing third noise is denoised based on the second diffusion module to obtain the eleventh feature vector; the second model parameters are adjusted based on the eighth, ninth, tenth, and eleventh feature vectors; the second model parameters are used to reflect the model parameters of the third machine learning model.
[0108] The eighth eigenvector can be the eigenvector obtained by mapping the third input information to the first space. The ninth eigenvector is the eigenvector obtained by parsing the eighth eigenvector in the first space. The tenth eigenvector can be the eigenvector obtained by mapping the third input information to the first space. The third noise can be Gaussian noise obtained by randomly sampling from a Gaussian distribution. Correspondingly, the eleventh eigenvector can be understood as the eigenvector obtained by denoising the eighth eigenvector after concatenating the third noise.
[0109] Specifically, the third input information is input to the fifth encoding module, which maps the third input information into the first space to obtain the eighth feature vector. Then, the eighth feature vector is parsed and processed by the fifth decoding module to obtain the ninth feature vector. Finally, the third input information is input to the first processing module, which maps the third input information into the first space to obtain the tenth feature vector.
[0110] The third noise is randomly sampled from a Gaussian distribution. The eighth feature vector is then concatenated with noise based on this third noise to obtain the concatenated eighth feature vector. It should be noted that block-level noise addition can be applied to the eighth feature vector. Therefore, the number of feature blocks corresponding to the eighth feature vector can be obtained using the following formula for block-level noise addition: ; in, The length of the eighth eigenvector is represented by the vector length. This represents the length of the feature block obtained by dividing the eighth feature vector into blocks. This indicates the number of feature blocks obtained by dividing the eighth feature vector into blocks.
[0111] Furthermore, for a noisy feature block, it includes all clean feature blocks preceding the current feature block and the noisy feature block itself. For the b-th noisy feature block, it can be represented as: ; in, This indicates that the gradient has stopped. This represents the b-th feature block after adding noise. This represents all clean feature blocks preceding the current b-th feature block. Let represent the visible set consisting of the current b-th noisy feature block and all clean feature blocks preceding the current b-th feature block. This visibility constraint ensures the causal dependencies between feature blocks, guaranteeing the accuracy of the denoising order.
[0112] The eighth feature vector, after being concatenated with the third noise, is denoised block by block by the second diffusion module to obtain the eleventh feature vector. Based on the eighth, ninth, tenth, and eleventh feature vectors, the loss corresponding to the third machine learning model is calculated, and the parameters of the second model are adjusted based on the loss. When the training error of the loss function of the third machine learning model converges, the adjustment of the second model parameters is complete, and the trained third machine learning model is obtained. By calculating the loss based on the output results of different modules, the model parameters of the third machine learning model are adjusted in a targeted manner according to the corresponding loss, ensuring the training accuracy of the third machine learning model.
[0113] Furthermore, based on the eighth, ninth, tenth, and eleventh feature vectors, the parameters of the second model are adjusted, including: obtaining a fourth loss based on the eighth and ninth feature vectors; obtaining a fifth loss based on the eighth and tenth feature vectors; obtaining a sixth loss based on the eleventh and twelfth feature vectors, where the twelfth feature vector indicates the feature vector corresponding to the expected content, which is related to the third input information; adjusting the second model parameters of the fifth encoding module and the fifth decoding module based on the fourth loss; adjusting the second model parameters of the fifth encoding module based on the fifth loss; and adjusting the second model parameters of the second diffusion module based on the sixth loss.
[0114] The fourth loss characterizes the degree of difference between the eighth and ninth feature vectors. It allows for parameter tuning of the fifth encoding and decoding modules, ensuring the accuracy of the latent space mapping in the fifth encoding module and the parsing accuracy of the fifth decoding module through regularized latent learning. The twelfth feature vector indicates the feature vector corresponding to the desired content. The fifth loss characterizes the degree of difference between the eighth and tenth feature vectors. It helps suppress potential drift in the fifth encoding module and avoids abnormal model parameter tuning.
[0115] The twelfth eigenvector can be obtained by truncating the gradient of the eighth eigenvector. The sixth loss characterizes the degree of difference between the eleventh and twelfth eigenvectors. The parameters of the second diffusion module are adjusted using the sixth loss to enable the second diffusion module to learn block-level conditional priors, thus ensuring the denoising accuracy of the second diffusion module.
[0116] The following describes the model parameter adjustment process. After the fifth encoding module performs latent space mapping based on the third input information to obtain the eighth feature vector, and the fifth decoding module analyzes and processes the eighth feature vector to obtain the ninth feature vector, loss processing is performed based on the eighth and ninth feature vectors to obtain the fourth loss. The fourth loss can be characterized by the following function: ; in, This indicates the fourth loss. This represents the weighting coefficient corresponding to the fourth loss. It is used to characterize the latent space mapping process of the third input information. This indicates the third input information. This represents the eighth eigenvector in the first space. This indicates that the eighth eigenvector in the first space is analyzed to obtain the ninth eigenvector. Indicates BERT-style mask loss, This represents the weighting coefficients corresponding to the mask loss.
[0117] For the eighth and tenth eigenvectors, the fifth loss can be obtained using the following function: ; in, Indicates the fifth loss. This represents the weighting coefficient corresponding to the fifth loss. This is used to characterize the latent space mapping process performed by the fifth encoding module on the third input information to obtain the eighth feature vector. This indicates the third input information. This represents the eighth eigenvector in the first space. This indicates that the first processing module performs latent space mapping processing on the third input information to obtain the tenth feature vector. This indicates the handling of KL divergence loss.
[0118] Accordingly, a loss is applied to the eleventh feature vector output by the second diffusion module and the twelfth feature vector obtained through gradient truncation to obtain the sixth loss. Optionally, the sixth loss can be expressed as: ; in, This indicates the sixth loss. This represents the weighting coefficient corresponding to the sixth loss. This represents the loss value obtained based on the eleventh and twelfth eigenvectors.
[0119] The second model parameters of the fifth encoding and decoding modules are adjusted based on the fourth loss to ensure the accuracy of latent space mapping in the fifth encoding module and the parsing accuracy in the fifth decoding module. The second model parameters of the fifth encoding module are also adjusted based on the fifth loss to suppress potential drift and avoid abnormal parameter adjustments. Furthermore, the second model parameters of the second diffusion module are adjusted based on the sixth loss to ensure denoising accuracy. Targeted parameter adjustments are then made to the corresponding modules of the third machine learning model based on these multiple losses to ensure the model processing accuracy of the third machine learning model.
[0120] Optionally, the total loss can be obtained based on the fourth, fifth, and sixth losses, and the model parameters of the fifth encoding module, fifth decoding module, and second diffusion module of the third machine learning model can be adjusted based on the total loss. The total loss can be expressed as: ; in, This indicates the fourth loss. This indicates the sixth loss. Indicates the fifth loss. This indicates the total loss.
[0121] S1020. In response to the convergence condition of the third machine learning model, the fifth encoding module is used as the first encoding module, the second diffusion module is used as the first diffusion module, and the fifth decoding module is used as the first decoding module to obtain the first machine learning model.
[0122] Specifically, the third machine learning model reaches the convergence condition when the training error of the third machine learning model is less than the preset error, or when the error change trend tends to stabilize.
[0123] Specifically, when the third machine learning reaches the convergence condition, the first processing module in the third machine learning model is removed, and the fifth encoding module is used as the first encoding module, the second diffusion module is used as the first diffusion module, and the fifth decoding module is used as the first decoding module to obtain the first machine learning model.
[0124] The beneficial effects of the above method are as follows: Based on the third machine learning model and the third input information, the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module are adjusted. When the third machine learning model reaches the convergence condition, the fifth encoding module is used as the first encoding module, the second diffusion module as the first diffusion module, and the fifth decoding module as the first decoding module to obtain the first machine learning model. During the training process of the third machine learning model, multiple losses are calculated to adjust the parameters of the fifth encoding module and the fifth decoding module through the fourth loss, ensuring the accuracy of the latent space mapping of the fifth encoding module and the parsing accuracy of the fifth decoding module. The fifth loss can suppress the potential drift of the fifth encoding module and avoid abnormal model parameter adjustment. Furthermore, the parameters of the second diffusion module are adjusted through the sixth loss so that the second diffusion module learns block-level conditional priors, ensuring the denoising accuracy and efficiency of the second diffusion module. Moreover, by adjusting the modules of the trained third machine learning model, the first machine learning model is obtained, providing model support for subsequent processing of the first input information and ensuring the accuracy and efficiency of the output content.
[0125] Figure 11 This is a flowchart illustrating the noise reduction process of the second diffusion module provided in this paper. Based on the above scheme, the process of progressively reducing noise in the eighth feature vector after splicing the third noise using the second diffusion module can also be described. The specific noise reduction process for the second feature vector can be found in the detailed description in this paper. Technical features that are the same as or similar to those described above will not be repeated here.
[0126] S1110. Based on the second diffusion module, the eighth feature vector after splicing the third noise is denoised to obtain the thirteenth feature vector. The thirteenth feature vector is used to indicate the processing result of at least one block after the denoising process.
[0127] At least one block can be understood as a feature block in the eighth feature vector concatenated with third noise. The processing result can be a denoising result. Optionally, the eighth feature vector is divided into blocks to obtain at least one feature block. For at least one feature block, third noise concatenation is performed to obtain the eighth feature vector concatenated with third noise. Correspondingly, during denoising, the eighth feature vector concatenated with third noise can be denoised to denoise the feature block in the eighth feature vector concatenated with third noise, thus obtaining the thirteenth feature vector.
[0128] To train the denoising capability of the second diffusion module, block-based denoising can be performed based on the second diffusion module. Loss processing is then applied based on the feature vector obtained after each denoising step and the twelfth feature vector. Specifically, the block-based denoising process can be as follows: The eighth feature vector, which is concatenated with third noise, is denoised based on the second diffusion module to remove at least one block of the eighth feature vector containing the third noise, thus obtaining the processing result corresponding to at least one block, i.e., obtaining the thirteenth feature vector.
[0129] For example, let's take "I like to eat apples" as the third input information. By dividing the eighth feature vector corresponding to the third input information into blocks, we can obtain feature blocks corresponding to "I," "like," "eat," and "apple." Accordingly, to enable the second diffusion module to have accurate denoising capabilities, we can first perform third noise concatenation on the feature block corresponding to "apple," and based on the feature blocks corresponding to "I," "like," "eat," and the concatenated third noise feature block corresponding to "apple," we obtain the eighth feature vector after concatenating the third noise. The second diffusion module then performs denoising on the eighth feature vector after concatenating the third noise to remove the third noise corresponding to the feature block corresponding to "apple," obtaining the thirteenth feature vector.
[0130] S1120. Based on the first processing unit, loss processing is performed on the thirteenth feature vector and the twelfth feature vector to obtain the seventh loss, and the parameters of the second diffusion module are adjusted based on the seventh loss.
[0131] The first processing unit can be a unit within the second diffusion module, used to perform loss processing on the thirteenth and twelfth feature vectors. The seventh loss is used to characterize the degree of difference between the thirteenth and twelfth feature vectors, and can reflect the block denoising capability of the second diffusion module.
[0132] Specifically, the thirteenth and twelfth feature vectors are processed by the first processing unit to obtain the seventh loss. The model parameters of the second diffusion module are then adjusted using the seventh loss to obtain the adjusted second diffusion module.
[0133] S1130. Based on the parameter-adjusted second diffusion module, the thirteenth feature vector of the spliced fourth noise is denoised to obtain the fourteenth feature vector.
[0134] The fourth noise can be Gaussian noise randomly sampled from a Gaussian distribution. The fourteenth eigenvector can be used to indicate the processing result of at least one block obtained after denoising the thirteenth eigenvector of the concatenated fourth noise.
[0135] Specifically, the second diffusion module with adjusted parameters performs noise reduction processing on the thirteenth feature vector spliced with the fourth noise to remove at least one block of the thirteenth feature vector spliced with the fourth noise, and obtains the processing result corresponding to at least one block, that is, obtains the fourteenth feature vector.
[0136] For example, in conjunction with the above example, if the thirteenth feature vector obtained by removing the third noise includes: the feature block corresponding to "I", the feature block corresponding to "like", the feature block corresponding to "eat", and the feature block corresponding to "apple", then the feature block corresponding to "I" can be subjected to fourth noise concatenation processing. Based on the feature blocks corresponding to "I", "like", "eat", and "apple" concatenated with the fourth noise, the thirteenth feature vector after concatenating with the fourth noise is obtained. The thirteenth feature vector after concatenating with the fourth noise by the second diffusion module is then subjected to denoising processing to remove the fourth noise corresponding to the feature block corresponding to "I", thus obtaining the fourteenth feature vector.
[0137] S1140. In response to the fourteenth eigenvector satisfying the first condition, the fourteenth eigenvector is taken as the eleventh eigenvector.
[0138] The first condition can be that the fourteenth feature vector has no feature blocks that can be spliced with noise.
[0139] Specifically, if there are no feature blocks with spliced Gaussian noise in the fourteenth feature vector, the fourteenth feature vector satisfies the first condition. Then, the fourteenth feature vector is used as the eleventh feature vector, and loss processing is performed based on the eleventh and twelfth feature vectors.
[0140] If there is a feature block in the fourteenth feature vector that can be spliced with Gaussian noise, then the fourteenth feature vector is taken as the thirteenth feature vector, and the above noise splicing process and the noise reduction process based on the second diffusion module after parameter adjustment are repeated until there is no feature block in the fourteenth feature vector that can be spliced with Gaussian noise. If the fourteenth feature vector satisfies the first condition, then the fourteenth feature vector is taken as the eleventh feature vector.
[0141] The beneficial effects of the above method are as follows: The eighth feature vector of the concatenated third noise is denoised using the second diffusion module to obtain the thirteenth feature vector; the thirteenth and twelfth feature vectors are subjected to loss processing using the first processing unit to obtain the seventh loss, and the parameters of the second diffusion module are adjusted based on the seventh loss; the thirteenth feature vector of the concatenated fourth noise is denoised using the parameter-adjusted second diffusion module to obtain the fourteenth feature vector; in response to the fourteenth feature vector satisfying the first condition, the fourteenth feature vector is used as the eleventh feature vector, and loss processing is performed based on the eleventh and twelfth feature vectors. By training the second diffusion module using the above iterative denoising and noise reduction method, the learning difficulty of the second diffusion module can be reduced, and the denoising error of the second diffusion module can be gradually corrected. While ensuring the denoising accuracy of the second diffusion module, the content corresponding to multiple denoised feature blocks can be obtained at once, improving the denoising efficiency of the second diffusion module.
[0142] In one scenario, the fifth encoding module includes a sixth and a seventh encoding module, and the fifth decoding module includes a sixth and a seventh decoding module. The sixth encoding module is associated with the sixth decoding module and is used to process the third input information in the text modality. The seventh encoding module is associated with the seventh decoding module and is used to process the third input information in the non-text modality. The text modality is different from the non-text modality.
[0143] See Figure 12 , Figure 12 This is an example diagram of the model architecture for the third machine learning model. Figure 12 The third machine learning model includes a first processing module 1201, a fifth encoding module 1202, a fifth decoding module 1203, and a second diffusion module 1204.
[0144] The fifth encoding module includes a sixth encoding module and a seventh encoding module, and the fifth decoding module includes a sixth decoding module and a seventh decoding module. The sixth encoding module performs latent space mapping on the third input information in the text modality to convert discrete information into feature vectors in a continuous feature space. The sixth decoding module analyzes the output features of the sixth encoding module. The seventh encoding module performs latent space mapping on the third input information in the non-text modality to convert discrete information into feature vectors in a continuous feature space. The seventh decoding module analyzes the output features of the seventh encoding module.
[0145] Figure 13This is a flowchart illustrating the model parameter adjustment process for the third machine learning model provided in this paper. Based on the above, the adjustment process for the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module can also be described. For a detailed description of the first module adjustment process, please refer to the description in this paper. Technical features that are the same as or similar to those described above will not be repeated here.
[0146] S1310, In response to the correlation between the third input information and the text modality, the second model parameters of the sixth encoding module, the sixth decoding module, and the second diffusion module are adjusted based on the third input information and the third machine learning model.
[0147] Specifically, if the third input information is text-based, the sixth encoding module maps the third input information to the first space to obtain the eighth feature vector corresponding to the text modality. The sixth decoding module then analyzes and processes the eighth feature vector to obtain the ninth feature vector corresponding to the text modality. Based on the eighth and ninth feature vectors in the text modality, the fourth loss corresponding to the text modality is obtained, and the model parameters of the sixth encoding and sixth decoding modules are adjusted based on the fourth loss.
[0148] Furthermore, the first processing module maps the input information in the text modality to the first space to obtain the tenth feature vector corresponding to the text modality. The second diffusion module performs noise reduction processing on the eighth feature vector in the text modality after concatenating with third noise to obtain the eleventh feature vector. Based on the eighth and tenth feature vectors in the text modality, a fifth loss is obtained, and the model parameters of the sixth encoding module are adjusted based on the fifth loss. Additionally, based on the eleventh and twelfth feature vectors in the text modality, a sixth loss is obtained, and the model parameters of the second diffusion module are adjusted based on the sixth loss.
[0149] S1320, In response to the correlation between the third input information and the non-text modality, the second model parameters of the seventh encoding module, the seventh decoding module, and the second diffusion module are adjusted based on the third input information and the third machine learning model.
[0150] Specifically, if the third input information is in a non-textual modality, the seventh encoding module maps the third input information to the first space to obtain the eighth feature vector corresponding to the non-textual modality. The seventh decoding module then analyzes the eighth feature vector corresponding to the non-textual modality to obtain the ninth feature vector. Based on the eighth and ninth feature vectors in the non-textual modality, a fourth loss corresponding to the non-textual modality is obtained, and the model parameters of the seventh encoding and seventh decoding modules are adjusted based on this fourth loss.
[0151] Furthermore, the first processing module maps the input information from the non-textual modality to the first space, obtaining the tenth feature vector corresponding to the non-textual modality. The second diffusion module performs noise reduction on the eighth feature vector from the non-textual modality after concatenating with third noise, obtaining the eleventh feature vector. Based on the eighth and tenth feature vectors from the non-textual modality, a fifth loss is obtained, and the model parameters of the seventh encoding module are adjusted based on this fifth loss. Additionally, based on the eleventh and twelfth feature vectors from the non-textual modality, a sixth loss is obtained, and the model parameters of the second diffusion module are adjusted based on this sixth loss.
[0152] It should also be noted that, when the third input information includes both text-based and non-text-based input information, the sixth encoding module maps the third input information into the first space to obtain the eighth feature vector corresponding to the text-based mode. Similarly, the seventh encoding module maps the third input information into the first space to obtain the eighth feature vector corresponding to the non-text-based mode.
[0153] Accordingly, the eighth feature vector corresponding to the text modality is analyzed using the sixth decoding module to obtain the ninth feature vector corresponding to the text modality. The eighth feature vector corresponding to the non-text modality is analyzed using the seventh decoding module to obtain the ninth feature vector corresponding to the non-text modality. Based on the eighth and ninth feature vectors in the text modality and the eighth and ninth feature vectors in the non-text modality, the fifth loss is calculated, and the model parameters of the sixth encoding module, sixth decoding module, seventh encoding module, and seventh decoding module are adjusted based on the fourth loss.
[0154] Furthermore, based on the first processing module, the input information in the text mode and the input information in the non-text mode are respectively mapped to the first space to obtain the tenth feature vector corresponding to the text mode and the tenth feature vector corresponding to the non-text mode.
[0155] Based on the eighth and tenth feature vectors in the text modality and the eighth and tenth feature vectors in the non-text modality, a fifth loss is obtained, and the model parameters of the sixth and seventh encoding modules are adjusted based on the fifth loss. Correspondingly, based on the eleventh feature vector in the text modality, the eleventh feature vector in the non-text modality, and the twelfth feature vector, a sixth loss is obtained, and the model parameters of the second diffusion module are adjusted based on the sixth loss.
[0156] S1330, In response to the convergence condition of the third machine learning model, the fifth encoding module is used as the first encoding module, the second diffusion module is used as the first diffusion module, and the fifth decoding module is used as the first decoding module to obtain the first machine learning model.
[0157] Specifically, when the third machine learning model reaches the convergence condition, a first encoding module is obtained based on the adjusted parameters of the sixth and seventh encoding modules, and a second diffusion module is used as the first diffusion module. A first decoding module is also obtained based on the adjusted parameters of the sixth and seventh decoding modules. A first machine learning model is then obtained based on the first encoding module, the first decoding module, and the first diffusion module.
[0158] The beneficial effects of the above method are as follows: In response to the correlation between the third input information and the text modality, the second model parameters of the sixth encoding module, the sixth decoding module, and the second diffusion module are adjusted based on the third input information and the third machine learning model; in response to the correlation between the third input information and the non-text modality, the second model parameters of the seventh encoding module, the seventh decoding module, and the second diffusion module are adjusted based on the third input information and the third machine learning model; when the third machine learning model reaches the convergence condition, the fifth encoding module is used as the first encoding module, the second diffusion module as the first diffusion module, and the fifth decoding module as the first decoding module to obtain the first machine learning model. By adjusting the parameters of the encoding module, decoding module, and second diffusion module under different modalities based on the third input information under different modalities, the model parameter adjustment becomes more targeted and flexible, which is beneficial to improving the data processing accuracy of the encoding module, decoding module, and second diffusion module under the corresponding modality.
[0159] Figure 14 This is a flowchart illustrating the model parameter adjustment process for the third machine learning model provided in this paper. Building upon the above, to improve training efficiency, the process of adjusting only the second model parameters of the second diffusion module can also be described. For a detailed description of this second module adjustment process, please refer to the article. Technical features identical or similar to those described above will not be repeated here.
[0160] S1410, In response to the third input information being related to a text modality or a non-text modality, adjust the second model parameters of the second diffusion module based on the third input information and the third machine learning model.
[0161] Specifically, in order to improve model training efficiency and obtain the first machine learning model more quickly, the model parameters of the fifth encoding module and the fifth decoding module of the third machine learning model can be adjusted without adjusting them. Instead, the model parameters of the second diffusion module of the third machine learning model can be adjusted based on the third input information.
[0162] This can be understood as follows: Based on the modality of the third input information, at least one encoding module in the fifth encoding module that is compatible with the third input information is obtained. Then, based on at least one encoding module in the fifth encoding module, latent space mapping processing is performed on the third input information to obtain the eighth feature vector corresponding to the third input information. Furthermore, based on the second diffusion module, the eighth feature vector after concatenating the third noise is denoised to obtain the eleventh feature vector. Based on the eleventh and twelfth feature vectors, a sixth loss is obtained, and the model parameters of the second diffusion module are adjusted based on the sixth loss.
[0163] S1420. In response to the convergence condition of the third machine learning model, the fifth encoding module is used as the first encoding module, the second diffusion module is used as the first diffusion module, and the fifth decoding module is used as the first decoding module to obtain the first machine learning model.
[0164] Specifically, when the third machine learning model reaches the convergence condition, the fifth encoding module is used as the first encoding module, the fifth decoding module is used as the first decoding module, and the second diffusion module with adjusted parameters is used as the first diffusion module. Based on the first encoding module, the first decoding module, and the first diffusion module, the first machine learning model is obtained.
[0165] The beneficial effects of the above method are as follows: Responding to the correlation between the third input information and the textual or non-textual modality, the second model parameters of the second diffusion module are adjusted based on the third input information and the third machine learning model; when the third machine learning model reaches the convergence condition, the first machine learning model is obtained based on the fifth encoding module, the second diffusion module, and the fifth decoding module. By adjusting only the model parameters of the second diffusion module, the amount of training data can be reduced, computational overhead can be lowered, the model training efficiency of the third machine learning model can be improved, and the model training speed of the first machine learning model can be guaranteed, thus enabling the first machine learning model to be applied more quickly.
[0166] Figure 15 This is a schematic diagram of the structure of a content generation device in one scenario, such as... Figure 15 As shown, the device includes: an information receiving module 1510, an information processing module 1520, and a content output module 1530.
[0167] The information receiving module 1510 is used to acquire first input information, which includes at least text information; the information processing module 1520 is used to obtain a first feature vector based on a first encoding module, to perform noise reduction processing on a second feature vector based on a first diffusion module to obtain a third feature vector, and to parse the third feature vector based on a first decoding module. The first feature vector is used to indicate the feature vector of the first input information in a first space, and the second feature vector is used to indicate the vector after concatenating the first feature vector with first noise. The first encoding module, the first diffusion module, and the first decoding module are modules in a first machine learning model; the content output module 1530 is used to output first content, which reflects the parsing result of the third feature vector.
[0168] In one scenario, optionally, the information processing module includes: a third feature vector acquisition unit, configured to perform noise reduction processing on the second feature vector based on the first diffusion module to obtain a fourth feature vector; to concatenate second noise into the fourth feature vector to obtain a fifth feature vector; and to process the fifth feature vector based on the first diffusion module to obtain a third feature vector.
[0169] In another scenario, optionally, the first encoding module includes a second encoding module and a third encoding module, and the first decoding module includes a second decoding module and a third decoding module; the second encoding module is associated with the second decoding module and is used to process the first input information in a text modality; the third encoding module is associated with the third decoding module and is used to process the first input information in a non-text modality.
[0170] In another scenario, optionally, the above-mentioned device further includes: a first model parameter acquisition module, comprising: a second input information acquisition unit, configured to acquire second input information, the second input information including first information and second information, the second information being used to indicate information after at least a portion of the content in the first information is masked; an information encoding unit, configured to process the second input information based on the fourth encoding module in the second machine learning model to obtain a sixth feature vector and a seventh feature vector of the second input information in a first space, the sixth feature vector being related to the first information and the seventh feature vector being related to the second information; an information decoding unit, configured to process the sixth feature vector and the seventh feature vector based on the fourth decoding module in the second machine learning model to obtain second content, the second content including third information and fourth information, the third information being used to characterize the reconstructed content of the sixth feature vector and the fourth information being used to characterize the predicted content of the seventh feature vector; and a first model parameter acquisition unit, configured to adjust the parameters of the fourth encoding module and the fourth decoding module based on the first information and the second content, and to use the model parameters of the fourth encoding module and the fourth decoding module when the convergence condition is met as the first model parameters.
[0171] In another scenario, optionally, the first model parameter acquisition unit is configured to obtain a first loss based on the first information and the third information, the first loss reflecting the reconstruction loss of the second machine learning model; obtain a second loss based on the first information and the fourth information, the second loss reflecting the prediction loss of the second machine learning model; obtain a third loss based on the first information, the second information, and the third information; and adjust the parameters of the fourth encoding module and the fourth decoding module based on the first loss, the second loss, and the third loss to obtain the first model parameters.
[0172] Optionally, in another scenario, the device further includes: a third machine learning model adjustment module, comprising: a second model parameter adjustment unit, configured to adjust the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module based on the third machine learning model and third input information, wherein the third machine learning model includes a second diffusion module, a first processing module, a fifth encoding module, and a fifth decoding module, wherein the fifth encoding module is the same as the fourth encoding module, the fifth decoding module is the same as the fourth decoding module, the first processing module includes the fourth encoding module, and is configured to instruct the adjustment of the second model parameters of the second diffusion module based on the output result of the second diffusion module; and a first machine learning model acquisition unit, configured to, in response to the third machine learning model reaching a convergence condition, use the fifth encoding module as the first encoding module, the second diffusion module as the first diffusion module, and the fifth decoding module as the first decoding module to obtain the first machine learning model.
[0173] In another scenario, optionally, the second model parameter adjustment unit includes: a ninth feature vector acquisition subunit, used to map the third input information to the first space based on the fifth encoding module to obtain an eighth feature vector, and to parse the eighth feature vector based on the fifth decoding module to obtain a ninth feature vector; a tenth feature vector acquisition subunit, used to map the third input information to the first space based on the first processing module to obtain a tenth feature vector; an eleventh feature vector acquisition subunit, used to perform noise reduction processing on the eighth feature vector after splicing third noise based on the second diffusion module to obtain an eleventh feature vector; and a second model parameter adjustment subunit, used to adjust the second model parameters based on the eighth feature vector, the ninth feature vector, the tenth feature vector, and the eleventh feature vector; the second model parameters are used to reflect the model parameters of the third machine learning model.
[0174] In another scenario, optionally, the second model parameter adjustment subunit is used to obtain a fourth loss based on the eighth and ninth feature vectors; a fifth loss based on the eighth and tenth feature vectors; and a sixth loss based on the eleventh and twelfth feature vectors, wherein the twelfth feature vector is used to indicate the feature vector corresponding to the desired content, which is related to the third input information; adjust the second model parameters of the fifth encoding module and the fifth decoding module based on the fourth loss; adjust the second model parameters of the fifth encoding module based on the fifth loss; and adjust the second model parameters of the second diffusion module based on the sixth loss.
[0175] In another scenario, optionally, the eleventh feature vector acquisition subunit is used to perform noise reduction processing on the eighth feature vector after splicing the third noise based on the second diffusion module to obtain a thirteenth feature vector, wherein the thirteenth feature vector is used to indicate the processing result of obtaining at least one block after noise reduction processing; the first processing unit performs loss processing on the thirteenth feature vector and the twelfth feature vector to obtain a seventh loss, and adjusts the parameters of the second diffusion module based on the seventh loss; the second diffusion module with adjusted parameters performs noise reduction processing on the thirteenth feature vector after splicing the fourth noise to obtain a fourteenth feature vector; in response to the fourteenth feature vector satisfying a first condition, the fourteenth feature vector is used as the eleventh feature vector.
[0176] In another scenario, optionally, the fifth encoding module includes a sixth encoding module and a seventh encoding module, and the fifth decoding module includes a sixth decoding module and a seventh decoding module. The sixth encoding module is associated with the sixth decoding module and is used to process the third input information of the text modality. The seventh encoding module is associated with the seventh decoding module and is used to process the third input information of the non-text modality. The second model parameter adjustment subunit is used to adjust the second model parameters of the sixth encoding module, the sixth decoding module, and the second diffusion module based on the third input information and the third machine learning model in response to the correlation between the third input information and the text modality; or, in response to the correlation between the third input information and the non-text modality, adjust the second model parameters of the seventh encoding module, the seventh decoding module, and the second diffusion module based on the third input information and the third machine learning model.
[0177] In another scenario, optionally, the fifth encoding module includes a sixth encoding module and a seventh encoding module, and the fifth decoding module includes a sixth decoding module and a seventh decoding module. The sixth encoding module is associated with the sixth decoding module and is used to process the third input information of the text modality. The seventh encoding module is associated with the seventh decoding module and is used to process the third input information of the non-text modality. The second model parameter adjustment subunit is used to adjust the second model parameters of the second diffusion module based on the third input information and the third machine learning model, in response to the third input information being associated with the text modality or the non-text modality.
[0178] The beneficial effects of the above-mentioned device are as follows: It acquires first input information, which includes at least text information; obtains a first feature vector based on a first encoding module; obtains a third feature vector by denoising a second feature vector based on a first diffusion module; and parses the third feature vector based on a first decoding module. The first feature vector indicates the feature vector of the first input information in a first space, and the second feature vector indicates the vector obtained by concatenating the first feature vector with first noise. The first encoding module, the first diffusion module, and the first decoding module are modules in a first machine learning model. It outputs first content, which reflects the parsing result of the third feature vector. Therefore, it can accurately and quickly process the first input information based on the first machine learning model, solving the problem that existing machine learning models can only process individual data units one by one, i.e., only one data unit's corresponding content can be generated each time, resulting in low content generation efficiency. This achieves the technical effect of improving content generation efficiency.
[0179] The content generation apparatus provided herein can execute any of the content generation methods provided herein, and has the corresponding functional modules and beneficial effects of executing the methods.
[0180] It is worth noting that the various units and modules included in the above-mentioned device are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this document.
[0181] The following is for reference. Figure 16 This document illustrates a schematic diagram of an electronic device (e.g., a terminal device or server) 1600 suitable for implementing the above-described methods. The terminal device referred to herein may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 16 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments described herein.
[0182] like Figure 16As shown, electronic device 1600 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1602 or a program loaded from storage device 1608 into random access memory (RAM) 1603. The RAM 1603 also stores various programs and data required for the operation of electronic device 1600. The processing unit 1601, ROM 1602, and RAM 1603 are interconnected via bus 1604. An input / output (I / O) interface 1605 is also connected to bus 1604.
[0183] Typically, the following devices can be connected to the I / O interface 1605: input devices 1606 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1607 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1608 including, for example, magnetic tape, hard disk, etc.; and communication devices 1609. Communication device 1609 allows electronic device 1600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 16 An electronic device 1600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0184] In particular, according to embodiments of this document, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, the technical solutions of this document include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via communication device 1609, or installed from storage device 1608, or installed from ROM 1602. When the computer program is executed by processing device 1601, it performs the functions defined in the methods of the embodiments of this document.
[0185] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0186] The electronic device provided in this embodiment and the content generation method provided in the above technical solutions belong to the same inventive concept. Technical details not described in detail in this document can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0187] This article provides a computer storage medium on which a computer program is stored, which, when executed by a processor, implements the content generation method provided in the above embodiments.
[0188] It should be noted that the computer-readable medium mentioned above can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, also known as flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0189] Based on one or more scenarios described herein, a content generation method is provided, including: Obtain first input information, wherein the first input information includes at least text information; The first feature vector is obtained based on the first encoding module, the second feature vector is denoised based on the first diffusion module to obtain the third feature vector, and the third feature vector is parsed based on the first decoding module. The first feature vector is used to indicate the feature vector of the first input information in the first space, and the second feature vector is used to indicate the vector after the first feature vector is concatenated with the first noise. The first encoding module, the first diffusion module and the first decoding module are modules in the first machine learning model. Output the first content, which reflects the parsing result of the third feature vector.
[0190] Based on one or more scenarios described in this paper, Example 2 provides a content generation method, wherein the method of obtaining a third feature vector by denoising a second feature vector based on a first diffusion module includes: Based on the first diffusion module, the second feature vector is denoised to obtain the fourth feature vector; The fourth feature vector is concatenated with the second noise to obtain the fifth feature vector; The third feature vector is obtained by processing the fifth feature vector based on the first diffusion module.
[0191] According to one or more scenarios described herein, Example 3 provides a content generation method, wherein the first encoding module includes a second encoding module and a third encoding module, and the first decoding module includes a second decoding module and a third decoding module; The second encoding module is associated with the second decoding module and is used to process the first input information in the text modality; the third encoding module is associated with the third decoding module and is used to process the first input information in the non-text modality.
[0192] Based on one or more scenarios described herein, Example 4 provides a content generation method that obtains the first model parameters of the first encoding module and the first decoding module in the following manner: Obtain second input information, which includes first information and second information, wherein the second information is used to indicate information after at least a portion of the first information is masked; The second input information is processed by the fourth encoding module in the second machine learning model to obtain the sixth feature vector and the seventh feature vector of the second input information in the first space. The sixth feature vector is related to the first information and the seventh feature vector is related to the second information. The sixth and seventh feature vectors are processed by the fourth decoding module in the second machine learning model to obtain second content. The second content includes third information and fourth information. The third information is used to characterize the reconstructed content of the sixth feature vector, and the fourth information is used to characterize the predicted content of the seventh feature vector. Based on the first information and the second content, the parameters of the fourth encoding module and the fourth decoding module are adjusted, and the model parameters of the fourth encoding module and the fourth decoding module when the convergence condition is met are used as the first model parameters.
[0193] Based on one or more scenarios described herein, Example 5 provides a content generation method, wherein adjusting the parameters of the fourth encoding module and the fourth decoding module based on the first information and the second content includes: Based on the first information and the third information, a first loss is obtained, which is used to reflect the reconstruction loss of the second machine learning model; Based on the first information and the fourth information, a second loss is obtained, which is used to reflect the prediction loss of the second machine learning model; Based on the first information, the second information, and the third information, a third loss is obtained; Based on the first loss, the second loss, and the third loss, the parameters of the fourth encoding module and the fourth decoding module are adjusted to obtain the first model parameters.
[0194] Based on one or more scenarios described herein, Example 6 provides a content generation method that trains a third machine learning model in the following manner to obtain a first machine learning model based on the trained third machine learning model: Based on the third machine learning model and the third input information, the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module are adjusted. The third machine learning model includes the second diffusion module, the first processing module, the fifth encoding module, and the fifth decoding module. The fifth encoding module is the same as the fourth encoding module, and the fifth decoding module is the same as the fourth decoding module. The first processing module includes the fourth encoding module and is used to indicate the adjustment of the second model parameters of the second diffusion module based on the output result of the second diffusion module. In response to the convergence condition of the third machine learning model, the fifth encoding module is used as the first encoding module, the second diffusion module is used as the first diffusion module, and the fifth decoding module is used as the first decoding module to obtain the first machine learning model.
[0195] Based on one or more scenarios described herein, Example 7 provides a content generation method, wherein the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module are adjusted based on the third machine learning model and the third input information, including: The third input information is mapped to the first space based on the fifth encoding module to obtain the eighth feature vector. The eighth feature vector is then parsed based on the fifth decoding module to obtain the ninth feature vector. The first processing module maps the third input information to the first space to obtain the tenth feature vector; Based on the second diffusion module, the eighth feature vector after splicing the third noise is denoised to obtain the eleventh feature vector. The parameters of the second model are adjusted based on the eighth feature vector, the ninth feature vector, the tenth feature vector, and the eleventh feature vector.
[0196] Based on one or more scenarios described herein, Example 8 provides a content generation method, wherein adjusting the second model parameters based on the eighth feature vector, the ninth feature vector, the tenth feature vector, and the eleventh feature vector includes: Based on the eighth feature vector and the ninth feature vector, a fourth loss is obtained; Based on the eighth feature vector and the tenth feature vector, the fifth loss is obtained; Based on the eleventh and twelfth feature vectors, a sixth loss is obtained, wherein the twelfth feature vector is used to indicate the feature vector corresponding to the expected content, and the expected content is related to the third input information; The second model parameters of the fifth encoding module and the fifth decoding module are adjusted based on the fourth loss, the second model parameters of the fifth encoding module are adjusted based on the fifth loss, and the second model parameters of the second diffusion module are adjusted based on the sixth loss.
[0197] Based on one or more scenarios described in this paper, Example 9 provides a content generation method in which the eighth feature vector after splicing third noise is denoised based on the second diffusion module to obtain the eleventh feature vector, including: Based on the second diffusion module, the eighth feature vector after splicing the third noise is denoised to obtain the thirteenth feature vector, which is used to indicate the processing result of at least one block after the denoising process. The first processing unit performs loss processing on the thirteenth and twelfth feature vectors to obtain the seventh loss, and adjusts the parameters of the second diffusion module based on the seventh loss. Based on the parameter-adjusted second diffusion module, the thirteenth feature vector of the spliced fourth noise is denoised to obtain the fourteenth feature vector; In response to the fourteenth feature vector satisfying the first condition, the fourteenth feature vector is used as the eleventh feature vector.
[0198] According to one or more scenarios in this paper, Example 10 provides a content generation method, wherein the fifth encoding module includes a sixth encoding module and a seventh encoding module, the fifth decoding module includes a sixth decoding module and a seventh decoding module, the sixth encoding module is related to the sixth decoding module and is used to process the third input information of the text modality, and the seventh encoding module is related to the seventh decoding module and is used to process the third input information of the non-text modality; The adjustment of the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module based on the third machine learning model and the third input information includes one of the following: In response to the correlation between the third input information and the text modality, the second model parameters of the sixth encoding module, the sixth decoding module, and the second diffusion module are adjusted based on the third input information and the third machine learning model; or, In response to the correlation between the third input information and the non-text modality, the second model parameters of the seventh encoding module, the seventh decoding module, and the second diffusion module are adjusted based on the third input information and the third machine learning model.
[0199] Based on one or more scenarios described herein, Example 11 provides a content generation method, wherein the fifth encoding module includes a sixth encoding module and a seventh encoding module, and the fifth decoding module includes a sixth decoding module and a seventh decoding module. The sixth encoding module is associated with the sixth decoding module and is used to process the third input information of the text modality, and the seventh encoding module is associated with the seventh decoding module and is used to process the third input information of the non-text modality. The adjustment of the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module based on the third machine learning model and the third input information includes: In response to the third input information being related to the text modality or non-text modality, the second model parameters of the second diffusion module are adjusted based on the third input information and the third machine learning model.
[0200] According to one or more scenarios described herein, Example Twelve provides a content generation apparatus, comprising: The information receiving module is used to acquire first input information, wherein the first input information includes at least text information; The information processing module is used to obtain a first feature vector based on the first encoding module, to perform noise reduction processing on the second feature vector based on the first diffusion module to obtain a third feature vector, and to parse the third feature vector based on the first decoding module. The first feature vector is used to indicate the feature vector of the first input information in the first space, and the second feature vector is used to indicate the vector after concatenating the first feature vector with the first noise. The first encoding module, the first diffusion module, and the first decoding module are modules in the first machine learning model. The content output module is used to output first content, which reflects the parsing result of the third feature vector.
[0201] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0202] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0203] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: Obtain first input information, wherein the first input information includes at least text information; The first feature vector is obtained based on the first encoding module, the second feature vector is denoised based on the first diffusion module to obtain the third feature vector, and the third feature vector is parsed based on the first decoding module. The first feature vector is used to indicate the feature vector of the first input information in the first space, and the second feature vector is used to indicate the vector after the first feature vector is concatenated with the first noise. The first encoding module, the first diffusion module and the first decoding module are modules in the first machine learning model. Output the first content, which reflects the parsing result of the third feature vector.
[0204] Computer program code for performing the operations described herein can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0205] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this document. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0206] The modules or units described herein can be implemented in software or hardware. The names of modules or units do not necessarily limit the functionality of the module or unit itself; for example, a second result acquisition unit can also be described as a "second result receiving unit".
[0207] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include at least one of the following: Field-Programmable Gate Array (FPGA), Application-Specific Integrated Circuit (ASIC), Application-Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), etc.
[0208] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (flash memory), optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0209] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.
[0210] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this document. Certain features described in the context of individual implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.
[0211] Although the subject matter has been described using a programming language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims.
Claims
1. A content generation method, comprising: Obtain first input information, wherein the first input information includes at least text information; The first feature vector is obtained based on the first encoding module, the second feature vector is denoised based on the first diffusion module to obtain the third feature vector, and the third feature vector is parsed based on the first decoding module. The first feature vector is used to indicate the feature vector of the first input information in the first space, and the second feature vector is used to indicate the vector after the first feature vector is concatenated with the first noise. The first encoding module, the first diffusion module and the first decoding module are modules in the first machine learning model. Output the first content, which reflects the parsing result of the third feature vector.
2. The content generation method according to claim 1, wherein obtaining the third feature vector by performing noise reduction processing on the second feature vector based on the first diffusion module includes: Based on the first diffusion module, the second feature vector is denoised to obtain the fourth feature vector; The fourth feature vector is concatenated with the second noise to obtain the fifth feature vector; The third feature vector is obtained by processing the fifth feature vector based on the first diffusion module.
3. The content generation method according to claim 1, wherein the first encoding module includes a second encoding module and a third encoding module, and the first decoding module includes a second decoding module and a third decoding module; The second encoding module is associated with the second decoding module and is used to process the first input information in the text modality; the third encoding module is associated with the third decoding module and is used to process the first input information in the non-text modality.
4. The content generation method according to claim 1, wherein the first model parameters of the first encoding module and the first decoding module are obtained in the following manner: Obtain second input information, which includes first information and second information, wherein the second information is used to indicate information after at least a portion of the content in the first information is masked; The second input information is processed by the fourth encoding module in the second machine learning model to obtain the sixth feature vector and the seventh feature vector of the second input information in the first space. The sixth feature vector is related to the first information and the seventh feature vector is related to the second information. The sixth and seventh feature vectors are processed by the fourth decoding module in the second machine learning model to obtain second content. The second content includes third information and fourth information. The third information is used to characterize the reconstructed content of the sixth feature vector, and the fourth information is used to characterize the predicted content of the seventh feature vector. Based on the first information and the second content, the parameters of the fourth encoding module and the fourth decoding module are adjusted, and the model parameters of the fourth encoding module and the fourth decoding module when the convergence condition is met are used as the first model parameters.
5. The content generation method according to claim 4, wherein adjusting the parameters of the fourth encoding module and the fourth decoding module based on the first information and the second content includes: Based on the first information and the third information, a first loss is obtained, which is used to reflect the reconstruction loss of the second machine learning model; Based on the first information and the fourth information, a second loss is obtained, which is used to reflect the prediction loss of the second machine learning model; Based on the first information, the second information, and the third information, a third loss is obtained; Based on the first loss, the second loss, and the third loss, the parameters of the fourth encoding module and the fourth decoding module are adjusted to obtain the first model parameters.
6. The content generation method according to claim 1, wherein the third machine learning model is trained in the following manner to obtain the first machine learning model based on the trained third machine learning model: Based on the third machine learning model and the third input information, the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module are adjusted. The third machine learning model includes the second diffusion module, the first processing module, the fifth encoding module, and the fifth decoding module. The fifth encoding module is the same as the fourth encoding module, and the fifth decoding module is the same as the fourth decoding module. The first processing module includes the fourth encoding module and is used to indicate the adjustment of the second model parameters of the second diffusion module based on the output result of the second diffusion module. In response to the convergence condition of the third machine learning model, the fifth encoding module is used as the first encoding module, the second diffusion module is used as the first diffusion module, and the fifth decoding module is used as the first decoding module to obtain the first machine learning model.
7. The content generation method according to claim 6, wherein adjusting the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module based on the third machine learning model and the third input information includes: The third input information is mapped to the first space based on the fifth encoding module to obtain the eighth feature vector. The eighth feature vector is then parsed based on the fifth decoding module to obtain the ninth feature vector. The first processing module maps the third input information to the first space to obtain the tenth feature vector; Based on the second diffusion module, the eighth feature vector after splicing the third noise is denoised to obtain the eleventh feature vector. The parameters of the second model are adjusted based on the eighth feature vector, the ninth feature vector, the tenth feature vector, and the eleventh feature vector.
8. The content generation method according to claim 7, wherein adjusting the second model parameters based on the eighth feature vector, the ninth feature vector, the tenth feature vector, and the eleventh feature vector includes: Based on the eighth feature vector and the ninth feature vector, a fourth loss is obtained; Based on the eighth feature vector and the tenth feature vector, the fifth loss is obtained; Based on the eleventh and twelfth feature vectors, a sixth loss is obtained, wherein the twelfth feature vector is used to indicate the feature vector corresponding to the expected content, and the expected content is related to the third input information; The second model parameters of the fifth encoding module and the fifth decoding module are adjusted based on the fourth loss, the second model parameters of the fifth encoding module are adjusted based on the fifth loss, and the second model parameters of the second diffusion module are adjusted based on the sixth loss.
9. The content generation method according to claim 7, wherein the step of performing noise reduction processing on the eighth feature vector after splicing the third noise based on the second diffusion module to obtain the eleventh feature vector includes: Based on the second diffusion module, the eighth feature vector after splicing the third noise is denoised to obtain the thirteenth feature vector, which is used to indicate the processing result of at least one block after the denoising process. The first processing unit performs loss processing on the thirteenth and twelfth feature vectors to obtain the seventh loss, and adjusts the parameters of the second diffusion module based on the seventh loss. Based on the parameter-adjusted second diffusion module, the thirteenth feature vector of the spliced fourth noise is denoised to obtain the fourteenth feature vector; In response to the fourteenth feature vector satisfying the first condition, the fourteenth feature vector is used as the eleventh feature vector.
10. The content generation method according to claim 6, wherein the fifth encoding module includes a sixth encoding module and a seventh encoding module, the fifth decoding module includes a sixth decoding module and a seventh decoding module, the sixth encoding module is related to the sixth decoding module and is used to process the third input information of the text modality, and the seventh encoding module is related to the seventh decoding module and is used to process the third input information of the non-text modality; The adjustment of the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module based on the third machine learning model and the third input information includes one of the following: In response to the correlation between the third input information and the text modality, the second model parameters of the sixth encoding module, the sixth decoding module, and the second diffusion module are adjusted based on the third input information and the third machine learning model; or, In response to the correlation between the third input information and the non-text modality, the second model parameters of the seventh encoding module, the seventh decoding module, and the second diffusion module are adjusted based on the third input information and the third machine learning model.
11. The content generation method according to claim 6, wherein the fifth encoding module includes a sixth encoding module and a seventh encoding module, and the fifth decoding module includes a sixth decoding module and a seventh decoding module; the sixth encoding module is associated with the sixth decoding module and is used to process the third input information of the text modality; and the seventh encoding module is associated with the seventh decoding module and is used to process the third input information of the non-text modality. The adjustment of the second model parameters of the fifth encoding module, the fifth decoding module, and the second diffusion module based on the third machine learning model and the third input information includes: In response to the third input information being related to the text modality or non-text modality, the second model parameters of the second diffusion module are adjusted based on the third input information and the third machine learning model.
12. A content generation apparatus, comprising: The information receiving module is used to acquire first input information, wherein the first input information includes at least text information; The information processing module is used to obtain a first feature vector based on the first encoding module, to perform noise reduction processing on the second feature vector based on the first diffusion module to obtain a third feature vector, and to parse the third feature vector based on the first decoding module. The first feature vector is used to indicate the feature vector of the first input information in the first space, and the second feature vector is used to indicate the vector after concatenating the first feature vector with the first noise. The first encoding module, the first diffusion module, and the first decoding module are modules in the first machine learning model. The content output module is used to output first content, which reflects the parsing result of the third feature vector.
13. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the content generation method as described in any one of claims 1-11.
14. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the content generation method as described in any one of claims 1-11.
15. A computer program product comprising a computer program that, when executed by a processor, implements the content generation method as described in any one of claims 1-11.