Target model training method, multimodal data processing method and apparatus therefor, electronic device and program
By training a target model with multimodal data using an N-layer spread network, the method addresses the challenge of incomplete image generation from complex descriptions, enhancing the model's controllability and image quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2024-11-21
- Publication Date
- 2026-04-27
AI Technical Summary
Existing text-to-image diffusion models struggle to fully capture the details and essence of complex descriptions due to the inherent brevity of text, leading to incomplete image generation.
A method for training a target model using multimodal data, including image and text features, through an N-layer spread network, enhancing the model's ability to process richer descriptive information and improve controllability.
The solution enriches the model's usage scenarios, allowing for precise control over image generation, meeting user customization needs and improving the quality of generated images.
Smart Images

Figure 0007852017000001 
Figure 0007852017000002 
Figure 0007852017000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to the fields of technologies such as computer vision, deep learning, and large-scale models, and can be applied to scenarios such as Artificial Intelligence Generated Content (AIGC).
Background Art
[0002] Diffusion models from text to image (Text to Image, T2I) have made remarkable progress in recent years, and this model technology has shown remarkable capabilities in generating highly faithful images from text descriptions. With the continuous innovation of deep learning and artificial intelligence technologies, existing model technologies can convert simple text descriptions into detailed and realistic images. This conversion not only provides a powerful tool for creative workers but also brings a never-before-seen visual experience to users.
Summary of the Invention
Problems to be Solved by the Invention
[0003] The present disclosure provides a method for training a target model, a method and apparatus for multi-modal data processing, an electronic device, and a program.
Means for Solving the Problems
[0004] In one aspect of the present disclosure, a method for training a target model is provided, and the method includes In another aspect of the present disclosure, a method for training a target model is provided, and the method includes The method involves inputting sample data into a preset model to obtain initial multimodal features of the sample data, wherein the sample data includes at least sample text and at least one entity image, the sample text includes at least one entity text, the entity image is an image corresponding to the entity text, the initial multimodal features include image features of the entity image and text features of the entity text, and the image features of the entity image are after the text features of the entity text corresponding to the entity image. If the parameters of the image / text encoder of the preset model are fixed, the method involves training the N-layer spread network in the preset model using the initial multimodal features and preset noise features to obtain a target model, where N is a positive integer.
[0005] Another aspect of this disclosure provides a multimodal data processing method, which includes: The method involves inputting multimodal prompt information into a pre-trained target model to obtain target multimodal features of the multimodal prompt information, wherein the multimodal prompt information includes at least a target text and at least one target entity image, the target text includes at least one target entity text, the target entity image is an image corresponding to the target entity text, and the target multimodal features include image features of the target entity image and text features of the target entity text, with the image features of the target entity image following the text features of the target entity text corresponding to the target entity image. The method involves using an N-layer goal diffusion network in the aforementioned goal model to perform feature diffusion on the goal multimodal features to obtain a goal inference result, including images and / or text, related to the multimodal prompt information, wherein N is a positive integer.
[0006] Another aspect of this disclosure provides a training device for a target model, the device being: A training unit is used to input sample data into a preset model to obtain initial multimodal features of the sample data, wherein the sample data includes at least sample text and at least one entity image, the sample text includes at least one entity text, the entity image is an image corresponding to the entity text, the initial multimodal features include image features of the entity image and text features of the entity text, the image features of the entity image are after the text features of the entity text corresponding to the entity image, and, if the parameters of the image / text encoder of the preset model are fixed, to perform model training on an N-layer spread network in the preset model using the initial multimodal features and preset noise features to obtain a target model, where N is a positive integer. The system includes a storage unit for outputting the aforementioned target model.
[0007] In another aspect of this disclosure, a multimodal data processing device is provided, the device is A reasoning unit used in the following: inputting multimodal prompt information into a pre-trained target model to obtain target multimodal features of the multimodal prompt information, wherein the multimodal prompt information includes at least a target text and at least one target entity image, the target text includes at least one target entity text, the target entity image is an image corresponding to the target entity text, the target multimodal features include image features of the target entity image and text features of the target entity text, the image features of the target entity image are after the text features of the target entity text corresponding to the target entity image, and performing feature diffusion on the target multimodal features using an N-layer target diffusion network in the target model to obtain a target reasoning result including an image and / or text related to the multimodal prompt information, where N is a positive integer; The system includes an output unit for outputting the aforementioned target reasoning results.
[0008] In another aspect of this disclosure, an electronic device is provided, which is At least one processor, The system comprises at least one processor and memory that is communicated with, The memory stores instructions that are executable by the at least one processor, and when executed by the at least one processor, the instructions cause one of the methods in the embodiments of the present disclosure to be performed.
[0009] Another aspect of the present disclosure provides a non-temporary computer-readable storage medium that stores computer instructions for causing a computer to perform any one of the methods of the embodiments of the present disclosure.
[0010] In another aspect of the present disclosure, a program is provided, when executed by a processor, for performing any of the methods in the embodiments of the present disclosure.
[0011] Thus, the solution of this disclosure utilizes the generated initial multimodal features and preset noise features to train an N-layer spread spectrum network in a preset model. This allows the trained target model to process multimodal information, enriching the model's usage scenarios, effectively improving its ability to grasp and control details, and enhancing the controllability of image generation.
[0012] It should be understood that the information contained herein is not intended to describe any key points or important features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Further details of other features of this disclosure will be provided in the specification below.
[0013] The attached drawings are for the purpose of better understanding the solutions of this disclosure and do not constitute a limitation of this disclosure. [Brief explanation of the drawing]
[0014] [Figure 1] It is one of the schematic flowcharts of the model training method according to an embodiment of the present disclosure. [Figure 2] It is one of the schematic diagrams showing the initial multimodal features according to an embodiment of the present disclosure. [Figure 3(a)] It is a schematic diagram showing the configuration of the diffusion network according to an embodiment of the present disclosure. [Figure 3(b)] It is a schematic diagram showing the series connection of the diffusion network according to an embodiment of the present disclosure. [Figure 4] It is the second of the schematic flowcharts of the model training method according to an embodiment of the present disclosure. [Figure 5(a)] It is the second of the schematic diagrams showing the initial multimodal features according to an embodiment of the present disclosure. [Figure 5(b)] It is a schematic diagram showing the instruction information according to an embodiment of the present disclosure. [Figure 6(a)] It is the third of the schematic flowcharts of the model training method according to an embodiment of the present disclosure. [Figure 6(b)] It is the third of the schematic diagrams showing the initial multimodal features according to an embodiment of the present disclosure. [Figure 6(c)] It is a schematic diagram showing the scenario in a specific example of the model training method according to an embodiment of the present disclosure. [Figure 7(a)] It is a schematic diagram showing the configuration of the dual information stream diffusion network according to another embodiment of the present disclosure. [Figure 7(b)] It is a schematic diagram showing the series connection of the dual information stream diffusion network according to another embodiment of the present disclosure. [Figure 8] It is the first of the schematic flowcharts of the multimodal data processing method according to an embodiment of the present disclosure. [Figure 9] It is the second of the schematic flowcharts of the multimodal data processing method according to an embodiment of the present disclosure. [Figure 10] It is a schematic diagram showing the application flow in a specific example of the multimodal data processing method according to an embodiment of the present disclosure. [Figure 11]It is a schematic diagram showing the configuration of a model training device according to an embodiment of the present disclosure. [Figure 12] It is a schematic diagram showing the configuration of a multimodal data processing device according to an embodiment of the present disclosure. [Figure 13] It is a block diagram of an electronic device for implementing a model training method or a multimodal data processing method according to an embodiment of the present disclosure.
Embodiments for Carrying out the Invention
[0015] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the accompanying drawings. These drawings include various details of the embodiments of the present disclosure for the purpose of assisting understanding, and these should be considered merely exemplary. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, descriptions of well-known features and structures are omitted in the following description for the sake of clarity and brevity.
[0016] The term "and / or" in the present disclosure merely describes the relationship between related objects and indicates that three types of relationships may exist. For example, A and / or B refers to three situations where A exists alone, A and B exist simultaneously, and B exists alone. The term "at least one" in the present disclosure indicates any one or any combination of at least two among a plurality. For example, at least one of A, B, and C indicates that any one or a plurality of elements can be selected from the set consisting of A, B, and C. The terms "first" and "second" in the present disclosure are used to refer to and distinguish a plurality of similar technical terms, and do not mean to limit the order or only limit to two. For example, the first feature and the second feature mean the existence of two types / two features, the first feature may be one or more, and the second feature may be one or more.
[0017] Furthermore, to better illustrate this disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that this disclosure is equally implementable without these specific details. In some examples, methods, means, components, and circuits that are well known to those skilled in the art are not described in detail in order to emphasize the essence of this disclosure.
[0018] The following describes related technologies of the embodiments of this disclosure, which can be optionally combined with the technical solutions of the embodiments of this disclosure as any solution, and which are all included within the scope of protection of the embodiments of this disclosure.
[0019] Text-to-image (T2I) diffusion models have made remarkable progress in recent years, demonstrating a remarkable ability to generate high-fidelity images from text descriptions. Existing T2I diffusion models can intelligently generate highly relevant images based on text prompts provided by the user. The emergence of these methods has greatly enriched visual representation options, making text-to-image conversion more direct and efficient. This conversion not only provides creative workers with a powerful tool but also delivers users an unprecedented visual experience.
[0020] While these methods have achieved remarkable results in image generation, they face a fundamental challenge: the inherent brevity of text descriptions themselves. When attempting to describe complex details, specific entities, or intricate scenarios in short characters, the limitations of language mean that the generated image may not fully capture all the details and essence of the original description.
[0021] To overcome this challenge and further improve the controllability of image generation, this disclosure provides a novel method for faithfully generating images using a generic visual language (VL) as input. Here, the generic visual language may include richer descriptive information such as product illustrations, photographs, and sketches. This information can provide the image generation model with more context and detailed guidance. This method allows for more precise control over the content and style of the generated images, meeting the user's greater customization needs.
[0022] Before detailing the solutions presented in this disclosure, we will briefly introduce some concepts that may be relevant to model training.
[0023] Regarding fine-tuning, while training a model from scratch involves identifying the model that needs training and then making adjustments based on that model, fine-tuning can save a large amount of computational resources and time, thereby improving computational efficiency and accuracy.
[0024] Regarding latent space, it is a method of representing image compression that extracts important features from an image and removes unimportant features.
[0025] Part 1 will provide a detailed explanation of model training methods.
[0026] The solution presented here proposes a model training method for multimodal information, enabling the model to process multimodal information, improving the controllability of the model in image generation, and ultimately effectively enhancing the user experience.
[0027] Furthermore, the preset model of the solution disclosed herein is based on a Multimodal Diffusion Transformer (MMDiT, or simply MultiModal-DiT) architecture. Compared to a U-net architecture model, the solution disclosed herein can support not only text-to-image generation tasks but also multimodal tasks (tasks that generate text and / or images based on input multimodal information), thus improving the applicability of the model in each scenario.
[0028] Specifically, Figure 1 is a schematic flowchart of a model training method according to one embodiment of the present invention. This method can be selectively applied to electronic devices such as personal computers, servers, and server clusters.
[0029] Furthermore, the method includes at least a portion of the following: As shown in Figure 1, it includes the following:
[0030] In step S101, sample data is input into the preset model to obtain the initial multimodal features of the sample data.
[0031] For example, in one case, the initial multimodal features of sample data can be obtained using image and text encoders in a preset model.
[0032] In this example, the sample data includes at least one sample text and at least one actual image (which can be described as M1 actual images for convenience, as will be explained later). Furthermore, the sample text includes at least one actual text (which can be described as M2 actual texts for convenience, as will be explained later). Here, the actual image is the image corresponding to the actual text.
[0033] Here, M1 and M2 are positive integers greater than or equal to 1. Furthermore, in one example, the values of M1 and M2 may be different, for example, M1 may be less than M2, which may correspond to a scenario where not all sample texts necessarily correspond to a single entity image. Alternatively, in another example, the values of M1 and M2 may be the same, which may correspond to a scenario where one entity text corresponds to a single entity image.
[0034] Here, the relationship between the entity text and the entity image may be one-to-one or one-to-many, and the solutions of this disclosure are not limited thereto. Accordingly, the solutions of this disclosure do not limit the values of M1 and M2.
[0035] Furthermore, in one example, the initial multimodal feature includes image features of an entity image and text features of an entity text, where the image features of the entity image follow the text features of the entity text corresponding to the entity image; in other words, in this example, the initial multimodal feature is obtained by inserting at least the image features of the entity image after the text features of the entity text corresponding to the entity image.
[0036] In step S102, if the parameters of the image / text encoder of the preset model are fixed, the initial multimodal features and preset noise features are used to train (e.g., fine-tune) the N-layer spread spectrum network in the preset model to obtain the target model. For example, a target model containing at least N target spread spectrum networks is obtained.
[0037] Here, the target model can infer a target reasoning result, including images and / or text, related to the multimodal prompt information, based on the input multimodal prompt information.
[0038] Furthermore, N is a positive integer greater than or equal to 1.
[0039] Furthermore, the training method of the solution disclosed herein can be applied primarily to scenarios for training a spreading network, in which case, step S101 specifically allows the input of sample data to be used in the current step into the preset model in order to obtain initial multimodal features of the sample data to be used in the current step, using the image / text encoder in the preset model. Accordingly, step S102 specifically allows the model training of the N-layer spreading network in the preset model using the initial multimodal features and the preset noise features to be used in the current step in order to complete the training in the current step.
[0040] Here, “current step” can be any of the total number of steps required to train the spreading network (or spreading model) (also known as the preset number of steps). In other words, the solution of this disclosure can be applied to any step in training a spreading network.
[0041] Furthermore, the term "step" here refers to a concept in the diffusion network training process, specifically a discrete point in time (or discrete step) experienced during the process of transforming the original data from its original state to a noisy state (which can be called a forward process, for example, in a forward process the model adds noise to the data stepwise, and this forward process can be divided into multiple discrete steps, each step of which adds a certain amount of noise to the data), and also to the process of restoring the data from the noisy state back to its original state (or original data) (a reverse process, for example, in a reverse process the model attempts to restore the data to its original state by progressively removing the noise, and this reverse process can also be divided into multiple steps).
[0042] Thus, the solution of this disclosure utilizes the generated initial multimodal features and preset noise features to train an N-layer spread spectrum network in a preset model, enabling the trained target model to process multimodal information, enriching the model's usage scenarios, effectively improving the model's ability to grasp and control details, and enhancing the controllability of image generation.
[0043] Furthermore, because the solution in this disclosure makes full use of multimodal features of mixed image and text input, i.e., a general-purpose visual language, the expressed features have richer descriptive information, and this descriptive information can provide more context and detailed guidance for reasoning, such as image generation. In this way, it can further satisfy the user's higher reasoning needs, and the trained model can be equipped with multi-entity processing and multi-entity generation capabilities, further improving the user experience.
[0044] Herein, as an example, the preset model described above may be a generative model, for example, specifically an image generation model, and the solution of this disclosure is not specifically limited to this.
[0045] In addition, when the sample data specifically includes sample text containing a plurality of entity texts and a plurality of entity images, and there is a corresponding entity image for each entity text, based on the normal browsing order of the texts in the sample text, the initial multi-modal features can be represented as "[text features of sub-text 1 (which may be understood as non-entity text) in the sample text, text features of entity text 1, image features of the entity image corresponding to entity text 1; text features of sub-text 2 (which may be understood as non-entity text) in the sample text, text features of entity text 2, image features of the entity image corresponding to entity text 2,...]". The above feature representation method is merely an example. In actual applications, it should be understood that the specific positions of the text features of the non-entity text and the text features of the entity text in the initial multi-modal features can be determined based on the positions of the non-entity text and the entity text in the sample text.
[0046] For example, as shown in Figure 2, the sample text included in the sample data is "Lin Daiyu took a selfie at the Forbidden City" (Note: In Japanese, it means "Lin Daiyu is taking a selfie at the Forbidden City"). This sample text is an entity text containing "Lin Daiyu" (which can be denoted as entity text 1) and "the Forbidden City" (which can be denoted as entity text 2). The non-entity texts included in this sample text are "at" and "took a selfie". Furthermore, the sample data includes entity image 1 corresponding to "Lin Daiyu" (i.e., entity text 1) and entity image 2 corresponding to "the Forbidden City" (i.e., entity text 2). At this time, the image feature 1 extracted from entity image 1 can be inserted after the text feature 1 of "Lin Daiyu", and the image feature 2 extracted from entity image 2 can be inserted after the text feature 3 of "the Forbidden City", and the following initial multi-modal features can be obtained as follows.
[0047] [Text feature 1 of "Lin Daiyu", image feature 1 of the corresponding entity image 1 of "Lin Daiyu"; text feature 2 of "at"; text feature 3 of "the Forbidden City", image feature 2 of the corresponding entity image 2 of "the Forbidden City"; text feature 4 of "took a selfie"].
[0048] Thus, the image-text alignment strategy described above allows for the alignment of information from different modalities, enabling the realization of mixed image-text feature representation and contributing to increased controllability in subsequent image generation.
[0049] Furthermore, in the above example, in order to more clearly distinguish between text features and image features in the initial multimodal features, the method disclosed herein may also include a header field (head) before the inserted image features (e.g., a header). For example, taking the example shown in Figure 2, by concatenating a header field (head) before image feature 1 and a header field (head) before image feature 2, it is possible to distinguish between text features and image features in the initial multimodal features, further improving the controllability of image generation and enhancing the inference effect of the model.
[0050] Furthermore, in one specific example, the feature representation method for image features of an actual image obtained using an image-text encoder includes at least one of the following: positional features of the actual image, segmentation features of a segmentation diagram corresponding to the actual image, image features of a cropped diagram of the actual image, image depth features of the actual image, and local features of the actual image (for example, face features, or even human face features).
[0051] In other words, when constructing the initial multimodal features, the inserted image features can take on multiple forms. Thus, compared to using conventional single red-green-blue (RGB) features, the solution of this disclosure can support features in more representational forms, enriching the representational forms of the initial multimodal features, improving the broad applicability of the model, and effectively enhancing the model's semantic comprehension ability. It also possesses better generalization capabilities when faced with different types of inputs, laying the foundation for generating high-quality reasoning results that meet user needs.
[0052] Furthermore, in one specific example, the feature representation method for the image features of the real image can be determined as follows, and the initial multimodal features of the sample data obtained above specifically include the following.
[0053] In Method 1, using the image / text encoder in the preset model, one or more feature representation methods are selected from the above-mentioned methods based on the sample features of the sample data (for example, based on the task features of the generation task indicated by the sample data) to obtain the image features of the actual images contained in the sample data, and consequently, the initial multimodal features.
[0054] In Method 2, one or more of the above feature representation methods are randomly selected using the image / text encoder in the preset model to obtain the image features of the actual images included in the sample data.
[0055] Furthermore, by selecting multiple types from the above feature representation methods using Method 1 or Method 2, it is possible to expand a single sample data to obtain multiple initial multimodal features. In this way, the sample volume for model training is effectively enriched, and the training efficiency of the model training is further improved.
[0056] During the model training phase, either of the two methods described above can be selected and implemented. Furthermore, the method selected may be the same or different depending on the sample data, and the solutions of this disclosure are not limited thereto.
[0057] Furthermore, to further improve the generalization ability of the model, the feature representation methods selected may differ depending on the sample data. In this way, the trained target model can accommodate any reasoning needs, enriching the usage scenarios and further improving the user experience.
[0058] Furthermore, the solutions in this disclosure do not impose any specific restrictions on how the above-mentioned feature representation methods are selected, and can be adjusted, for example, in a scenario based on the requirements of the actual scenario or the requirements of training.
[0059] The entities described in the embodiments of this disclosure may be any form of image, such as a division diagram, a positional coordinate diagram, a crop diagram, or a line art diagram, and the embodiments of this disclosure are not limited thereto.
[0060] Thus, this disclosure allows for the selection of appropriate image feature representation methods based on the actual conditions of the sample data, enabling the construction of diversified initial multimodal features. Compared to using conventional single RGB features, the solution of this disclosure can support features in a wider range of representation formats and has more diversified initial multimodal features. In this way, the semantic comprehension ability of the model can be enhanced in subsequent model training, the controllability of the model in generating image content can be improved, and more complex generation tasks can be handled better by utilizing different feature representation methods, thereby further improving the reasoning efficiency of the model and laying the foundation for an improved user experience.
[0061] In one specific example of the solution of this disclosure, each of the N-layer spreading networks is a dual information stream spreading network, and in one example, the dual information stream spreading network may specifically include a first information stream for processing initial multimodal features (for example, in one example, it may specifically be a condition information stream or may be called a multimodal condition information stream) and a second information stream for compressing the initial multimodal features based on preset noise features in order to represent the initial multimodal features in hidden layer space (for example, in one example, it may specifically be a latent layer space stream).
[0062] Here, a dual information stream architecture is employed, and one branch of this dual information stream can effectively handle multimodal features. In this way, the semantic comprehension and controllability of the model are enhanced, ensuring that the model can effectively handle multimodal information input and complete generation tasks (e.g., image generation tasks or text reasoning tasks) under multimodal conditions, further improving the model's reasoning ability and enriching the model's usage scenarios.
[0063] Furthermore, the feature dimension of the features represented in the hidden layer space is smaller than that of the initial multimodal features. In this way, it is possible to effectively enhance the inference effect while also maintaining inference efficiency.
[0064] For example, as shown in Figure 3(a), the first information stream is used to process the input initial multimodal features, and the second information stream is used to compress the initial multimodal features based on preset noise features. Specifically, in one example, initial multimodal features with a feature dimension of 128 × 128 × 3 are mapped to the hidden layer space and compressed to obtain new features with a feature dimension of 32 × 32 × 3, where the initial multimodal features reside in the hidden layer space. Furthermore, preset noise features can also be introduced in the above process, which effectively enhances the model's reasoning ability, allows for more effective learning of data features, and enables the generation of higher quality reasoning results.
[0065] Thus, the solution of this disclosure can perform model reasoning using a dual information street spreading network, and one branch of the dual information street can effectively process multimodal features. In this way, the model can effectively respond to multimodal information input, understand the generation task from different angles, and retain more detail and semantic information in the subsequent task generation process, efficiently generating reasoning results that match the user's expectations. Furthermore, by introducing noise to other branches of the dual information stream, the diversity and stability of the model in content generation can be improved.
[0066] Furthermore, in one specific example, when the value of N is 2 or greater, the N spreading networks are connected in series. As shown in Figure 3(b), the corresponding output of the first information stream in the i-th spreading network (where i is a natural number greater than 0 and less than or equal to N-1) among the N spreading networks is the input of the first information stream in the (i+1)-th spreading network among the N spreading networks, and the output of the second information stream in the i-th spreading network among the N spreading networks is the input of the second information stream in the (i+1)-th spreading network among the N spreading networks. In this way, by utilizing the N spreading networks connected in series, the feature representation ability of the model can be enhanced, further improving the overall performance of the model, and effectively increasing the generalization ability and robustness of the model, thus laying the foundation for further enhancing the user experience.
[0067] Figure 4 is a schematic flowchart, part two, of a model training method according to one embodiment of the present disclosure. This method can be selectively applied to electronic devices such as personal computers, servers, and server clusters. The relevant aspects of the method shown in Figures 1 to 3(a) and 3(b) above can be applied in this embodiment, and it should be understood that the relevant aspects will not be repeated.
[0068] Furthermore, the method includes at least a portion of the following: As shown in Figure 4, it includes the following:
[0069] In step S401, sample data is input into the preset model to obtain the initial multimodal features of the sample data.
[0070] Here, the sample data includes at least sample text and at least one entity image. Furthermore, the sample text includes at least one entity text, and the entity image is an image corresponding to the entity text.
[0071] Furthermore, the initial multimodal feature includes image features of the entity image and text features of the entity text, where the image features of the entity image follow the text features of the corresponding entity text; that is, in one example, the initial multimodal feature is obtained by inserting the image features of the entity image at least after the text features of the corresponding entity text.
[0072] Details regarding the sample data can be found in the explanation above and will not be repeated here.
[0073] In step S402, instruction information is inserted into the initial multimodal feature.
[0074] For example, if it is determined that at least a portion of an entity image needs to be image-adjusted based on sample data, such as when a portion of a particular entity image needs to be adjusted, or when some of a particular entity image needs to be image-adjusted, then instruction information can be inserted into the initial multimodal features, for example, into the feature sequence representing the initial multimodal features.
[0075] For example, in the case of task generation type sample data, if the generation task indicated by the sample data requires image adjustment of at least a portion of the entity image (e.g., adjusting a specific region of an entity image, or adjusting some of a specific entity image out of all the entity images), instruction information can be inserted into the initial multimodal feature.
[0076] Here, the instruction information can specifically include image features of the entity image that needs adjustment and adjustment instructions. Furthermore, the adjustment instructions can specifically include actions such as adding, deleting, changing attributes, swapping, scaling, and moving. In this way, introducing instruction information into model training can effectively improve the model's multimodal instruction capability.
[0077] In step S403, if the parameters of the image / text encoder of the preset model are fixed, the initial multimodal features (i.e., the initial multimodal features into which the instruction information was concatenated (or inserted) in step S402) and the preset noise features are used to train the N-layer spread network in the preset model to obtain the target model.
[0078] Here, the target model can infer a target reasoning result, including images and / or text, related to the multimodal prompt information, based on the input multimodal prompt information. Furthermore, N is a positive integer greater than or equal to 1.
[0079] Here, explanations of diffusion networks and target models can be found in the explanations above and will not be repeated here.
[0080] Thus, the solution of this disclosure allows for the addition of instruction information to the obtained initial multimodal features and the use of the initial multimodal features with added instruction information to train a preset model. In this way, the instruction capabilities of the model are effectively improved through the instruction information, which in turn improves the model's ability to understand the generation task shown in the sample data, allowing the generated images to be better adapted to the generation task, enabling precise control over the generated image content, and further enhancing the user experience.
[0081] Furthermore, in one specific example, the following method can be employed to insert instruction information. Specifically, inserting the instruction information into the initial multimodal feature (for example, step S402 described above) is specifically: This includes concatenating the instruction information to the header of a feature sequence representing the initial multimodal features.
[0082] For example, taking the feature sequence of the initial multimodal feature shown in Figure 2 as an example, if the generation task indicated by the sample data requires image adjustment of a portion of the actual image, instruction information can be concatenated to the header of the feature sequence representing the initial multimodal feature, as shown in Figure 5(a). Furthermore, as shown in Figure 5(b), this instruction information can be represented by a sequence of two parts: one part is a text feature sequence corresponding to the adjustment instruction (for example, it can be written as Instruct Text Tokens), and the other part is an image feature sequence corresponding to the image features of the actual image that need adjustment (for example, it can be written as Global Image Tokens).
[0083] Furthermore, similar to the example shown in Figure 2, a header field `head` may be added before image features (e.g., Global Image Tokens) to distinguish between text features and image features. This allows for differentiation between text and image features in the instruction information, which in turn helps to further improve the controllability of image generation and the inference effect of the model.
[0084] Thus, the solution of this disclosure provides a refinement method for linking initial multimodal features with instruction information, thereby effectively improving the instructional capabilities of the model and enabling precise control over the image content generated during model training, ultimately laying the foundation for generating images that meet user needs.
[0085] Figure 6(a) is a schematic flowchart, part three, of a model training method according to one embodiment of the present disclosure. This method can be selectively applied to electronic devices such as personal computers, servers, and server clusters. The relevant aspects of the methods shown in Figures 1 to 5 above can be applied to this embodiment, and it should be understood that these relevant aspects will not be repeated.
[0086] Furthermore, the method includes at least a portion of the following: As shown in Figure 6(a), it includes the following:
[0087] In step S601, sample data is input into the preset model to obtain the initial multimodal features of the sample data.
[0088] Here, the sample data includes at least sample text and at least one entity image. Furthermore, the sample text includes at least one entity text, the entity image is the image corresponding to the entity text, and the initial multimodal feature includes the image feature of the entity image and the text feature of the entity text, where the image feature of the entity image follows the text feature of the entity text corresponding to the entity image.
[0089] Details regarding the sample data can be found in the explanation above and will not be repeated here.
[0090] In step S602, a mask operation (e.g., a random mask operation) is performed on some of the features in the initial multimodal feature (or initial multimodal feature in which instruction information is concatenated) to obtain a masked initial multimodal feature.
[0091] For example, taking the feature sequence of the initial multimodal feature shown in Figure 5(a) as an example, in this case, a random mask operation is performed on some of the features in the feature sequence representing the initial multimodal feature. For example, some tokens (word elements) in the initial multimodal feature are randomly masked with a certain probability. As shown in Figure 6(b), text feature 1 in the initial multimodal feature is masked (corresponding to mask case 1), or image feature 1 in the initial multimodal feature is masked (corresponding to mask case 2), thereby obtaining a masked initial multimodal feature. Here, the solution of this disclosure can mask at least some of the features of the initial multimodal feature, thereby enabling the processing of finer-grained features in the first information street thereafter, and consequently enabling finer-grained modeling. This effectively improves the model's ability to grasp and control details, improves the model's learning ability for fine-grained tasks, and lays the foundation for further improving the accuracy of inference results.
[0092] In step S603, if the parameters of the image / text encoder of the preset model are fixed, the masked initial multimodal features are used as input to the first information stream of the first layer of the N-layer spread network, and the preset noise features are used as input to the second information stream of the first layer of the N-layer spread network. Feature diffusion is then performed using the N-layer spread network to obtain the overall prediction result.
[0093] Here, the overall prediction results can specifically include the first prediction result and the second prediction result.
[0094] Furthermore, the first prediction result is the output of the first information stream of the final layer of the N-layer spread network, and the first prediction result can at least represent the predicted mask position corresponding to the predicted random mask operation. That is, the first prediction result is used to predict the position of the mask operation. For example, in one example, the first prediction result is used to represent the feature to be masked, and / or to represent the specific position of the feature to be masked in the initial multimodal features before the mask operation. In this way, feature reconstruction is realized, the model learns finer-grained generation tasks, achieves finer-grained modeling, and ultimately helps to have the ability to generate multimodal information (for example, generating bounding boxes (Bboxes) and face ID for entities indicated by given entity text, or generating text information for an entity image from a given entity image).
[0095] Furthermore, the second prediction result is the output of the second information stream of the diffusion network in the final layer of the N-layer diffusion network, and this second prediction result can specifically represent the prediction result in the hidden layer space of the generation task indicated by the predicted sample data. That is, the second prediction result is used to represent the representation of the prediction result in the hidden layer space, and this prediction result refers to the generation result required for the task indicated by the sample data. For example, in one example, this prediction result represents the prediction result in the hidden layer space output by the model in the current step.
[0096] In step S604, the loss value of the target loss function is obtained based on the first prediction result and the second prediction result.
[0097] Here, the target loss function is used to calculate the difference between the predicted mask position and the actual mask position, and is used to calculate the difference between the predicted result in the hidden layer space of the generation task shown by the sample data and the actual result in the hidden layer space (e.g., the actual result in the hidden layer space of the current step).
[0098] Here, the actual results in the hidden layer space can be understood as the theoretical results that should be theoretically achieved in the current step.
[0099] In step S605, based on the loss value of the target loss function, at least some of the tunable network parameters in the N-layer spreading network are adjusted until the number of training iterations (e.g., the number of training iterations up to the current step) and / or the loss value of the target loss function satisfy the model training requirement, and then the training for the current step is completed, thereby obtaining the target model.
[0100] In the scenario of training a spreading network, steps S601 to S605 above are specifically applied to any of the total number of steps required to train the spreading network (or spreading model), and therefore, steps S601 to S605 specifically apply to:
[0101] The sample data used in the current step is input into a preset model, and the initial multimodal features of the sample data used in the current step are obtained using the image / text encoder in the preset model.
[0102] This involves performing a masking operation on some of the features of the initial multimodal features (or initial multimodal features formed by concatenating instruction information) to obtain masked initial multimodal features,
[0103] When the parameters of the image / text encoder are fixed, the masked initial multimodal features are used as input to the first information stream of the first layer of the N-layer spread network, and the preset noise features used in the current step are used as input to the second information stream of the first layer of the N-layer spread network. Feature diffusion is then performed using the N-layer spread network to obtain the overall prediction result corresponding to the training in the current step.
[0104] Based on the overall prediction results, we obtain the loss value of the target loss function, i.e., the training loss value for the current step.
[0105] This includes adjusting the tunable network parameters of at least some of the N-layer spreading network based on the training loss value of the current step to complete the training of the current step, and further obtaining the target model by completing the training of all steps.
[0106] For example, as shown in Figure 6(c), first, initial multimodal features are obtained using the image / text encoder in the preset model. A random mask operation is performed on some of the features in the initial multimodal features to obtain masked initial multimodal features. Next, the masked initial multimodal features are used as input to the first information stream of the first layer of the N-layer spread network in the preset model, and the preset noise features are used as input to the second information stream of the first layer of the spread network. Feature diffusion is performed using the N-layer spread network, thereby obtaining the output of the first information stream of the final layer of the spread network in the N-layer spread network (i.e., the first prediction result) and the output of the second information stream of the final layer of the spread network in the N-layer spread network (i.e., the second prediction result). Then, the loss value of the target loss function is calculated based on the first and second prediction results. Finally, at least some of the tunable network parameters of the N-layer spread network in the preset model are adjusted based on the loss value of the target loss function. In this way, training for the current step is performed, and upon completion of training for all steps, the target model is obtained.
[0107] Thus, the solution of this disclosure provides a refinement method for training an N-layer spread network in a preset model, efficiently training the target model. This training method effectively improves the model's learning ability for fine-grained tasks and its ability to grasp and control details, thereby improving the model's controllability in generating image content. The resulting image content can be better adapted to user-specified generation tasks, meeting the user's need for greater customization and further enhancing the user experience.
[0108] Furthermore, in one specific example of the solution of this disclosure, as shown in Figure 7(a), the dual information stream spreading network described above may include a first condition information stream module and a second condition information stream module corresponding to the first information stream, and a first latent layer spatial stream module and a second latent layer spatial stream module corresponding to the second information stream.
[0109] Furthermore, in this example, in a dual information stream spreading network, the input to the first information stream is the input to the first conditional information stream module, and the output to the second conditional information stream module is the output to the first information stream. Accordingly, the input to the second information stream is the input to the first latent layer spatial stream module, and the output to the second latent layer spatial stream module is the output to the second information stream.
[0110] For example, in the first layer of a dual information stream spreading network in an N-layer dual information stream spreading network, the input to the first information stream is an initial multimodal feature, and the input to the second information stream is a preset noise feature. In this case, the initial multimodal feature is the input to the first conditional information stream module in the first layer of the dual information stream spreading network, and the preset noise feature is the input to the first latent layer spatial stream module in the first layer of the dual information stream spreading network. Accordingly, the output of the second conditional information stream module in the first layer of the dual information stream spreading network is the output of the first information stream in the first layer of the dual information stream spreading network, and the output of the second latent layer spatial stream module in the first layer of the dual information stream spreading network is the output of the second information stream in the first layer of the dual information stream spreading network.
[0111] Here, the first conditional information stream module in the above example mainly includes a layernorm layer, a modulo operation layer, and a linear layer, processing from top to bottom (e.g., processing information from top to bottom). Accordingly, the first latent layer space stream module may also include a layernorm layer, a modulo operation layer, and a linear layer, processing from top to bottom (e.g., processing information from top to bottom).
[0112] Furthermore, the second conditional information stream module in the above example may mainly include linear layers, normalization layers, modulo layers, and multi-layer perceptrons (MLPs) that process information from top to bottom (e.g., processing information from top to bottom). Accordingly, the second latent layer spatial stream module may also mainly include linear layers, normalization layers, modulo layers, and multi-layer perceptrons (MLPs) that process information from top to bottom (e.g., processing information from top to bottom).
[0113] It should be noted that the above is merely an example, and in actual applications, the processing layers used in each information street module and each spatial flow module can be configured according to the actual needs, and the solutions disclosed herein are not limited to this.
[0114] Thus, this disclosure provides a specific architecture for a dual information stream spreading network, in which one branch of the architecture can effectively handle multimodal features, allowing the model to understand the generation task from different angles, further retaining more detail and semantic information in the subsequent task generation process, and enabling precise control over image content, reducing information loss, and ultimately ensuring that the generated image matches the user's expectations, thereby improving the user experience. Another branch of the architecture introduces noise, improving the diversity and stability of the model's content generation, thus effectively improving the overall flexibility and extensibility of the model, and ultimately helping to improve the user experience.
[0115] Furthermore, in one specific example, feature diffusion using the N-layer diffusion network described above can be performed as shown in Figure 7(b), specifically, The following steps involve spreading features using the i-th layer (in this case, i is an integer greater than 0 and less than or equal to N) of an N-layer spread network (or N-layer dual information stream spread network). In other words, the processing logic of the i-th layer spread network is as follows:
[0116] In step S701, the output of the first information stream in the i-1th layer diffusion network (i.e., the output of the first information stream in the diffusion network of the previous layer of the current layer's diffusion network, for example, specifically the output of the second condition information stream module in the i-1th layer's diffusion network) is input to the first condition information stream module in the i-th layer's diffusion network (which can be understood as the current layer's diffusion network) for processing, and the output of the second information stream in the i-1th layer's diffusion network (i.e., the output of the second information stream in the diffusion network of the previous layer of the current layer's diffusion network, for example, the output of the second latent layer information stream module in the i-1th layer's diffusion network) is input to the first latent layer spatial stream module in the i-th layer's diffusion network for processing.
[0117] In step S702, the processed features of the first condition information street module in the i-th layer diffusion network and the processed features of the first latent layer spatial stream module in the i-th layer diffusion network are manipulated at the element level to connect the processed features of the first condition information street module and the processed features of the first latent layer spatial stream module.
[0118] In step S703, the combined features are subjected to self-attentional processing to obtain fused features.
[0119] In step S704, the features of the fused features corresponding to the first information stream are input to the second condition information stream module in the i-th layer diffusion network to obtain the output of the second condition information stream module, and the features of the fused features corresponding to the second information stream are input to the second latent layer spatial stream module in the i-th layer diffusion network to obtain the output of the second latent layer spatial stream module.
[0120] In this example, when i takes the value 1, the input to the first condition information stream module in the i-1th layer spread network is a masked initial multimodal feature (or an initial multimodal feature formed by concatenating masked instruction information), and the input to the first latent layer spatial stream module in the i-1th layer spread network is a preset noise feature.
[0121] Thus, the solution of this disclosure provides a specific method for spreading features using an N-layer dual information stream spreading network, which in this way can effectively capture and propagate key features and details in text and / or images, enable the model to understand the generation task from different angles, retain more details and semantic information in subsequent task generation processes, achieve precise control over image content, ensure that the generated images match user expectations, improve the diversity and stability of the model's content generation, effectively enhance the overall flexibility and extensibility of the model, ultimately help meet users' higher customization needs and improve the user experience.
[0122] Part Two will provide a detailed explanation of the model's usage scenarios.
[0123] Figure 8 is a schematic flowchart of a multimodal data processing method according to one embodiment of the present disclosure. This method can be selectively applied to electronic devices such as personal computers, servers, and server clusters.
[0124] Furthermore, the method includes at least a portion of the following: As shown in Figure 8, it includes the following:
[0125] In step S801, multimodal prompt information is input to a pre-trained target model to obtain the target multimodal features of the multimodal prompt information.
[0126] For example, in one scenario, multimodal prompt information can be input into a pre-trained (or post-trained) target model, and then the target multimodal features of the multimodal prompt information can be obtained using the image / text encoder in the target model.
[0127] In this example, the multimodal prompt information includes at least the target text and at least one target entity image (for convenience, as will be explained later, this can be referred to as M3 target entity images). Furthermore, the target text includes at least one target entity text (for convenience, as will be explained later, this can be referred to as M3 target entity texts). In this case, the target entity image is the image corresponding to the target entity text.
[0128] Here, M3 and M4 are positive integers greater than or equal to 1. Furthermore, in one example, the values of M3 and M4 may be different, for example, M3 may be less than M4, which may correspond to a scenario where not all sample texts necessarily correspond to a single entity image. Alternatively, in another example, the values of M3 and M4 may be the same, which may correspond to a scenario where one target entity text corresponds to one target entity image.
[0129] Here, the target entity text and target entity image may be one-to-one or one-to-many, and the solutions of this disclosure are not limited thereto. Accordingly, the solutions of this disclosure do not limit the values of M3 and M4.
[0130] Furthermore, in one example, the target multimodal feature includes image features of a target entity image and text features of a target entity text, wherein the image features of the target entity image are after the text features of the target entity text corresponding to the target entity image; in other words, in this example, the target multimodal feature is obtained by inserting at least the image features of the target entity image after the text features of the target entity text corresponding to the target entity image.
[0131] Furthermore, the target model in this example was trained based on one of the model training methods described above.
[0132] In step S802, feature diffusion is performed on the target multimodal features using the N-layer target diffusion network in the target model to obtain a target inference result, including images and / or text, related to the multimodal prompt information.
[0133] Here, N is a positive integer greater than or equal to 1.
[0134] Thus, the solution of this disclosure can obtain a target reasoning result, including images and / or text, related to the multimodal prompt information, based on the input multimodal prompt information, using a target model. Compared to conventional solutions, the target model of the solution of this disclosure has multi-entity processing capability and multi-entity generation capability, enriching and improving the user experience.
[0135] Furthermore, because the solution disclosed herein makes full use of multimodal prompt information for mixed image and text input, in scenarios requiring image generation, the solution disclosed herein can provide more context and detailed guidance, thereby effectively improving the controllability of image content generation and further satisfying the user's higher reasoning needs.
[0136] Furthermore, in the model training stage described above, the model training method can be applied primarily to training the "current step" of the spreading network; in other words, the model training method can be applied primarily to training any step of the spreading network. On the other hand, in the model usage stage of this example, in order to obtain the target reasoning result described above, it is necessary to perform all steps using the target spreading network.
[0137] Furthermore, if the multimodal prompt information specifically consists of sample text containing multiple target entity texts and multiple target entity images, and each target entity text has a corresponding target entity image, then based on the normal browsing order of the text in the target text, the target multimodal feature can be represented as follows: "text features of subtext 1 (which may be understood as target non-entity text) in the target text, text features of target entity text 1, image features of the target entity image corresponding to target entity text 1; text features of subtext 2 (which may be understood as target non-entity text) in the target text, text features of target entity text 2, image features of the target entity image corresponding to target entity text 2;..." It should be understood that the above feature representation method is merely illustrative, and in actual applications, the specific positions of the text features of target non-entity texts and target entity texts in the target multimodal feature can be determined based on the positions of the target non-entity texts and target entity texts in the target text.
[0138] In the example described above, referring to the example shown in Figure 2, the image-text alignment strategy can align information of different modalities in multimodal prompt information, thus enabling the realization of mixed image-text feature representation and helping to increase the controllability of subsequent image generation.
[0139] Furthermore, in the above example, to more clearly distinguish between text and image features in the target multimodal feature, the solution of this disclosure may also include a header field (head) before the inserted image feature (e.g., a header), which helps to distinguish between text and image features in the target multimodal feature, further improve the controllability of image generation, and enhance the inference effect of the model. A specific example can be found in Figure 2, which will not be repeated here.
[0140] Furthermore, in one specific example, the following method can be employed to obtain the target multimodal features. Specifically, obtaining the target multimodal features of the multimodal prompt information described above (step S801 above) can be done in more detail, Using the image / text encoder in the target model, one or more feature representation methods are selected from the data features of the multimodal prompt information (for example, the task features of the target generation task indicated by the multimodal prompt information), such as positional features of the target entity image, segmentation features of the segmentation diagram corresponding to the target entity image, image features of the cropped diagram of the target entity image, image depth features of the target entity image, and local features of the target entity image (for example, face features, or even human face features), to obtain the image features of the target entity image included in the multimodal prompt information and obtain the target multimodal features.
[0141] In other words, the solution of this disclosure can determine a method of representing image features that is suitable for the task needs based on the task needs of the input multimodal prompt information. In other words, the solution of this disclosure can flexibly handle various forms of image features, thus enriching the representation formats of target multimodal features, possessing better generalization ability when faced with different types of task needs, and subsequently laying the foundation for generating high-quality reasoning results that meet user needs, and further laying the foundation for improving the user experience.
[0142] The target entity images described in the solutions of this disclosure may be any form of image, such as a division diagram, a location coordinate diagram, a crop diagram, or a line art diagram, and the solutions of this disclosure are not limited to these.
[0143] Thus, the solution of this disclosure can select an appropriate method of representing image features based on the actual situation of the input multimodal prompt information, and can construct diversified target multimodal features. Compared to using conventional single RGB features, the solution of this disclosure can support features in more representation formats and can obtain even more diversified target multimodal features. In this way, it helps to better handle more complex generation tasks, further meets the higher reasoning needs of users, and ultimately lays the foundation for improving the user experience.
[0144] Furthermore, in one specific example, before performing feature diffusion on the target multimodal features using the N-layer target diffusion network in the target model, it is also possible to adjust the target multimodal features based on the target generation task indicated in the multimodal prompt information. Specifically, Target instruction information is inserted into the target multimodal features, where the target instruction information includes image features of the target entity image requiring adjustment and target adjustment instructions.
[0145] For example, if it is determined that at least a portion of a target entity image needs to be image-adjusted based on target multimodal features, such as adjusting a portion of a specific target entity image or adjusting several portions of a specific target entity image, target instruction information can be inserted into the feature sequence representing the target multimodal features. Furthermore, feature diffusion can be performed again on the target multimodal features into which the target instruction information has been inserted, using the N-layer target diffusion network in the target model.
[0146] For example, in one instance, if the multimodal prompt information is prompt information for task generation type information, then target instruction information can be inserted into the target multimodal feature if the target generation task indicated in the multimodal prompt information requires image adjustment of at least a portion of the target entity image (for example, adjusting a portion of a specific target entity image, or image adjustment of several specific target entity images).
[0147] Here, the target instruction information includes image features of the target entity image that needs adjustment and target adjustment instructions. Furthermore, the target adjustment instructions may specifically be adjustment instructions such as adding, deleting, changing attributes, exchanging, scaling, or moving. In this way, the target instruction information is effectively used to instruct the model's reasoning process, further enhance the model's ability to understand the target generation task indicated by the multimodal prompt information, and subsequently adapt the generated image to the target generation task, thereby achieving precise control over the generated image content and further improving the user experience.
[0148] Furthermore, in one specific example, target instruction information may be inserted as follows, specifically by inserting the target instruction information into the target multimodal feature described above, specifically, The target instruction information is concatenated to the header of the target feature sequence representing the target multimodal features.
[0149] For example, as in the example in Figure 5(a), if the target generation task indicated in the multimodal prompt information needs to adjust a portion of the target entity image, the target instruction information can be concatenated to the header of the target feature sequence representing the target multimodal feature. Furthermore, as in the example shown in Figure 5(b), the target instruction information can be represented by a sequence of two parts: one part is a sequence of text features corresponding to the target adjustment instruction (for example, it can be written as Instruct Text Tokens), and the other part is a sequence of image features corresponding to the image features of the entity image that need adjustment (for example, it can be written as Global Image Tokens).
[0150] Furthermore, similar to the example shown in Figure 2, by adding a header field `head` before image features (e.g., Global Image Tokens) to distinguish between text and image features, it is possible to differentiate between text and image features in the target instruction information, which in turn helps to further improve the controllability of image generation and enhance the inference effect of the model.
[0151] Thus, the solution of this disclosure provides a refinement method for concatenating target multimodal features and target instruction information, thereby effectively improving the instructional capabilities of the model and enabling precise control over the image content generated during model reasoning, ultimately laying the foundation for generating images that meet the user's needs.
[0152] In one specific example of the present disclosure, each of the N-layer goal spread networks is a dual information street goal spread network, and in one example, the dual information stream goal spread network may specifically include a first goal information stream for processing goal multimodal features (for example, the first goal information stream may specifically be a condition information stream, or may be called a first goal multimodal condition information stream) and a second goal information stream for compressing goal multimodal features based on preset goal noise features in order to represent the goal multimodal features in hidden layer space (for example, in one example, it may specifically be a latent layer space stream)).
[0153] Here, a dual information stream architecture is employed, and one branch of this dual information stream can effectively handle multimodal features. In this way, the semantic comprehension and controllability of the model are enhanced, and the model can effectively ensure that it can efficiently complete goal generation tasks (e.g., image generation tasks or text reasoning tasks) under multimodal conditions, further improving the model's reasoning ability and enriching the model's usage scenarios.
[0154] Furthermore, the feature dimension of the features represented in the hidden layer space is smaller than that of the target multimodal features. In this way, it is possible to effectively enhance the inference effect while also maintaining inference efficiency.
[0155] For example, in one example, similar to the example shown in Figure 3(a), in a dual-information-street goal diffusion network, the first goal information street is used to process the input goal multimodal features (or goal multimodal features formed by concatenating goal instruction information), and the second goal information street is used to compress the goal multimodal features based on pre-set goal noise features. For example, in one example, a goal multimodal feature with a feature dimension of 128 × 128 × 3 is mapped to the hidden layer space and compressed to obtain a new feature in the hidden layer space with a feature dimension of 32 × 32 × 3. Furthermore, preset goal noise features can also be introduced in the above process, which effectively enhances the model's reasoning ability, allows for more effective learning of data features, and enables the generation of higher-quality reasoning results.
[0156] Thus, the solution of this disclosure can perform model reasoning using a goal-spreading network of dual information streams, and one branch of the dual information stream can effectively process multimode features. In this way, the model can effectively support multimode information input, understand the generation task from different angles, and retain more detail and semantic information in the subsequent task generation process, enabling the efficient generation of reasoning results that match the user's expectations. Furthermore, introducing noise into the other branch of the dual information stream can improve diversity and stability in content generation.
[0157] Furthermore, in one specific example, when the value of N is 2 or greater, the N spreading networks are connected in series. Additionally, similar to the example shown in Figure 3(b), the corresponding output of the first target information stream in the i-th target spreading network (where i is a natural number greater than 0 and less than or equal to N-1) among the N spreading networks is the input to the first information stream in the (i+1)-th target spreading network among the N spreading networks, and the output of the second target information stream in the i-th target spreading network among the N spreading networks is the input to the second target information stream in the (i+1)-th target spreading network among the N spreading networks. In this way, by utilizing the N spreading networks connected in series, the feature representation ability of the model can be enhanced, further improving the overall performance of the model, and effectively increasing the model's generalization ability and robustness, thus laying the foundation for further enhancing the user experience.
[0158] Figure 9 is a schematic flowchart, part two, of a multimodal data processing method according to one embodiment of the present invention. This method can be selectively applied to electronic devices such as personal computers, servers, and server clusters. The relevant aspects of the method shown in Figure 8 above can be applied to this embodiment, and it should be understood that the relevant aspects will not be repeated.
[0159] Furthermore, the method includes at least a portion of the following: As shown in Figure 9, it includes the following:
[0160] In step S901, multimodal prompt information is input to the target model to obtain the target multimodal features of the multimodal prompt information.
[0161] Here, the aforementioned target model is trained based on the model training method described above.
[0162] Furthermore, the multimodal prompt information includes at least a target text and at least one target entity image, the target text includes at least one target entity text, the target entity image is an image corresponding to the target entity text, the target multimodal feature includes image features of the target entity image and text features of the target entity text, the image features of the target entity image follow the text features of the target entity text corresponding to the target entity image.
[0163] The content related to multimodal prompt information can be found in the example above and will not be repeated here.
[0164] In step S902, the target multimodal feature (or target multimodal feature formed by concatenating target instruction information) is used as the input to the first target information stream of the first layer of the N-layer target spread network, and the preset target noise feature is used as the input to the second target information stream of the first layer of the N-layer target spread network, and feature spread is performed using the N-layer target spread network to obtain an initial reasoning result.
[0165] Here, the initial reasoning result is obtained based on the output of the second target information stream of the final layer of the N-layer target diffusion network, and represents the predicted result in the hidden layer space of the predicted target generation task; that is, the initial reasoning result represents the representation of the prediction result in the hidden layer space, and the prediction result here specifically refers to the generation result required for the target generation task shown in the multimodal prompt information.
[0166] In step S903, if the initial reasoning result includes image features, the image features in the initial reasoning result are decoded using the image decoder in the target model to obtain the target reasoning result.
[0167] For example, as shown in Figure 10, first, multimodal prompt information input to the target object is input to the target model, and target multimodal features are obtained using the image / text encoder in the target model. Next, the target multimodal features are used as input to the first target information stream of the first layer of the N-layer target spread network, and the preset target noise features are used as input to the second target information stream of the first layer of the N-layer target spread network, and feature diffusion is performed using the N-layer target spread network. For example, feature diffusion is performed several times using the preset steps, and the result output by the second target information stream of the final layer of the N-layer target spread network (i.e., the output result of the second target information stream of the final layer of the final step), i.e., the initial reasoning result, is obtained. Finally, the final target reasoning result is obtained based on the initial reasoning result, or, if the initial reasoning result includes image features, the image features in the initial reasoning result are decoded using the image decoder in the target model to obtain the target reasoning result.
[0168] Thus, the solution of this disclosure provides a subdivision scheme for reasoning using an N-layer goal diffusion network in the target model, which enables precise control over image content generation, allowing the resulting image content to better suit the user-specified generation task, satisfying the user's higher customization needs and thereby further improving the user experience.
[0169] Furthermore, in one specific example of the solution of this disclosure, the dual information stream target diffusion network described above may include a first target condition information stream module and a second target condition information stream module corresponding to a first target information stream, and a first target latent layer spatial stream module and a second target latent layer spatial stream module corresponding to a second target information stream.
[0170] Furthermore, similar to the example shown in Figure 7(a), in this example, in the dual information stream target spreading network, the input to the first target information stream is the input to the first target condition information stream module, the output to the second target condition information stream module is the output to the first target information stream, and accordingly, the input to the second target information stream is the input to the first target latent layer spatial stream module, and the output to the second target latent layer spatial stream module is the output to the second target information stream.
[0171] For example, in the first layer of a dual information stream target spread network in an N-layer dual information stream target spread network, the input to the first target information stream is a target multimodal feature, and the input to the second target information stream is a preset target noise feature. In this case, the target multimodal feature is the input to the first target condition information stream module in the first layer of the dual information stream target spread network, and the preset target noise feature is the input to the first target latent layer spatial stream module in the first layer of the dual information stream target spread network. Accordingly, the output of the second target condition information stream module in the first layer of the dual information stream target spread network is the output of the first target information stream in the first layer of the dual information stream target spread network, and the output of the second target latent layer spatial stream module in the first layer of the dual information stream target spread network is the output of the second target information stream in the first layer of the dual information stream target spread network.
[0172] Thus, this disclosure provides a specific architecture for a dual-information street-target diffusion network, in which one branch of the architecture can effectively handle multimodal features, allowing the model to understand the generation task from different angles, further retaining more detail and semantic information in the subsequent task generation process, and enabling precise control over image content, reducing information loss, and ultimately ensuring that the generated images meet user expectations and improve the user experience. Another branch of the architecture introduces noise, improving the diversity and stability of the model's content generation, thus effectively improving the overall flexibility and scalability of the model and providing strong support for improving the user experience.
[0173] Furthermore, in one specific example, feature diffusion is performed using the N-layer target diffusion network described above, and specifically, similar to the example shown in Figure 7(b), The following steps involve spreading features using the i-th layer (where i is an integer greater than 0 and less than or equal to N) of an N-layer goal-oriented spread network (or N-layer dual information stream goal-oriented spread network). In other words, the processing logic of the i-th layer's goal-oriented spread network is as follows:
[0174] In step S1101, the output of the first target information stream in the i-1th layer's target diffusion network (i.e., the output of the first target information stream in the previous layer's target diffusion network of the current layer's target diffusion network, for example, specifically the output of the second target condition information stream module in the i-1th layer's target diffusion network) is input to the first target condition information stream module in the i-th layer's target diffusion network (which can be understood as the current layer's target diffusion network) for processing, and the output of the second target information stream in the i-1th layer's target diffusion network (i.e., the output of the second target information stream in the previous layer's target diffusion network of the current layer's target diffusion network, for example, the output of the second target latent layer information stream module in the i-1th layer's target diffusion network) is input to the first target latent layer spatial stream module in the i-th layer's target diffusion network for processing.
[0175] In step S1102, the processed features of the first target condition information street module in the i-th layer's target diffusion network and the processed features of the first target latent layer spatial stream module in the i-th layer's target diffusion network are manipulated at the element level to connect the processed features of the first target condition information street module and the processed features of the first target latent layer spatial stream module.
[0176] In step S1103, the combined features are processed by self-attention to obtain the target fused features.
[0177] In step S1104, the feature corresponding to the first target information stream among the target fusion features is input to the second target condition information stream module in the i-th layer of the target diffusion network to obtain the output of the second target condition information stream module, and the feature corresponding to the second target information stream among the target fusion features is input to the second target latent layer spatial stream module in the i-th layer of the target diffusion network to obtain the output of the second target latent layer spatial stream module.
[0178] In this example, when i takes the value 1, the input to the first target condition information stream module in the i-1th layer of the target spread network is a target multimodal feature (or a target multimodal feature formed by concatenating target instruction information), and the input to the first target latent layer spatial stream module in the i-1th layer of the target spread network is a preset noise feature.
[0179] Furthermore, for the target diffusion network of the final layer, the output of the second target latent layer spatial stream module in this target diffusion network of the final layer can be directly used as the initial reasoning result.
[0180] Thus, the solution of this disclosure provides a specific method for spreading features using an N-layer dual information stream target diffusion network, which can effectively capture and propagate key features and details in text and / or images, allowing the model to understand the generation task from different angles, retain more detail and semantic information in subsequent task generation processes, achieve precise control over image content, ensure that the generated images match the user's expectations, improve the diversity and stability of the model's content generation, and further help to meet the user's higher customization needs and improve the user experience.
[0181] The solution of this disclosure also provides a model training device, as shown in Figure 11, which is A training unit 1101 is used to input sample data into a preset model and obtain initial multimodal features of the sample data using an image / text encoder in the preset model, wherein the sample data includes at least sample text and at least one real image, the sample text includes at least one real text, the real image is an image corresponding to the real text, the initial multimodal features include image features of the real image and text features of the real text, the image features of the real image are after the text features of the real text corresponding to the real image, and if the parameters of the image / text encoder in the preset model are fixed, model training is performed on an N-layer spread network in the preset model using the initial multimodal features and preset noise features to obtain a target model, where N is a positive integer. The system includes a storage unit 1102 for outputting the aforementioned target model.
[0182] In one specific example of the solution disclosed herein, the training unit is: It is further used to insert instruction information, including image features of the real image that needs adjustment and adjustment instructions, into the initial multimodal features. For example, in one example, instruction information is inserted into the initial multimodal features when it is necessary to adjust the real image using the initial multimodal features and preset noise features before model training the N-layer spread spectrum network in the preset model.
[0183] In one specific example of the solution disclosed herein, the training unit specifically, It is used to concatenate the instruction information to the header of the feature sequence representing the initial multimodal features.
[0184] In one specific example of the solution of the present disclosure, the N-layer spread network is a dual information stream spread network, the dual information stream spread network includes a first information stream for processing initial multimodal features and a second information stream for compressing the initial multimodal features based on preset noise features in order to represent the initial multimodal features in hidden layer space.
[0185] In one specific example of the solution of this disclosure, when the value of N is 2 or greater, the N spreading networks are connected in series, the corresponding output of the first information stream in the i-th spreading network among the N spreading networks is the input of the first information stream in the (i+1)-th spreading network among the N spreading networks, and the output of the second information stream in the i-th spreading network among the N spreading networks is the input of the second information stream in the (i+1)-th spreading network among the N spreading networks.
[0186] In one specific example of the solution disclosed herein, the training unit is: This method is further used to obtain masked initial multimodal features by performing a masking operation on some of the features among the initial multimodal features.
[0187] In one specific example of the solution disclosed herein, the training unit specifically, The method involves performing feature diffusion using an N-layer diffusion network, with masked initial multimodal features as input to the first information stream of the diffusion network in the first layer of the N-layer diffusion network, and preset noise features as input to the second information stream of the diffusion network in the first layer of the N-layer diffusion network, to obtain a total prediction result, wherein the total prediction result includes: a first prediction result, which is the output of the first information stream of the diffusion network in the final layer of the N-layer diffusion network and represents at least the predicted mask position corresponding to the predicted mask operation; and a second prediction result, which is the output of the second information stream of the diffusion network in the final layer of the N-layer diffusion network and represents the prediction result in the hidden layer space of the generation task indicated by the predicted sample data. The process involves obtaining the loss value of the target loss function based on the first and second prediction results, where the target loss function is used to calculate the difference between the predicted mask position and the actual mask position, and is used to calculate the difference between the predicted result in the hidden layer space of the generation task shown by the sample data and the actual result in the hidden layer space. It is used to adjust at least some of the tunable network parameters in the N-layer spreading network based on the loss value of the target loss function.
[0188] In one specific example of the solution of this disclosure, the dual information stream spreading network includes a first conditional information stream module and a second conditional information stream module corresponding to a first information stream, and a first latent layer spatial stream module and a second latent layer spatial stream module corresponding to a second information stream, Here, the input to the first information stream is the input to the first conditional information stream module, and the output to the second conditional information stream module is the output to the first information stream. The input to the second information stream is the input to the first latent layer space stream module, and the output to the second latent layer space stream module is the output to the second information stream.
[0189] In one specific example of the solution disclosed herein, the training unit specifically, The following steps are used to perform feature diffusion using the i-th layer of the N-layer diffusion network: The aforementioned step is, The output of the first information stream in the i-1th layer diffusion network is input to the first conditional information stream module in the i-th layer diffusion network for processing, and the output of the second information stream in the i-1th layer diffusion network is input to the first latent layer spatial stream module in the i-th layer diffusion network for processing, This involves manipulating the processed features of the first condition information stream module in the i-th layer diffusion network and the processed features of the first latent layer spatial stream module in the i-th layer diffusion network at the element level, thereby concatenating the processed features of the first condition information stream module and the processed features of the first latent layer spatial stream module. The features after stitching together are processed by self-attention to obtain fused features, This includes inputting a feature from the fused features corresponding to the first information stream into the second condition information stream module in the i-th layer's spread network to obtain the output of the second condition information stream module, and inputting a feature from the fused features corresponding to the second information stream into the second latent layer spatial stream module in the i-th layer's spread network to obtain the output of the second latent layer spatial stream module.
[0190] In one specific example of the solution of this disclosure, the feature representation scheme for the image features of an entity image includes at least one of the positional features of the entity image, the segmentation features of a segmentation diagram corresponding to the entity image, the image features of a cutout diagram of the entity image, the image depth features of the entity image, and the local features of the entity image.
[0191] In one specific example of the solution disclosed herein, the training unit specifically, Using the image and text encoders in the preset model, one or more feature representation methods are selected from the above based on the sample features of the sample data to obtain the image features of the actual images contained in the sample data, thereby obtaining the initial multimodal features of the sample data. Alternatively, it can be used to randomly select one or more of the above feature representation methods to obtain image features of the actual images included in the sample data, thereby obtaining the initial multimodal features of the sample data.
[0192] The specific functions and illustrative descriptions of each unit of the apparatus according to the embodiments of this disclosure can be found in the relevant descriptions of the corresponding steps in the embodiments of the method described above, and will not be repeated here.
[0193] The solution of this disclosure also provides a multimodal data processing device, as shown in Figure 12, which is A reasoning unit 1201 is used to input multimodal prompt information into a pre-trained target model to obtain target multimodal features of the multimodal prompt information, wherein the multimodal prompt information includes at least a target text and at least one target entity image, the target text includes at least one target entity text, the target entity image is an image corresponding to the target entity text, the target multimodal features include image features of the target entity image and text features of the target entity text, the image features of the target entity image are after the text features of the target entity text corresponding to the target entity image, and feature diffusion is performed on the target multimodal features using an N-layer target diffusion network in the target model to obtain a target reasoning result including an image and / or text related to the multimodal prompt information, where N is a positive integer. The system includes an output unit 1202 for outputting the aforementioned target reasoning results.
[0194] In one specific example of the solution of the present disclosure, the reasoning unit is further used to insert target instruction information, which includes image features of target entity images that need adjustment and target adjustment instructions, into the target multimodal feature. For example, in one example, target instruction information is inserted into the target multimodal feature when at least some of the target entity images need image adjustment.
[0195] In one specific example of the solution disclosed herein, the reasoning unit specifically, It is used to concatenate the target instruction information to the header of the target feature sequence representing the target multimodal features.
[0196] In one specific example of the solution of the present disclosure, the N-layer target spread network is a dual information stream target spread network, the dual information stream target spread network includes a first target information stream for processing target multimodal features and a second target information stream for compressing the target multimodal features based on a preset target noise feature in order to represent the target multimodal features in hidden layer space.
[0197] In one specific example of the solution of this disclosure, when the value of N is 2 or greater, the N target spreading networks are connected in series, the corresponding output of the first target information stream in the i-th target spreading network among the N target spreading networks is the input of the first target information stream in the (i+1)-th target spreading network among the N target spreading networks, and the output of the second target information stream in the i-th spreading network among the N target spreading networks is the input of the second target information stream in the (i+1)-th target spreading network among the N target spreading networks.
[0198] In one specific example of the solution disclosed herein, the reasoning unit specifically, The method involves using the N-layer target spread network to perform feature diffusion, with the aforementioned target multimodal features as input to the first target information stream of the first layer of the target spread network in the N-layer target spread network, and the preset target noise features as input to the second target information stream of the first layer of the target spread network in the N-layer target spread network, thereby obtaining an initial reasoning result. The initial reasoning result is obtained based on the output of the second target information stream of the final layer of the target spread network in the N-layer target spread network, and represents the predicted result in the hidden layer space of the predicted target generation task. If the initial reasoning result includes image features, the image features in the initial reasoning result are decoded using the image decoder in the target model to obtain the target reasoning result.
[0199] In one specific example of the solution of this disclosure, the dual information stream target spreading network includes a first target condition information stream module and a second target condition information stream module corresponding to a first target information stream, and a first target latent layer spatial stream module and a second target latent layer spatial stream module corresponding to the second target information stream, Here, the input to the first target information stream is the input to the first target condition information stream module, and the output to the second target condition information stream module is the output to the first target information stream. The input to the second target information stream is the input to the first target latent layer spatial stream module, and the output to the second target latent layer spatial stream module is the output to the second target information stream.
[0200] In one specific example of the solution disclosed herein, the reasoning unit specifically, The following steps are used to perform feature diffusion using the i-th layer of the N-layer target diffusion network: The aforementioned step is, The output of the first target information stream in the i-1th layer target diffusion network is input to the first target condition information stream module in the i-th layer target diffusion network for processing, and the output of the second target information stream in the i-1th layer target diffusion network is input to the first target latent layer spatial stream module in the i-th layer target diffusion network for processing, This involves manipulating the processed features of the first target condition information stream module in the i-th layer's target diffusion network and the processed features of the first target latent layer spatial stream module in the i-th layer's target diffusion network at the element level, thereby concatenating the processed features of the first target condition information stream module and the processed features of the first target latent layer spatial stream module. The features after stitching together are processed by self-attention to obtain the target fused features. This includes inputting the feature corresponding to the first target information stream from among the target fusion features into the second target condition information stream module in the i-th layer of the target diffusion network to obtain the output of the second target condition information stream module, and inputting the feature corresponding to the second target information stream from among the target fusion features into the second target latent layer spatial stream module in the i-th layer of the target diffusion network to obtain the output of the second target latent layer spatial stream module.
[0201] In one specific example of the solution disclosed herein, the reasoning unit specifically, The image and text encoder in the target model is used to obtain the target multimodal features by selecting one or more feature representation schemes from among those including positional features of the target entity image, segmentation features of the segmentation diagram corresponding to the target entity image, image features of the cutout diagram of the target entity image, image depth features of the target entity image, and local features of the target entity image, based on the data features of the multimodal prompt information.
[0202] The specific functions and illustrative descriptions of each unit of the apparatus according to the embodiments of this disclosure can be found in the relevant descriptions of the corresponding steps in the embodiments of the method described above, and will not be repeated here.
[0203] The technical solutions disclosed herein, in which users' personal information is collected, stored, and applied, comply with the provisions of applicable laws and regulations and do not violate public order and morality.
[0204] According to embodiments of the present disclosure, the present disclosure further provides electronic devices, non-temporary computer-readable storage media, and program products.
[0205] Figure 13 is a block diagram of an electronic device 1300 for implementing an embodiment of the present disclosure. The electronic device refers to various types of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other compatible computers. The electronic device further refers to various types of mobile devices, such as personal digital assistants, cellular phones, intelligent phones, wearable devices, and other similar computer devices. The components, their connections, and functions described in this disclosure are illustrative and do not limit the implementation of anything described or specified in this disclosure.
[0206] As shown in Figure 13, device 1300 includes a computing unit 1301 that can perform various appropriate operations and processes based on computer program instructions stored in read-only memory (ROM) 1302 or computer program instructions loaded from storage unit 1308 into random access memory (RAM) 1303. RAM 1303 can further store various programs and data necessary for the operation of device 1300. The computing unit 1301, ROM 1302, and RAM 1303 are connected to each other via bus 1304. An input / output (I / O) interface 1305 is also connected to bus 1304.
[0207] Multiple components in device 1300 are connected to an I / O interface 1305, which includes an input unit 1306 such as a keyboard and mouse, an output unit 1307 such as various displays and speakers, a storage unit 1308 such as a magnetic disk or optical disk, and a communication unit 1309 such as a network card, modem, or wireless communication transceiver. The communication unit 1309 allows device 1300 to exchange information / data with other devices via computer networks such as the Internet and / or various carrier networks.
[0208] The computing unit 1301 may be a variety of general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, a computing unit that executes various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 performs each of the methods and processes described above, such as a model training method or a multimodal data processing method. For example, in some embodiments, the model training method or the multimodal data processing method can be implemented as a computer software program tangibly contained in a machine-readable medium such as a storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed into device 1300 via ROM 1302 and / or communication unit 1309. When the computer program is loaded into RAM 1303 and executed by the computing unit 1301, one or more steps of the model training method or multimodal data processing method described above can be performed. In addition, in other embodiments, the computing unit 1301 may be configured to perform a model training method or a multimodal data processing method by any other suitable method (e.g., firmware).
[0209] Various embodiments of the systems or technologies described in this disclosure can be implemented by digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standards (ASSPs), systems-on-a-chip (SOCs), complex-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. Each of these embodiments may be implemented by one or more computer programs that run and / or interpret on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0210] Program code for performing the methods of this disclosure can be written in any combination of one or more programming languages. These program codes are provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programming data processing device, so that when the program code is executed by the processor or controller, it can perform the functions / operations defined in the flowcharts and / or block diagrams. The program code may run entirely in a mainscan, partially in a mainscan, partially as an independent soft encapsulation and partially in a remote mainscan, or entirely in a remote mainscan or server.
[0211] In this disclosure, machine-readable media may be tangible media containing or storing programs used by or in conjunction with instruction execution systems, devices, or equipment. Machine-readable media may be machine-readable signal media or machine-readable storage media. Machine-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any suitable combination of the foregoing. Further specific examples of machine-readable storage media include electrical connections by one or more wires, portable computer disk cartridges, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any combination of the foregoing.
[0212] To provide user interaction, a computer may implement the systems and technologies described herein, which may include a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor), a keyboard and pointing device for the user to provide input to the computer (e.g., a mouse or trackball). Other types of devices may also be used to provide user interaction; for example, the feedback provided to the user may be any form of sensor feedback (e.g., visual feedback, auditory feedback, or haptic feedback), and input from the user may be accepted in any form (e.g., acoustic input, voice input, haptic input).
[0213] The systems and technologies described herein can be implemented in computing systems that include background components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include front-end components (e.g., user computers having a graphics user interface or network browser, through which users can interact with embodiments of the systems and technologies described herein), or in any combination of such background components, middleware components, or front-end components. Components of the system can be connected to one another via digital data communication in any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0214] A computer system can include a client and a server. Typically, the client and server are geographically separated and interact via a communication network. The client-server relationship is created by a computer program that operates on the corresponding computer. The server may be a cloud server, a server in a distributed system, or a server incorporating blockchain technology, etc.
[0215] It should be understood that steps can be newly ranked, added, or deleted using the various forms of flows shown above. For example, each step described in this disclosure may be executed in parallel, sequentially, or in a different order. This disclosure is not limited to this, as long as the technical solutions disclosed herein can achieve the desired results.
[0216] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations, and substitutions are possible due to design considerations and other factors. Any changes, equivalent substitutions, and improvements within the gist and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A computer equipped with a processor The method involves inputting sample data into a preset model to obtain initial multimodal features of the sample data, wherein the sample data includes at least sample text and at least one entity image, the sample text includes at least one entity text, the entity image is an image corresponding to the entity text, and the initial multimodal features include image features of the entity image and text features of the entity text, with the image features of the entity image following the text features of the entity text corresponding to the entity image. If the parameters of the image and text encoders of the preset model are fixed, the method involves training the N-layer spreading network in the preset model using the initial multimodal features and preset noise features to obtain the target model, where N is a positive integer. Training methods for the target model.
2. The training method for the aforementioned target model is: This further includes inserting instruction information, which includes image features of the entity image requiring adjustment and adjustment instructions, into the initial multimodal features. A method for training the target model described in claim 1.
3. Inserting the aforementioned instruction information into the initial multimodal feature means This includes concatenating the instruction information to the header of a feature sequence representing the initial multimodal features. A method for training the target model according to claim 2.
4. Each of the aforementioned N-layer spread network is a dual information stream spread network, and each dual information stream spread network includes a first information stream for processing initial multimodal features and a second information stream for compressing the initial multimodal features based on preset noise features in order to represent the initial multimodal features in hidden layer space. A method for training the target model described in claim 1.
5. If the value of N is 2 or greater, the N spreading networks are connected in series, the corresponding output of the first information stream in the i-th spreading network among the N spreading networks is the input of the first information stream in the (i+1)-th spreading network among the N spreading networks, and the output of the second information stream in the i-th spreading network among the N spreading networks is the input of the second information stream in the (i+1)-th spreading network among the N spreading networks. A method for training the target model described in claim 4.
6. The training method for the aforementioned target model is: The method further includes performing a masking operation on some of the initial multimodal features to obtain masked initial multimodal features. A method for training the target model described in claim 4.
7. Using the initial multimodal features and preset noise features, model training is performed on the N-layer spread network in the preset model. The method involves performing feature diffusion using an N-layer diffusion network, with masked initial multimodal features as input to the first information stream of the diffusion network in the first layer of the N-layer diffusion network, and preset noise features as input to the second information stream of the diffusion network in the first layer of the N-layer diffusion network, to obtain a total prediction result, wherein the total prediction result includes: a first prediction result, which is the output of the first information stream of the diffusion network in the final layer of the N-layer diffusion network and represents at least the predicted mask position corresponding to the predicted mask operation; and a second prediction result, which is the output of the second information stream of the diffusion network in the final layer of the N-layer diffusion network and represents the prediction result in the hidden layer space of the generation task indicated by the predicted sample data. The process involves obtaining the loss value of the target loss function based on the first and second prediction results, where the target loss function is used to calculate the difference between the predicted mask position and the actual mask position, and is used to calculate the difference between the predicted result in the hidden layer space of the generation task shown by the sample data and the actual result in the hidden layer space. This includes adjusting at least some of the tunable network parameters in the N-layer spreading network based on the loss value of the target loss function, A method for training the target model described in claim 6.
8. The dual information stream spreading network includes a first conditional information stream module and a second conditional information stream module corresponding to a first information stream, and a first latent layer spatial stream module and a second latent layer spatial stream module corresponding to a second information stream, Here, the input to the first information stream is the input to the first conditional information stream module, and the output to the second conditional information stream module is the output to the first information stream. The input to the second information stream is the input to the first latent layer space stream module, and the output to the second latent layer space stream module is the output to the second information stream. A method for training the target model described in claim 7.
9. Performing feature diffusion using the aforementioned N-layer diffusion network is, The following steps include performing feature diffusion using the i-th layer of the N-layer diffusion network: The aforementioned step is, The output of the first information stream in the i-1th layer diffusion network is input to the first conditional information stream module in the i-th layer diffusion network for processing, and the output of the second information stream in the i-1th layer diffusion network is input to the first latent layer spatial stream module in the i-th layer diffusion network for processing, The processed features of the first condition information stream module in the i-th layer diffusion network and the processed features of the first latent layer spatial stream module in the i-th layer diffusion network are manipulated at the element level to connect the processed features of the first condition information stream module and the processed features of the first latent layer spatial stream module. The features after stitching together are processed by self-attention to obtain fused features, This includes inputting the features corresponding to the first information stream from among the fused features into the second condition information stream module in the i-th layer's spread network to obtain the output of the second condition information stream module, and inputting the features corresponding to the second information stream from among the fused features into the second latent layer spatial stream module in the i-th layer's spread network to obtain the output of the second latent layer spatial stream module. A method for training the target model according to claim 8.
10. The feature representation method for the image features of a real image includes at least one of the following: the positional features of the real image, the segmentation features of the segmentation diagram corresponding to the real image, the image features of the cropped diagram of the real image, the image depth features of the real image, and the local features of the real image. A method for training the target model described in claim 1.
11. Obtaining the initial multimodal features of the aforementioned sample data is possible. Using the image and text encoders in the preset model, one or more feature representation methods are selected from the aforementioned methods based on the sample features of the sample data to obtain the image features of the actual images contained in the sample data, thereby obtaining the initial multimodal features of the sample data. Alternatively, this may include randomly selecting one or more feature representation methods from the aforementioned methods to obtain image features of the actual images included in the sample data, thereby obtaining the initial multimodal features of the sample data. A method for training a target model according to claim 10.
12. A computer equipped with a processor The method involves inputting multimodal prompt information into a pre-trained target model to obtain target multimodal features of the multimodal prompt information, wherein the multimodal prompt information includes at least a target text and at least one target entity image, the target text includes at least one target entity text, the target entity image is an image corresponding to the target entity text, and the target multimodal features include image features of the target entity image and text features of the target entity text, with the image features of the target entity image following the text features of the target entity text corresponding to the target entity image. Using the N-layer target diffusion network in the target model, feature diffusion is performed on the target multimodal features to obtain a target inference result, including images and / or text, related to the multimodal prompt information, where N is a positive integer. A method for processing multimodal data.
13. The aforementioned multimodal data processing method is: This further includes inserting target instruction information, which includes image features of the target entity image requiring adjustment and target adjustment instructions, into the target multimodal feature. The multimodal data processing method according to claim 12.
14. Inserting the aforementioned target instruction information into the target multimodal feature means This includes concatenating the target instruction information to the header of the target feature sequence representing the target multimodal feature, The multimodal data processing method according to claim 13.
15. The N-layer target spread network is a dual information stream target spread network, and the dual information stream target spread network includes a first target information stream for processing target multimodal features and a second target information stream for compressing target multimodal features based on preset target noise features in order to represent the target multimodal features in hidden layer space. The multimodal data processing method according to claim 12.
16. If the value of N is 2 or greater, the N target spreading networks are connected in series, the corresponding output of the first target information stream in the i-th target spreading network among the N target spreading networks is the input of the first target information stream in the (i+1)-th target spreading network among the N target spreading networks, and the output of the second target information stream in the i-th spreading network among the N target spreading networks is the input of the second target information stream in the (i+1)-th target spreading network among the N target spreading networks. The multimodal data processing method according to claim 15.
17. Using the N-layer target diffusion network in the aforementioned target model, feature diffusion is performed on the target multimodal features to obtain target inference results, including images and / or text, related to the multimodal prompt information. The method involves using the N-layer target spread network to perform feature diffusion, with the aforementioned target multimodal features as input to the first target information stream of the first layer of the target spread network in the N-layer target spread network, and the preset target noise features as input to the second target information stream of the first layer of the target spread network in the N-layer target spread network, thereby obtaining an initial reasoning result. The initial reasoning result is obtained based on the output of the second target information stream of the final layer of the target spread network in the N-layer target spread network, and represents the prediction result in the hidden layer space of the target generation task indicated in the multimodal prompt information. If the initial reasoning result includes image features, the image features in the initial reasoning result are decoded using the image decoder in the target model to obtain the target reasoning result, and this includes: The multimodal data processing method according to claim 15.
18. The dual information stream target diffusion network includes a first target condition information stream module and a second target condition information stream module corresponding to a first target information stream, and a first target latent layer spatial stream module and a second target latent layer spatial stream module corresponding to the second target information stream, Here, the input to the first target information stream is the input to the first target condition information stream module, and the output to the second target condition information stream module is the output to the first target information stream. The input to the second target information stream is the input to the first target latent layer spatial stream module, and the output to the second target latent layer spatial stream module is the output to the second target information stream. The multimodal data processing method according to claim 17.
19. Performing feature diffusion using the aforementioned N-layer target diffusion network is, The following steps include performing feature diffusion using the target diffusion network of the i-th layer of an N-layer target diffusion network: The aforementioned step is, The output of the first target information stream in the i-1th layer target diffusion network is input to the first target condition information stream module in the i-th layer target diffusion network for processing, and the output of the second target information stream in the i-1th layer target diffusion network is input to the first target latent layer spatial stream module in the i-th layer target diffusion network for processing, The processed features of the first target condition information stream module in the i-th layer of the target diffusion network and the processed features of the first target latent layer spatial stream module in the i-th layer of the target diffusion network are manipulated at the element level to connect the processed features of the first target condition information stream module and the processed features of the first target latent layer spatial stream module. The features after stitching together are processed by self-attention to obtain the target fused features. This includes inputting the feature corresponding to the first target information stream from among the target fusion features into the second target condition information stream module in the i-th layer of the target diffusion network to obtain the output of the second target condition information stream module, and inputting the feature corresponding to the second target information stream from among the target fusion features into the second target latent layer spatial stream module in the i-th layer of the target diffusion network to obtain the output of the second target latent layer spatial stream module. The multimodal data processing method according to claim 18.
20. Obtaining the target multimodal features of the aforementioned multimodal prompt information is: This includes obtaining the target multimodal features by using the image and text encoder in the target model, selecting one or more feature representation schemes from among those including the positional features of the target entity image, the segmentation features of the segmentation diagram corresponding to the target entity image, the image features of the cropped diagram of the target entity image, the image depth features of the target entity image, and the local features of the target entity image, based on the data features of the multimodal prompt information, thereby obtaining the target multimodal features. The multimodal data processing method according to claim 13.
21. A training device for the target model, A training unit is used to input sample data into a preset model to obtain initial multimodal features of the sample data, wherein the sample data includes at least sample text and at least one real image, the sample text includes at least one real text, the real image is an image corresponding to the real text, the initial multimodal features include image features of the real image and text features of the real text, the image features of the real image are after the text features of the real text corresponding to the real image, and, if the parameters of the image / text encoder of the preset model are fixed, to perform model training on an N-layer spread network in the preset model using the initial multimodal features and preset noise features to obtain a target model, where N is a positive integer. A memory unit for outputting the aforementioned target model is provided. A training device for the target model.
22. The aforementioned training unit is It is further used to insert instruction information, including image features of the entity image requiring adjustment and adjustment instructions, into the initial multimodal features. A training device for a target model according to claim 21.
23. The aforementioned training unit specifically, This is used to concatenate the instruction information to the header of the feature sequence representing the initial multimodal features. A training device for a target model according to claim 22.
24. Each of the aforementioned N-layer spread network is a dual information stream spread network, and each dual information stream spread network includes a first information stream for processing initial multimodal features and a second information stream for compressing the initial multimodal features based on preset noise features in order to represent the initial multimodal features in hidden layer space. A training device for a target model according to any one of claims 21 to 23.
25. If the value of N is 2 or greater, the N spreading networks are connected in series, the corresponding output of the first information stream in the i-th spreading network among the N spreading networks is the input of the first information stream in the (i+1)-th spreading network among the N spreading networks, and the output of the second information stream in the i-th spreading network among the N spreading networks is the input of the second information stream in the (i+1)-th spreading network among the N spreading networks. A training device for a target model according to claim 24.
26. The aforementioned training unit is This is further used to obtain masked initial multimodal features by performing a masking operation on some of the features among the aforementioned initial multimodal features. A training device for a target model according to claim 24.
27. The aforementioned training unit specifically, The method involves performing feature diffusion using an N-layer diffusion network, with masked initial multimodal features as input to the first information stream of the diffusion network in the first layer of the N-layer diffusion network, and preset noise features as input to the second information stream of the diffusion network in the first layer of the N-layer diffusion network, to obtain a total prediction result, wherein the total prediction result includes: a first prediction result, which is the output of the first information stream of the diffusion network in the final layer of the N-layer diffusion network and represents at least the predicted mask position corresponding to the predicted mask operation; and a second prediction result, which is the output of the second information stream of the diffusion network in the final layer of the N-layer diffusion network and represents the prediction result in the hidden layer space of the generation task indicated by the predicted sample data. The process involves obtaining the loss value of the target loss function based on the first and second prediction results, where the target loss function is used to calculate the difference between the predicted mask position and the actual mask position, and is used to calculate the difference between the predicted result in the hidden layer space of the generation task shown by the sample data and the actual result in the hidden layer space. Used to adjust at least some of the tunable network parameters in the N-layer spread spectrum network based on the loss value of the target loss function, A training device for a target model according to claim 26.
28. The dual information stream spreading network includes a first conditional information stream module and a second conditional information stream module corresponding to a first information stream, and a first latent layer spatial stream module and a second latent layer spatial stream module corresponding to a second information stream, Here, the input to the first information stream is the input to the first conditional information stream module, and the output to the second conditional information stream module is the output to the first information stream. The input to the second information stream is the input to the first latent layer space stream module, and the output to the second latent layer space stream module is the output to the second information stream. A training device for a target model according to claim 27.
29. The aforementioned training unit specifically, The following steps are used to perform feature diffusion using the i-th layer of the N-layer diffusion network: The aforementioned step is, The output of the first information stream in the i-1th layer diffusion network is input to the first conditional information stream module in the i-th layer diffusion network for processing, and the output of the second information stream in the i-1th layer diffusion network is input to the first latent layer spatial stream module in the i-th layer diffusion network for processing, The processed features of the first condition information stream module in the i-th layer diffusion network and the processed features of the first latent layer spatial stream module in the i-th layer diffusion network are manipulated at the element level to connect the processed features of the first condition information stream module and the processed features of the first latent layer spatial stream module. The features after stitching together are processed by self-attention to obtain fused features, This includes inputting the features corresponding to the first information stream from among the fused features into the second condition information stream module in the i-th layer's spread network to obtain the output of the second condition information stream module, and inputting the features corresponding to the second information stream from among the fused features into the second latent layer spatial stream module in the i-th layer's spread network to obtain the output of the second latent layer spatial stream module. A training device for a target model according to claim 28.
30. The feature representation method for the image features of a real image includes at least one of the following: the positional features of the real image, the segmentation features of the segmentation diagram corresponding to the real image, the image features of the cropped diagram of the real image, the image depth features of the real image, and the local features of the real image. A training device for a target model according to claim 21.
31. The aforementioned training unit specifically, Using the image and text encoders in the preset model, one or more feature representation methods are selected from the above based on the sample features of the sample data to obtain the image features of the actual images contained in the sample data, thereby obtaining the initial multimodal features of the sample data. Alternatively, it can be used to randomly select one or more of the above feature representation methods to obtain image features of the actual images included in the sample data, and to obtain the initial multimodal features of the sample data. A training device for a target model according to claim 30.
32. A multimodal data processing device, A reasoning unit used in the following: inputting multimodal prompt information into a pre-trained target model to obtain target multimodal features of the multimodal prompt information, wherein the multimodal prompt information includes at least a target text and at least one target entity image, the target text includes at least one target entity text, the target entity image is an image corresponding to the target entity text, the target multimodal features include image features of the target entity image and text features of the target entity text, the image features of the target entity image are after the text features of the target entity text corresponding to the target entity image, and performing feature diffusion on the target multimodal features using an N-layer target diffusion network in the target model to obtain a target reasoning result including an image and / or text related to the multimodal prompt information, where N is a positive integer; The system comprises an output unit for outputting the aforementioned target reasoning result, Multimodal data processing device.
33. The aforementioned reasoning unit, It is further used to insert target instruction information, which includes image features of the target entity image requiring adjustment and target adjustment instructions, into the target multimodal features. The multimodal data processing device according to claim 32.
34. The aforementioned reasoning unit specifically, This is used to concatenate the target instruction information to the header of the target feature sequence representing the target multimodal feature. The multimodal data processing device according to claim 33.
35. The N-layer target spread network is a dual information stream target spread network, and the dual information stream target spread network includes a first target information stream for processing target multimodal features and a second target information stream for compressing target multimodal features based on preset target noise features in order to represent the target multimodal features in hidden layer space. A multimodal data processing device according to any one of claims 32 to 34.
36. If the value of N is 2 or greater, the N target spreading networks are connected in series, the corresponding output of the first target information stream in the i-th target spreading network among the N target spreading networks is the input of the first target information stream in the (i+1)-th target spreading network among the N target spreading networks, and the output of the second target information stream in the i-th spreading network among the N target spreading networks is the input of the second target information stream in the (i+1)-th target spreading network among the N target spreading networks. The multimodal data processing device according to claim 35.
37. The aforementioned reasoning unit specifically, The method involves using the N-layer target spread network to perform feature diffusion, with the aforementioned target multimodal features as input to the first target information stream of the first layer of the target spread network in the N-layer target spread network, and the preset target noise features as input to the second target information stream of the first layer of the target spread network in the N-layer target spread network, thereby obtaining an initial reasoning result. The initial reasoning result is obtained based on the output of the second target information stream of the final layer of the target spread network in the N-layer target spread network, and represents the prediction result in the hidden layer space of the target generation task indicated in the multimodal prompt information. If the initial reasoning result includes image features, the image features in the initial reasoning result are decoded using the image decoder in the target model to obtain the target reasoning result, and this is used in the following: The multimodal data processing device according to claim 35.
38. The dual information stream target diffusion network includes a first target condition information stream module and a second target condition information stream module corresponding to a first target information stream, and a first target latent layer spatial stream module and a second target latent layer spatial stream module corresponding to the second target information stream, Here, the input to the first target information stream is the input to the first target condition information stream module, and the output to the second target condition information stream module is the output to the first target information stream. The input to the second target information stream is the input to the first target latent layer spatial stream module, and the output to the second target latent layer spatial stream module is the output to the second target information stream. The multimodal data processing device according to claim 37.
39. The aforementioned reasoning unit specifically, The following steps are used to perform feature diffusion using the i-th layer of the N-layer target diffusion network: The aforementioned step is, The output of the first target information stream in the i-1th layer target diffusion network is input to the first target condition information stream module in the i-th layer target diffusion network for processing, and the output of the second target information stream in the i-1th layer target diffusion network is input to the first target latent layer spatial stream module in the i-th layer target diffusion network for processing, The processed features of the first target condition information stream module in the i-th layer of the target diffusion network and the processed features of the first target latent layer spatial stream module in the i-th layer of the target diffusion network are manipulated at the element level to connect the processed features of the first target condition information stream module and the processed features of the first target latent layer spatial stream module. The features after stitching together are processed by self-attention to obtain the target fused features. This includes inputting the feature corresponding to the first target information stream from among the target fusion features into the second target condition information stream module in the i-th layer of the target diffusion network to obtain the output of the second target condition information stream module, and inputting the feature corresponding to the second target information stream from among the target fusion features into the second target latent layer spatial stream module in the i-th layer of the target diffusion network to obtain the output of the second target latent layer spatial stream module. The multimodal data processing device according to claim 38.
40. The aforementioned reasoning unit specifically, Using the image and text encoders in the target model, one or more feature representation schemes are selected from among those including positional features of the target entity image, segmentation features of the segmentation diagram corresponding to the target entity image, image features of the cropped diagram of the target entity image, image depth features of the target entity image, and local features of the target entity image, based on the data features of the multimodal prompt information, to obtain the image features of the target entity image included in the multimodal prompt information, thereby obtaining the target multimodal features. The multimodal data processing device according to claim 33.
41. At least one processor, The system comprises at least one processor and a memory that is communicated with by it, The memory stores instructions that can be executed by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is caused to perform the method according to any one of claims 1 to 20. Electronic devices.
42. A non-temporary computer-readable storage medium storing computer instructions that cause a computer to perform the method described in any one of claims 1 to 20.
43. A program for implementing the method described in any one of claims 1 to 20, which is executed by a processor in a computer.
Citation Information
Patent Citations
Multimedia event information detection method, device, server and storage medium
CN111858973B
Multi-modal data classification method, terminal equipment and storage medium
CN117421639A
Multi-modal pre-training model training method and device and storage medium
CN117875395A
Image processing method and device based on multi-modal information
CN117975211A