Modal Information Generation Method, Apparatus, Electronic Device, and Storage Medium

By using a binary modal transformation model to convert multimodal reference information into feature information under the target mode and fusion, the dependence problem of multimodal generation model on massive data is solved, and the optimization of multimodal content generation is achieved.

CN118113886BActive Publication Date: 2025-05-30SHUXING TECH (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311650381.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-05-30
Estimated Expiration
2043-12-04

AI Technical Summary

Technical Problem

The existing multimodal generation model requires a large amount of paired data or triple data during the training stage, which makes it difficult to obtain large amounts of triple or even quadruple data in actual scenarios, limiting the application of multimodal content generation.

Method used

By obtaining multimodal reference information, a binary modal transformation model is used for each modal reference information to convert it into feature information under the target mode, and these feature information are fused to generate multimodal fusion feature information, and finally the content information under the expected mode is generated.

Benefits of technology

It realizes multimodal content generation without massive triple or quadruple data, reduces the harsh conditions for content generation, and optimizes the application of multimodal content generation scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118113886B_ABST
    Figure CN118113886B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a method, apparatus, electronic device, and storage medium for generating modal information. The method includes: obtaining multi-modal reference information, where the multi-modal reference information includes modal reference information of at least two original modalities; for each modal reference information, processing the modal reference information according to a binary modal conversion model corresponding to the original modality and the target modality to obtain modal conversion feature information in the target modality; fusing the modal conversion feature information to obtain multi-modal fusion feature information of the multi-modal reference information in the target modality; and generating modal content information in the target modality according to the multi-modal fusion feature information. Each binary modal conversion model is converted to the target modality, realizing the use of multiple binary modal conversion models to replace a single multi-modal content generation model, reducing the problem of harsh content generation conditions, and realizing the optimization of the multi-modal content generation scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and particularly to a method, apparatus, electronic device, and storage medium for generating modal information, where the storage medium includes a computer-readable storage medium. Background Art

[0002] In the current social scenario, artificial intelligence content creation (AIGC) mainly focuses on images and videos. That is, users upload text content and picture content, and through a generation model, generate high-quality content results (such as pictures) that conform to the picture content and text description.

[0003] However, content generation schemes based on a single modality or a small number of modalities have problems with poor content generation accuracy. With the increase in the types of information, content generation schemes based on more modalities have become the main strategy to solve content generation accuracy.

[0004] Currently, for multi-modal generation models, during the training stage, relevant modal data needs to be reserved in advance. For example, for a "text-image" generation model, "text-image" paired data is required, and for other models that support multi-modal generation, such as a "text-audio-image" generation model, triple data in the "text-audio-image" format is required.

[0005] However, in actual scenarios, binary data is generally easy to obtain in large quantities, such as image + text, video + audio, audio + text. However, triple or even quadruple data often cannot be obtained in large quantities due to the lack of a certain modal data, which greatly limits the application of content generation schemes based on multi-modalities. Summary of the Invention

[0006] Embodiments of the present application provide a method, apparatus, electronic device, and storage medium for generating modal information, which can optimize content generation schemes based on multi-modalities and facilitate the application of content generation schemes based on multi-modalities.

[0007] Embodiments of the present application provide a method for generating modal information, including:

[0008] Obtain multi-modal reference information, where the multi-modal reference information includes modal reference information of at least two original modalities;

[0009] For each modal reference information, process the modal reference information according to the binary modal conversion model corresponding to the original modality and the target modality to obtain modal conversion feature information in the target modality;

[0010] Fuse the modal conversion feature information to obtain multi-modal fusion feature information corresponding to the multi-modal reference information in the target modality;

[0011] Generate modal content information in the desired modality according to the multi-modal fusion feature information.

[0012] Correspondingly, an embodiment of the present application further provides a modal information generation device, including:

[0013] An information acquisition module, configured to acquire multi-modal reference information, where the multi-modal reference information includes modal reference information of at least two original modalities;

[0014] A feature generation module, configured to process each modal reference information according to a binary modal conversion model corresponding to the original modality and the target modality, to obtain modal conversion feature information in the target modality;

[0015] A feature fusion module, configured to fuse the modal conversion feature information to obtain multi-modal fusion feature information of the multi-modal reference information in the target modality;

[0016] An information generation module, configured to generate modal content information in the desired modality according to the multi-modal fusion feature information.

[0017] Optionally, in some embodiments of the present application, the information generation module includes:

[0018] A determination unit, configured to acquire a modal information generation model corresponding to the target modality and the desired modality;

[0019] A generation unit, configured to generate modal content information corresponding to the multi-modal fusion feature information through the modal information generation model.

[0020] Optionally, in some embodiments of the present application, the desired modality includes the target modality, and the determination unit includes:

[0021] A first determination subunit, configured to set a target modality generation model corresponding to the target modality as a modal information generation model, where the target modality generation model is used to convert the multi-modal fusion feature information into modal content information of the target modality.

[0022] Optionally, in some embodiments of the present application, the target modality is used as the first modality, and the desired modality includes generating a second modality other than the first modality, and the determination unit includes:

[0023] A second determination subunit, configured to determine a transition modality generation model according to the first modality and the second modality, where the transition modality generation model is used to convert multi-modal fusion feature information of the first modality into modal content information of the second modality;

[0024] A third determination subunit, configured to set the transition mode generation model as a mode information generation model.

[0025] Wherein, in some embodiments of the present application, the information generation module further includes a scoring unit, and the scoring unit includes:

[0026] A scoring subunit, configured to perform an aesthetic score on the mode content information through an aesthetic scoring model to obtain an actual aesthetic score;

[0027] An optimization subunit, configured to, if the actual aesthetic score is less than an aesthetic score threshold, optimize the mode information generation model according to aesthetic features to obtain an optimized mode information generation model, where the optimized mode information generation model is used to regenerate new mode content information corresponding to the multi-modal fusion feature information.

[0028] Wherein, in some embodiments of the present application, the feature fusion module includes:

[0029] A feature fusion unit, configured to fuse the mode conversion feature information respectively corresponding to the mode reference information of the at least two original modes according to a fusion weight to obtain multi-modal fusion feature information of the multi-modal reference information in the target mode.

[0030] Wherein, in some embodiments of the present application, the device further includes a training module, and the training module includes:

[0031] An acquisition unit, configured to acquire at least two binary mode group sample information, where the binary mode group sample information is composed of sample mode reference information belonging to an original mode and sample mode target information belonging to a target mode;

[0032] A construction unit, configured to, for each original mode, construct an initial mode conversion model according to the original mode and the target mode, where the initial mode conversion model is used for conversion between the sample mode reference information of the original mode and the sample mode target information of the target mode;

[0033] A training unit, configured to perform training on each of the initial mode conversion models respectively and uniformly map them to the feature space corresponding to the target mode to obtain binary mode conversion models corresponding to the respective original modes.

[0034] In a third aspect, an embodiment of the present application further provides an electronic device, where the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the computer program is executed by the processor, the steps in the above mode information generation method are implemented.

[0035] Fourthly, an embodiment of the present application further provides a storage medium, which includes a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in the above-mentioned modality information generation method are implemented.

[0036] Fifthly, an embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners of the embodiment of the present application.

[0037] The embodiment of the present application obtains multi-modal reference information, where the multi-modal reference information includes modal reference information of at least two original modalities. For each piece of modal reference information, the modal reference information is processed according to a binary modality conversion model corresponding to the original modality and the target modality to obtain modality conversion feature information in the target modality. The modality conversion feature information is fused to obtain multi-modal fusion feature information corresponding to the multi-modal reference information in the target modality. According to the multi-modal fusion feature information, modality content information in the expected modality is generated. The modality feature of the modal reference information is converted and aligned from the original modality to the target modality through the binary modality conversion model. The modal reference information of each original modality is subjected to feature conversion processing through a plurality of binary modality conversion models respectively, and the obtained modality conversion feature information is fused, and modality content information is generated based on the fused multi-modal fusion feature information, so as to realize content generation based on multiple modalities. Among them, each binary modality conversion model is converted to the target modality, so as to realize the effect of using a plurality of binary modality conversion models to replace a single multi-modal content generation model for multi-modal content generation, without using a multi-modal content generation model with more stringent generation conditions or training conditions, reducing the problem of stringent content generation conditions, and realizing the optimization of the multi-modal content generation scheme. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 It is a schematic diagram of the scenario of the modality information generation method provided by the embodiment of the present application;

[0040] Figure 2It is a schematic flowchart of the modal information generation method provided by an embodiment of the present application;

[0041] Figure 3 It is a schematic flowchart of the image generation method provided by an embodiment of the present application;

[0042] Figure 4 It is a schematic structural diagram of the modal information generation device provided by an embodiment of the present application;

[0043] Figure 5 It is a schematic structural diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0044] Next, the technical solutions in the present application will be clearly and completely described in conjunction with the accompanying drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0045] An embodiment of the present application provides a modal information generation method, device, electronic device, and storage medium. Specifically, an embodiment of the present application provides a modal information generation device applicable to an electronic device, where the electronic device includes a terminal device or a server. The terminal device includes devices such as a laptop computer, a touch screen, a desktop computer, a mobile phone, or a personal computer (PC, Personal Computer). The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs, Content Delivery Networks), and big data and artificial intelligence platforms. The server can be directly or indirectly connected through wired or wireless communication methods.

[0046] Among them, in an embodiment of the present application, the modal information generation method can be executed by the terminal device, or by the server, or jointly executed by the terminal device and the server. For example, please refer to Figure 1 , Figure 1 It is a schematic diagram of the scenario where the terminal device and the server jointly execute the modal information generation method provided by an embodiment of the present application. Among them, the specific execution process of the terminal device and the server jointly executing the modal information generation method is as follows:

[0047] In response to an input operation, the terminal device 10 receives multimodal reference information, where the multimodal reference information includes modal reference information of multiple original modalities. Subsequently, the terminal device 10 sends the multimodal reference information to the server 11, and the server 11 generates modal content information based on the multimodal reference information. Specifically:

[0048] After receiving the multimodal reference information sent by the terminal device 10, for each modal reference information, the server 11 generates modal conversion feature information corresponding to the modal reference information in the target modality according to the binary modal conversion model of the original modality corresponding to the target modality. Subsequently, the modal conversion feature information corresponding to the multiple modal reference information is fused to obtain multimodal fusion feature information corresponding to the multimodal reference information in the target modality, and modal content information in the desired modality is generated based on the multimodal fusion feature information.

[0049] Next, the server 11 feeds back the modal content information to the terminal device 10, and the terminal device 10 displays the modal content information in the visualization interface of the terminal device 10.

[0050] Among them, the binary modal conversion model is obtained after mapping training based on the feature space corresponding to the target modality, that is, each binary modal conversion model is trained based on the same feature space corresponding to the target modality.

[0051] Among them, the modal information generation method provided in the embodiments of the present application involves machine learning in the field of artificial intelligence. The embodiments of the present application can optimize the content generation scheme based on multimodality, facilitating the application of the content generation scheme based on multimodality.

[0052] Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machine to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Among them, artificial intelligence software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0053] Among them, machine learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0054] For example, in the embodiments of the present application, each binary modality conversion model is trained through the feature space corresponding to the same target modality, achieving the effect of jointly training multiple binary modality conversion models using the same feature space. By inputting the modality reference information into the corresponding binary modality conversion model, the binary modality conversion model can extract features from the modality reference information and align the extracted modality conversion feature information to the target modality.

[0055] It can be understood that in the embodiments of the present application, the modality feature of the modality reference information is converted and aligned from the original modality to the target modality through the binary modality conversion model. Each binary modality conversion model performs feature conversion processing on the modality reference information of each original modality respectively, fuses the processed modality conversion feature information, and generates modality content information based on the fused multi-modal fusion feature information, realizing multi-modal content generation. Among them, each binary modality conversion model is converted to the target modality, achieving the effect of using multiple binary modality conversion models to replace a single multi-modal content generation model for multi-modal content generation, without using a multi-modal content generation model with more stringent generation conditions or training conditions, reducing the problem of stringent content generation conditions, and realizing the optimization of the multi-modal content generation solution.

[0056] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.

[0057] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of the modality information generation method provided by the embodiments of the present application. The specific process of this modality information generation method is as follows:

[0058] 101. Obtain multi-modal reference information, where the multi-modal reference information includes the modality reference information of at least two original modalities.

[0059] It can be understood that the multimodal reference information is the sum of the modal reference information of multiple original modalities. By receiving the modal reference information of each input original modality, the multimodal reference information is aggregated.

[0060] It should be noted that the original modality refers to the construction form or data type of the data. Different original modalities correspond to different data types. For example, in the embodiments of the present application, the original modality can be modalities such as text, image, audio, or video.

[0061] Among them, in the embodiments of the present application, the modal reference information is information used for reference to generate other content. For example, by inputting the original text or the original image to generate a new text or image. Correspondingly, the original text or the original image is the modal reference information.

[0062] For each original modality, a binary modality conversion model corresponding to the original modality is determined according to the target modality. The binary modality conversion model is obtained after mapping training based on the feature space corresponding to the target modality.

[0063] It can be understood that the binary modality conversion model is a conversion model or a generation model for binary modalities. For example, a generation model that generates an image from text, which is represented as a text-image model; or a generation model that generates an image from audio, which is represented as an audio-image model. In addition, the binary modality conversion model can also be a video-image model, an image-image model, an image-text model, or an image-audio model, etc.

[0064] Correspondingly, in the embodiments of the present application, the target modality can be any one of text, image, audio, or video. For example, in the embodiments of the present application, taking the target modality as an image as an example, correspondingly, the binary modality conversion model can be a text-image model, an image-image model, an audio-image model, or a video-image model. Correspondingly, the original modality is text, image, audio, or video.

[0065] Among them, after the target modality is selected, the binary modality conversion model for the target modality can be mapped and trained in the feature space of the target modality. And since each binary modality conversion model is trained based on the same feature space, the effect of joint training of each binary modality conversion model is achieved.

[0066] Correspondingly, since the binary modality conversion model is trained based on binary original modalities, only binary data groups need to be collected for individual training, and there is no need to collect ternary data groups or quaternary data groups, reducing the difficulty of data collection and realizing the optimization of multimodal content generation.

[0067] 102. For each modal reference information, process the modal reference information according to the binary modal conversion model of the original modality corresponding to the target modality to obtain the modal conversion feature information in the target modality.

[0068] It can be understood that by generating the modal conversion feature information through the binary modal conversion model for the target modality, the alignment of the feature information of the modal reference information with the features in the target modality is achieved.

[0069] 103. Fuse the modal conversion feature information to obtain the multi-modal fusion feature information of the multi-modal reference information in the target modality.

[0070] It can be understood that since the modal conversion feature information of each original modality is aligned with the features of the target modality, after fusing the modal conversion feature information corresponding to each original modality, the obtained multi-modal fusion feature information is still aligned with the features of the target modality.

[0071] Among them, by fusing the modal conversion feature information of each original modality to obtain the multi-modal fusion feature information, the feature representation of the multi-modal reference information in the target modality is realized by using the multi-modal fusion feature information, which is convenient for generating modal content according to the multi-modal fusion feature information.

[0072] 104. Generate the modal content information in the desired modality according to the multi-modal fusion feature information.

[0073] It can be understood that after obtaining the multi-modal fusion feature information representing the multi-modal reference information, the corresponding modal content information can be generated according to the multi-modal fusion feature information, realizing content generation based on multiple modalities.

[0074] It can be understood that in the embodiments of the present application, the modal features of the modal reference information are converted and aligned from the original modality to the target modality through the binary modal conversion model. The modal reference information of each original modality is respectively subjected to feature conversion processing through multiple binary modal conversion models, and the processed modal conversion feature information is fused, and the modal content information is generated based on the fused multi-modal fusion feature information, realizing content generation based on multiple modalities. Among them, each binary modal conversion model is converted to the target modality, achieving the effect of using multiple binary modal conversion models to replace a single multi-modal content generation model for multi-modal content generation, without using a multi-modal content generation model with more demanding generation conditions or training conditions, reducing the problem of demanding content generation conditions, and realizing the optimization of the multi-modal content generation scheme. Among them, taking the target modality as an intermediate transition, the fusion of the feature information of multiple modalities is realized, and then the modal content information for the desired modality is generated by using the target modality.

[0075] Correspondingly, in the embodiments of the present application, after obtaining the multi-modal fusion feature information, the generation of the modal content information can also be obtained by searching for content based on the multi-modal fusion feature information. For example, content similar to the multi-modal fusion feature information is searched as the modal content information. That is, generating the modal content information according to the multi-modal fusion feature information can be:

[0076] Obtain the pre-coding result of the candidate content in the target modality, where the pre-coding result is obtained by encoding the candidate content in advance;

[0077] Calculate the similarity between the multi-modal fusion feature information and the pre-coding result;

[0078] Use the candidate content with the highest similarity or a certain number of candidate contents with the top-ranked similarities as the modal content information.

[0079] Wherein, the pre-coding result is a representation of the feature information corresponding to the candidate content.

[0080] Wherein, through the matching of feature similarities, the search for candidate contents in the library corresponding to the target modality is realized, achieving the function of image search based on multiple modalities.

[0081] Correspondingly, for the generation of the modal content information in the desired modality, the multi-modal fusion feature information of the target modality can be converted based on the desired modality to obtain the desired feature information in the desired modality, and the modal content information in the desired modality is generated or matched based on the desired feature information.

[0082] Correspondingly, in the embodiments of the present application, after obtaining the multi-modal fusion feature information, the generation of the modal content information can also be realized according to the corresponding modal information generation model. Correspondingly, the modal information generation model can be selected and determined according to the content generation requirements. For example, the modal content information in the desired modality is generated based on the multi-modal fusion feature information of the target modality. That is, optionally, in some embodiments of the present application, the step of "generating the modal content information in the desired modality according to the multi-modal fusion feature information" includes:

[0083] Obtain the modal information generation model of the target modality corresponding to the desired modality;

[0084] Generate the modal content information corresponding to the multi-modal fusion feature information through the modal information generation model.

[0085] It can be understood that the content generation requirements can be flexibly adjusted or changed according to actual needs. It should be emphasized that the content generation requirements correspond to the modality ultimately corresponding to the modal content information. For example, the content generation requirements can be to generate content in the target modality, or to generate content in other modalities other than the target modality, such as content in the expected modality. That is, the expected modality can be the same as the target modality or other modalities other than the target modality. For example, if the target modality is the image modality, the expected modality can be the image modality, or the text modality or the video modality.

[0086] Among them, by selecting the corresponding modality information generation model according to the content generation requirements, the generation of the content information of the corresponding modality based on the multi-modal fusion feature information is realized, improving the diversity and flexibility of the generation of the modal content information.

[0087] For example, in the embodiment of the present application, if the expected modality includes the target modality, the step of "obtaining the modality information generation model corresponding to the target modality for the expected modality" includes:

[0088] Setting the target modality generation model corresponding to the target modality as the modality information generation model, where the target modality generation model is used to convert the multi-modal fusion feature information into the modality content information of the target modality.

[0089] Among them, since the content generation requirements correspond to the target modality, therefore, the generation of the corresponding modality content information can be directly realized by using the modality information generation model of the target modality. For example, if the target modality is an image, the multi-modal fusion feature information corresponding image can be directly generated through the image generation model.

[0090] Again, for example, in the embodiment of the present application, if the content generation requirement is to generate the modality content information of other modalities different from the target modality, then the generation model corresponding to the target modality and the other modality is selected to generate the modality content information of the other modality based on the multi-modal fusion feature information. That is, optionally, in some embodiments of the present application, the target modality is used as the first modality, and the expected modality includes generating a second modality other than the first modality. The step of "obtaining the modality information generation model corresponding to the target modality for the expected modality" includes:

[0091] Determining a transition modality generation model according to the first modality and the second modality, where the transition modality generation model is used to convert the multi-modal fusion feature information of the first modality into the modality content information of the second modality;

[0092] Setting the transition modality generation model as the modality information generation model.

[0093] It can be understood that if the content generation requirement is different from the first modality, the first modality can be regarded as a bridge for feature transformation to achieve feature alignment of the modality reference information of multiple original modalities. Subsequently, based on the generation models corresponding to the first modality and the second modality (transition modality generation model), multi-modal fusion feature information based on the first modality can be used to generate modality content information based on the second modality.

[0094] Correspondingly, in the embodiments of the present application, in order to improve the accuracy, effectiveness or visual effect of the modality content information, an aesthetic score is also given to the generated modality content information, and it is determined whether it is necessary to regenerate new modality content information based on the aesthetic score result. That is, optionally, in some embodiments of the present application, after the step of "generating the modality content information corresponding to the multi-modal fusion feature information through the modality information generation model", the method further includes:

[0095] Performing an aesthetic score on the modality content information through an aesthetic scoring model to obtain an actual aesthetic score;

[0096] If the actual aesthetic score is less than the aesthetic score threshold, the modality information generation model is optimized according to the aesthetic features to obtain an optimized modality information generation model, and the optimized modality information generation model is used to regenerate the new modality content information corresponding to the multi-modal fusion feature information.

[0097] For example, if the modality content information is an image, an aesthetic score can be given to the image, and when the aesthetic score threshold is not met, an image with a higher aesthetic score can be regenerated, so as to obtain higher-quality modality content information.

[0098] It can be understood that the aesthetic score threshold is a predefined value. The content generation requirement corresponds to different modalities, that is, the aesthetic scoring methods for the modality content information of different modalities can be different, and the aesthetic requirements are different, that is, the aesthetic score thresholds are also set differently.

[0099] Optionally, in the embodiments of the present application, when fusing multiple modality conversion feature information, fusion can also be performed based on the corresponding fusion weights to improve the accuracy of the fusion result. That is, optionally, in some embodiments of the present application, the step of "fusing each of the modality conversion feature information to obtain the multi-modal fusion feature information of the multi-modal reference information in the target modality" includes:

[0100] Fusing the modality conversion feature information respectively corresponding to the modality reference information of the at least two original modalities according to the fusion weights to obtain the multi-modal fusion feature information of the multi-modal reference information in the target modality.

[0101] It can be understood that in the embodiments of the present application, the number of original modalities included in the input multimodal reference information is not fixed, that is, it can be modality reference information including two original modalities, or modality reference information including three or four original modalities. Correspondingly, the fusion weight of each original modality is determined based on the number of original modalities.

[0102] Furthermore, the fusion weight corresponding to each original modality can also be configured according to actual needs. For example, by receiving the fusion weight configuration information input by the user, the fusion weight corresponding to each original modality is obtained.

[0103] Correspondingly, before applying the binary modality conversion model, training of the binary modality conversion model is also included. That is, optionally, in some embodiments of the present application, before the step of "for each original modality, determining a binary modality conversion model corresponding to the original modality according to the target modality", the method further includes:

[0104] Obtaining at least two pieces of binary modality group sample information, where the binary modality group sample information is composed of sample modality reference information belonging to the original modality and sample modality target information belonging to the target modality;

[0105] For each original modality, constructing an initial modality conversion model according to the original modality and the target modality, where the initial modality conversion model is used for the conversion between the sample modality reference information of the original modality and the sample modality target information of the target modality;

[0106] Training each of the initial modality conversion models separately and uniformly mapping them to the feature space corresponding to the target modality to obtain binary modality conversion models corresponding to each original modality.

[0107] It can be understood that by constructing an initial modality conversion model for the target modality and the original modality, the construction of a binary-based modality conversion model is realized, and each initial modality conversion model is trained based on the same feature space corresponding to the target modality, achieving the effect of joint training of each initial modality conversion model. Among them, the initial modality conversion model obtains a binary modality conversion model after training.

[0108] Among them, since the initial modality conversion model is a binary-based generation model, and the sample data required for training a binary-based generation model is also binary. Compared with ternary or quaternary data, the binary data in the embodiments of the present application has the advantage of being convenient to collect. By training through mapping to the same feature space instead of the traditional joint training of ternary or quaternary data, the embodiments of the present application can optimize the generation of multimodal content and facilitate the application of multimodal-based content generation schemes.

[0109] It can be understood that by using the target modality as a bridge for generating modal content, the alignment of features of multiple modalities is achieved, facilitating the subsequent generation of content between different modalities.

[0110] For example, the following takes the target modality as the image modality, and the original modalities include the text modality, image modality, audio modality, and video modality for illustration (some modalities may be missing in actual applications). Please refer to Figure 3 , Figure 3 is a schematic flowchart of the image generation method provided by the embodiments of the present application. Among them, the process of the image generation method includes:

[0111] 201. Receive the text, image, audio, and video input by the user;

[0112] 202. For the text and image, obtain the text features and image features of the text and image in the image modality through a text-image model;

[0113] Among them, the text-image model is also the text-to-image model.

[0114] 203. For the audio, obtain the audio features of the audio in the image modality through an audio-image model;

[0115] Among them, the audio-image model is also the audio-to-image model.

[0116] 204. For the video, obtain the video features of the video in the image modality through a video-image model;

[0117] Among them, the video-image model is also the video-to-image model.

[0118] 205. Fuse the text features, image features, audio features, and video features to obtain the multi-modal fusion feature information of the text, image, audio, and video in the image modality;

[0119] 206. Through an image generation model, obtain the target image corresponding to the multi-modal fusion feature information.

[0120] Among them, in the embodiments of the present application, the content generation requirement is to generate image content. Therefore, an image generation model is selected to generate the target image based on the multi-modal fusion feature information.

[0121] 207. Determine whether the target image meets the aesthetic standard. If not, return to step 206 for execution. If it meets, execute step 208;

[0122] 208. Output the target image.

[0123] Among them, in the embodiments of the present application, the text image model, the audio image model, and the video image model are all trained based on the feature space corresponding to the image.

[0124] Correspondingly, when fusing features, the fusion weights corresponding to each feature can also be configured, and the text features, image features, audio features, and video features are fused based on the fusion weights.

[0125] It can be understood that by converting text, images, audio, and video through the corresponding binary modality conversion models, the feature representations of text, images, audio, and video in the image modality are obtained, aligning the features of text, images, audio, and video with the image features, which facilitates the implementation of multi-modal content generation.

[0126] Among them, the binary modality conversion models are obtained through unified mapping and training based on the feature space of the image modality, so that when each binary modality conversion model is trained based on its corresponding binary data, the effect of joint training can also be achieved. And the separate training of each binary modality conversion model based on binary data avoids the acquisition of ternary data or quaternary data, reduces the implementation difficulty of the multi-modal content generation scheme, and facilitates the application of the multi-modal content generation scheme.

[0127] To facilitate better implementation of the modality information generation method of the present application, the present application also provides a modality information generation device based on the above modality information generation method. The meanings of the nouns are the same as those in the above modality information generation method, and the specific implementation details can refer to the descriptions in the method embodiments.

[0128] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the modality information generation device provided by the embodiments of the present application. The modality information generation device can be specifically as follows:

[0129] An information acquisition module 301, configured to acquire multi-modal reference information, where the multi-modal reference information includes modality reference information of at least two original modalities;

[0130] A feature generation module 302, configured to process each modality reference information according to the binary modality conversion model corresponding to the original modality to the target modality, to obtain modality conversion feature information in the target modality;

[0131] A feature fusion module 303, configured to fuse the modality conversion feature information to obtain multi-modal fusion feature information of the multi-modal reference information in the target modality;

[0132] An information generation module 304, configured to generate modality content information in the desired modality according to the multi-modal fusion feature information.

[0133] Optionally, in some embodiments of the present application, the information generation module 304 includes:

[0134] A determination unit, configured to obtain a modal information generation model corresponding to the desired modality for the target modality;

[0135] A generation unit, configured to generate modal content information corresponding to the multi-modal fusion feature information through the modal information generation model.

[0136] Optionally, in some embodiments of the present application, the desired modality includes the target modality, and the determination unit includes:

[0137] A first determination subunit, configured to set the target modality generation model corresponding to the target modality as the modal information generation model, where the target modality generation model is used to convert the multi-modal fusion feature information into modal content information of the target modality.

[0138] Optionally, in some embodiments of the present application, the target modality is used as the first modality, and the desired modality includes generating a second modality other than the first modality. The determination unit includes:

[0139] A second determination subunit, configured to determine a transition modality generation model according to the first modality and the second modality, where the transition modality generation model is used to convert the multi-modal fusion feature information of the first modality into modal content information of the second modality;

[0140] A third determination subunit, configured to set the transition modality generation model as the modal information generation model.

[0141] Wherein, in some embodiments of the present application, the information generation module further includes a scoring unit, and the scoring unit includes:

[0142] A scoring subunit, configured to perform an aesthetic score on the modal content information through an aesthetic scoring model to obtain an actual aesthetic score;

[0143] An optimization subunit, configured to, if the actual aesthetic score is less than the aesthetic score threshold, optimize the modal information generation model according to aesthetic features to obtain an optimized modal information generation model, where the optimized modal information generation model is used to regenerate new modal content information corresponding to the multi-modal fusion feature information.

[0144] Wherein, in some embodiments of the present application, the feature fusion module 303 includes:

[0145] A weight determination unit, configured to determine a fusion weight for each original modality according to the number of the original modalities;

[0146] A feature fusion unit, configured to fuse the modal conversion feature information corresponding to the modal reference information of the at least two original modalities according to the fusion weights, so as to obtain the multi-modal fusion feature information of the multi-modal reference information in the target modality.

[0147] Wherein, in some embodiments of the present application, the device further includes a training module, and the training module includes:

[0148] An acquisition unit, configured to acquire at least two binary modality group sample information, where the binary modality group sample information is composed of sample modal reference information belonging to the original modality and sample modal target information belonging to the target modality;

[0149] A construction unit, configured to construct an initial modal conversion model for each original modality according to the original modality and the target modality, where the initial modal conversion model is used for the conversion between the sample modal reference information of the original modality and the sample modal target information of the target modality;

[0150] A training unit, configured to train each of the initial modal conversion models respectively and map them uniformly to the feature space corresponding to the target modality, so as to obtain binary modality conversion models corresponding to each original modality.

[0151] In the embodiments of the present application, first, the information acquisition module 301 acquires multi-modal reference information, where the multi-modal reference information includes modal reference information of at least two original modalities. Subsequently, for each modal reference information, the feature generation module 302 processes the modal reference information according to the binary modality conversion model corresponding to the original modality and the target modality, so as to obtain the modal conversion feature information in the target modality. Then, the feature fusion module 303 fuses the modal conversion feature information to obtain the multi-modal fusion feature information of the multi-modal reference information in the target modality. Then, the information generation module 304 generates modal content information in the desired modality according to the multi-modal fusion feature information.

[0152] It can be understood that through the binary modality conversion model, the modal features of the modal reference information are converted and aligned from the original modality to the target modality. The modal reference information of each original modality is subjected to feature conversion processing through multiple binary modality conversion models respectively, and the processed modal conversion feature information is fused, and the modal content information is generated based on the fused multi-modal fusion feature information, so as to realize multi-modal content generation. Among them, each binary modality conversion model is converted to the target modality, so as to achieve the effect of using multiple binary modality conversion models to replace a single multi-modal content generation model for multi-modal content generation, without using a multi-modal content generation model with more stringent generation conditions or training conditions, reducing the problem of stringent content generation conditions, and realizing the optimization of the multi-modal content generation scheme.

[0153] In addition, the present application also provides an electronic device, such as Figure 5 shown, which shows a schematic structural diagram of the electronic device involved in the present application. Specifically:

[0154] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art can understand that Figure 5 the structure of the electronic device shown in does not constitute a limitation on the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0155] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401.

[0156] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0157] The electronic device further includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may further include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0158] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0159] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402, so as to implement the steps in any one of the modal information generation methods provided in the embodiments of the present application.

[0160] The embodiment of the present application obtains multi-modal reference information, which includes modal reference information of at least two original modes. For each modal reference information, the modal reference information is processed according to the binary modal conversion model corresponding to the original mode and the target mode to obtain modal conversion feature information in the target mode. The modal conversion feature information is fused to obtain multi-modal fusion feature information corresponding to the multi-modal reference information in the target mode. According to the multi-modal fusion feature information, modal content information in the desired mode is generated.

[0161] The modal feature of the modal reference information is converted and aligned from the original mode to the target mode through the binary modal conversion model. The modal reference information of each original mode is respectively subjected to feature conversion processing through a plurality of binary modal conversion models, and the processed modal conversion feature information is fused, and modal content information is generated based on the fused multi-modal fusion feature information, so as to realize content generation based on multiple modes. Among them, each binary modal conversion model performs conversion to the target mode, achieving the effect of using a plurality of binary modal conversion models to replace a single multi-modal content generation model for multi-modal content generation, without using a multi-modal content generation model with more stringent generation conditions or training conditions, reducing the problem of stringent content generation conditions, and realizing the optimization of the multi-modal content generation scheme.

[0162] For the specific implementation of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.

[0163] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer storage medium and loaded and executed by a processor.

[0164] For this reason, the present application provides a storage medium, which includes a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and the computer program can be loaded by a processor to execute the steps in any of the modal information generation methods provided by the present application.

[0165] For the specific implementation of each of the above operations, reference can be made to the previous embodiments and will not be elaborated here.

[0166] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0167] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the modal information generation methods provided by the present application, the beneficial effects that can be achieved by any of the modal information generation methods provided by the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.

[0168] The above has introduced in detail a modal information generation method, device, electronic device, and computer-readable storage medium provided by the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0169] Among them, it should be noted that in the specific implementation manner of the present application, it involves reference information in multiple modalities such as text, images, audio, or video that the user can input when using the content generation scheme corresponding to the modal information generation method, as well as the fusion weight parameter information of the modal conversion feature information for each modality during fusion, and aesthetic standard parameter data of the modal content information input by the user. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

Claims

1. A method for generating modal information, characterized in that, it includes: Obtain multimodal reference information, where the multimodal reference information includes modal reference information of at least two original modalities, and the modal reference information of the at least two original modalities includes at least two of text, image, audio, and video, and each original modality corresponds to a data type; For each modal reference information, process the modal reference information according to the binary modal conversion model of the original modality corresponding to the target modality to obtain modal conversion feature information. Among them, for different original modalities, through different binary modal conversion models, align the modal conversion feature information of the original modality with the features corresponding to the target modality; Fuse each modal conversion feature information based on the fusion weight to obtain the multimodal fusion feature information of the multimodal reference information in the target modality; Obtain the modal information generation model corresponding to the target modality and the expected modality; Generate modal content information corresponding to the multimodal fusion feature information through the modal information generation model; Among them, when the expected modality is the same as the target modality, the modal information generation model is a target modality generation model corresponding to the target modality, and the target modality generation model is used to convert the multimodal fusion feature information into the modal content information of the target modality; or, when the expected modality is other modalities except the target modality, the modal information generation model is a transition modality generation model, and the transition modality generation model is determined according to the target modality and the expected modality, and the transition modality generation model is used to convert the multimodal fusion feature information of the target modality into the modal content information of the expected modality; Each binary modal conversion model is trained based on the corresponding binary modal group sample information, and each binary modal group sample information of the binary modal conversion model of the original modality corresponding to the target modality includes the sample modal information of the original modality and the sample modal information of the target modality.

2. The modal information generation method according to claim 1, characterized in that, after generating the modal content information corresponding to the multimodal fusion feature information through the modal information generation model, the method further includes: Perform an aesthetic score on the modal content information through an aesthetic score model to obtain an actual aesthetic score; If the actual aesthetic score is less than the aesthetic score threshold, optimize the modal information generation model according to the aesthetic features to obtain an optimized modal information generation model, and the optimized modal information generation model is used to regenerate the new modal content information corresponding to the multimodal fusion feature information.

3. The modal information generation method according to claim 1, characterized in that, the fusing each modal conversion feature information to obtain the multimodal fusion feature information of the multimodal reference information in the target modality includes: Fuse the modal conversion feature information corresponding to the modal reference information of the at least two original modalities according to the fusion weights to obtain the multi-modal fusion feature information of the multi-modal reference information in the target modality.

4. The modal information generation method according to claim 1, wherein, the binary modal conversion model is trained through the following steps, including: Obtain at least two binary modal group sample information, which is composed of sample modal reference information belonging to the original modality and sample modal target information belonging to the target modality; For each original modality, construct an initial modal conversion model according to the original modality and the target modality, and the initial modal conversion model is used for the conversion between the sample modal reference information of the original modality and the sample modal target information of the target modality; Train each of the initial modal conversion models separately and map them uniformly to the feature space corresponding to the target modality to obtain the binary modal conversion models corresponding to the respective original modalities.

5. A modal information generation device, wherein, comprising: An information acquisition module, configured to acquire multi-modal reference information, where the multi-modal reference information includes modal reference information of at least two original modalities, and the modal reference information of the at least two original modalities includes at least two of text, image, audio, and video, and each original modality corresponds to a data type; A feature generation module, configured to process the modal reference information according to the binary modal conversion model corresponding to the original modality for each modal reference information to obtain modal conversion feature information, where, for different original modalities, different binary modal conversion models are used to align the modal conversion feature information of the original modality with the features corresponding to the target modality; A feature fusion module, configured to fuse the modal conversion feature information based on the fusion weights to obtain the multi-modal fusion feature information of the multi-modal reference information in the target modality; An information generation module, configured to obtain the modal information generation model corresponding to the desired modality of the target modality; generate the modal content information corresponding to the multi-modal fusion feature information through the modal information generation model; wherein, when the desired modality is the same as the target modality, the modal information generation model is a target modality generation model including the target modality corresponding to the target modality, and the target modality generation model is used to convert the multi-modal fusion feature information into the modal content information of the target modality; or, when the desired modality is other than the target modality, the modal information generation model is a transition modality generation model, and the transition modality generation model is determined according to the target modality and the desired modality, and the transition modality generation model is used to convert the multi-modal fusion feature information of the target modality into the modal content information of the desired modality; Each of the binary modality conversion models is trained based on the corresponding binary modality group sample information. Each binary modality group sample information of the binary modality conversion model corresponding to the original modality and the target modality includes the sample modality information of the original modality and the sample modality information of the target modality.

6. An electronic device, characterized in that, it includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps in the modality information generation method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium, characterized in that, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in the modality information generation method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • User portrait construction method and device and electronic equipment

    CN116150415A

  • Content generation method and device, model training method and device, electronic equipment and medium

    CN117094367A

  • Intelligent cataloging method for all-media news based on multi-modal information fusion understanding

    US20220270369A1