Multi-modal content generation method and device, readable medium and electronic equipment

By automatically generating multimodal content through a multi-style content generation model based on consistent initial materials, the problems of semantic consistency and process fragmentation in multimodal content generation are solved, and efficient, low-cost and stable multimodal content generation is achieved.

CN120723904APending Publication Date: 2025-09-30HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510832315.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies have semantic consistency risks in multimodal content generation. The generation process is fragmented and difficult to coordinate, and requires excessive human intervention, resulting in low production efficiency, high costs, and unstable product quality.

Method used

Generate content of different modalities based on consistent initial materials. Through a multi-style content generation model, automatically generate multimodal content, including summary text, descriptive images, descriptive audio, and descriptive videos, to reduce the risk of semantic consistency, unify and coordinate the generation process, and reduce manual intervention.

Benefits of technology

It achieves semantic consistency and improved production efficiency in multimodal content generation, reduces costs, and improves the stability of generation quality and production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723904A_ABST
    Figure CN120723904A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a multi-modal content generation method and device, a readable medium and electronic equipment. According to the method, an initial material can be collected from multiple data sources according to a theme label corresponding to a user side, and based on a content style corresponding to the user side, multi-modal content corresponding to the content style is generated on the basis of the initial material; wherein the multi-modal content can comprise at least two of abstract texts, description images, description audios and description videos. According to the scheme, in multi-modal content generation, contents of different modalities are generated based on consistent initial materials, so that the semantic consistency risk among multiple modalities is reduced, the generation process is unified and coordinated conveniently, the contents of different modalities can be automatically generated at the same time or in sequence according to business requirements, the cost is low, the production efficiency is high, and the method is suitable for large-scale popularization and application. Excessive manual intervention is not needed, and the generated content quality is stable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology. More specifically, embodiments of the present disclosure relate to a multimodal content generation method, a multimodal content generation device, a computer-readable storage medium, and an electronic device. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims, and no statement herein is admitted to be prior art by inclusion in this section.

[0003] Content generation can be based on artificial intelligence (AI) technology to generate different forms of content that conform to semantic logic, and can be widely used in fields such as text, images, voice and video.

[0004] Currently, content generation is usually single-modal. When it comes to multimodal content, there is a risk of semantic consistency. The generation process cannot be unified and coordinated, or requires excessive manual intervention, resulting in low production efficiency, high cost, and unstable quality. Summary of the Invention

[0005] In this context, embodiments of the present disclosure are intended to provide a multimodal content generation method, a multimodal content generation device, a computer-readable storage medium, and an electronic device.

[0006] According to a first aspect of an embodiment of the present disclosure, a multimodal content generation method is provided, which may include: collecting initial materials from multiple data sources based on topic tags corresponding to a user side; generating multimodal content based on the initial materials based on the content style corresponding to the user side; the multimodal content includes at least two of summary text, descriptive images, descriptive audio, and descriptive video.

[0007] Optionally, the multimodal content includes summary text, and based on the content style corresponding to the user side, the multimodal content is generated on the basis of the initial material, including: determining the target text generation model in the multi-style text generation model based on the content style corresponding to the user side; and generating summary text based on the initial material with the target text generation model.

[0008] Optionally, the multimodal content includes a description image, and based on the content style corresponding to the user side, the multimodal content is generated on the basis of the initial material, including: determining a target image generation model in a multi-style image generation model based on the content style corresponding to the user side; and generating a description image based on the initial material using the target image generation model.

[0009] Optionally, the multimodal content includes descriptive audio, and based on the content style corresponding to the user side, the multimodal content is generated on the basis of the initial material, including: determining the target audio generation model in the multi-style audio generation model based on the content style corresponding to the user side; and generating descriptive audio based on the target audio generation model on the basis of the initial material.

[0010] Optionally, the multimodal content includes descriptive audio, and based on the content style corresponding to the user side, the multimodal content is generated on the basis of the initial material, including: based on the content style corresponding to the user side, determining the target text generation model in the multi-style text generation model, and determining the target audio generation model in the multi-style audio generation model; based on the initial material, generating summary text with the target text generation model; based on the summary text, generating description audio with the target audio generation model.

[0011] Optionally, the multimodal content includes a descriptive video, and based on the content style corresponding to the user side, the multimodal content is generated on the basis of the initial material, including: determining a target video generation model in a multi-style video generation model based on the content style corresponding to the user side; generating a storyboard with the target video generation model based on the initial material, and obtaining a descriptive video based on the storyboard synthesis.

[0012] Optionally, the topic tag corresponding to the user side is determined according to the user side's first interaction behavior with historical multimodal content under different topic tags.

[0013] Optionally, the content style corresponding to the user side is determined according to the second interaction behavior of the user side to historical multimodal content with different content styles.

[0014] Optionally, based on the corresponding content style on the user side, after generating multimodal content based on the initial material, it also includes: dynamically selecting a corresponding distribution channel for distributing the multimodal content based on data traffic; the distribution channel includes a push channel and a pull channel.

[0015] Optionally, based on the content style corresponding to the user side, after generating multimodal content based on the initial material, it also includes: obtaining a third interactive behavior of the user side on the multimodal content; and updating at least one of the topic tag and content style based on the third interactive behavior.

[0016] Optionally, based on the corresponding content style on the user side, after generating multimodal content based on the initial material, it also includes: performing a quality review on the multimodal content; the quality review includes at least one of content compliance detection, quality score detection, and dynamic extraction detection.

[0017] According to a second aspect of an embodiment of the present disclosure, a multimodal content generation device is provided, which may include: a data acquisition module for collecting initial materials from multiple data sources based on topic tags corresponding to the user side; a content generation module for generating multimodal content based on the initial materials based on the content style corresponding to the user side; the multimodal content includes at least two of summary text, descriptive images, descriptive audio, and descriptive video.

[0018] Optionally, the multimodal content includes summary text, and the content generation module is specifically used to determine the target text generation model in the multi-style text generation model based on the content style corresponding to the user side; based on the initial material, the summary text is generated using the target text generation model.

[0019] Optionally, the multimodal content includes a description image, and a content generation module is specifically used to determine a target image generation model in a multi-style image generation model based on the content style corresponding to the user side; based on the initial material, a description image is generated using the target image generation model.

[0020] Optionally, the multimodal content includes descriptive audio, and the content generation module is specifically used to determine the target audio generation model in the multi-style audio generation model based on the content style corresponding to the user side; based on the initial material, the descriptive audio is generated using the target audio generation model.

[0021] Optionally, the multimodal content includes descriptive audio, and a content generation module is specifically used to determine a target text generation model in a multi-style text generation model based on the corresponding content style on the user side, and to determine a target audio generation model in a multi-style audio generation model; based on the initial material, a summary text is generated using the target text generation model; based on the summary text, a description audio is generated using the target audio generation model.

[0022] Optionally, the multimodal content includes a descriptive video and a content generation module, which is specifically used to determine a target video generation model in a multi-style video generation model based on the corresponding content style on the user side; based on the initial material, generate a storyboard with the target video generation model, and obtain a descriptive video based on the storyboard synthesis.

[0023] Optionally, the topic tag corresponding to the user side is determined according to the user side's first interaction behavior with historical multimodal content under different topic tags.

[0024] Optionally, the content style corresponding to the user side is determined according to the second interaction behavior of the user side to historical multimodal content with different content styles.

[0025] Optionally, the device further includes a content distribution module for dynamically selecting a corresponding distribution channel for distribution of the multimodal content based on data traffic; the distribution channel includes a push channel and a pull channel.

[0026] Optionally, the device further includes a data updating module for obtaining a third interactive behavior of the user side on the multimodal content; and updating at least one of the topic tag and the content style based on the third interactive behavior.

[0027] Optionally, the device further includes a quality audit module for performing a quality audit on the multimodal content; the quality audit includes at least one of content compliance detection, quality score detection, and dynamic extraction detection.

[0028] According to a third aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any one of the above-mentioned multimodal content generation methods is implemented.

[0029] According to a fourth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned multimodal content generation methods by executing the executable instructions.

[0030] According to the multimodal content generation method, multimodal content generation device, computer-readable storage medium and electronic device of the embodiment of the present disclosure, initial materials can be collected from multiple data sources based on the corresponding topic tags on the user side, and based on the corresponding content style on the user side, multimodal content corresponding to the content style can be generated on the basis of the initial materials; wherein, the multimodal content can include at least two of summary text, description image, description audio, and description video. In the generation of multimodal content, this solution generates content of different modalities based on consistent initial materials, thereby reducing the risk of semantic consistency between multiple modalities and facilitating the unified and coordinated generation process. It can automatically generate content of different modalities simultaneously or sequentially according to business needs, with low cost and high production efficiency, without excessive human intervention, and with stable quality of generated content. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:

[0032] Figure 1 FIG1 shows one of the flow charts of the multimodal content generation method according to an embodiment of the present disclosure;

[0033] Figure 2A schematic diagram of an operation interface of a multimodal content generation method according to an embodiment of the present disclosure is shown;

[0034] Figure 3 FIG2 shows a second flowchart of a multimodal content generation method according to an embodiment of the present disclosure;

[0035] Figure 4 A schematic diagram of an implementation architecture of a multimodal content generation method according to an embodiment of the present disclosure is shown;

[0036] Figure 5 A schematic diagram showing a multimodal content generation device according to an embodiment of the present disclosure is shown;

[0037] Figure 6 A schematic diagram of a storage medium according to an embodiment of the present disclosure is shown;

[0038] Figure 7 A schematic diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0039] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. Rather, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0040] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, apparatus, device, method, or computer program product. Therefore, the present disclosure may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software.

[0041] According to an embodiment of the present disclosure, a multimodal content generation method, a multimodal content generation device, a computer-readable storage medium, and an electronic device are provided.

[0042] In this document, any number of elements in the drawings is for illustration and not for limitation, and any naming is for distinction only and does not have any limiting meaning.

[0043] The principles and spirit of the present disclosure are described in detail below with reference to several representative embodiments of the present disclosure. SUMMARY OF THE INVENTION

[0045] As business content on internet platforms evolves towards multimodality, the demand for quality content generation is also increasing. Currently, independent modules can be used to generate content in different modalities. However, these modules lack effective coordination mechanisms, posing a risk of semantic inconsistency. Furthermore, the fragmented production processes of different modules hinder unified management and coordination. Alternatively, content in different modalities can be aligned or manually verified after generation. However, this process is cumbersome, complex, and inefficient, and the quality of the finished product depends on the initial generation, making stability uncertain.

[0046] Alternatively, multimodal content generation and planning can be led by humans, with AI assisting in specific steps of multimodal content production. However, manual intervention leads to lower production efficiency and higher costs, and the lack of an end-to-end quality control system leads to significant fluctuations in the quality of finished products.

[0047] In the embodiments of the present disclosure, when generating multimodal content, content of different modalities is generated based on consistent initial materials, which reduces the risk of semantic consistency between multimodalities and facilitates unified and coordinated management. Content of different modalities can also be automatically generated simultaneously or successively according to business needs, with low cost and high production efficiency, without the need for excessive human intervention, and the quality of generated content is stable.

[0048] After introducing the basic principles of the present invention, various non-limiting embodiments of the present invention are described in detail below.

[0049] Example application scenarios

[0050] It should be noted that the following application scenarios are only provided to facilitate understanding of the spirit and principles of the present invention, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0051] The multimodal content generation method of the embodiments of the present disclosure can be applied to a variety of application scenarios involving multimodal content generation, publishing, and sharing.

[0052] In one application scenario, a music platform may be involved. Typically, in this application scenario, music and podcasts can be generated, published, and shared, where multimodality can include text, images, audio, video, and the like. In this application scenario, the multimodal content generation method of the embodiment of the present disclosure can be used to generate multimodal content corresponding to the user-side content style based on the initial materials collected by the corresponding topic tags on the user side, such as generating summary text, describing images, describing audio, describing videos, and the like. The multimodal content can be music, podcasts, and the like to be shared.

[0053] In other application scenarios, short video platforms, instant messaging platforms, information platforms, etc. may also be involved, and the embodiments of the present disclosure do not impose specific restrictions on this.

[0054] Exemplary Methods

[0055] In combination with the above application scenarios, Figure 1 A multimodal content generating method according to an exemplary embodiment of the present disclosure will be described.

[0056] like Figure 1 One of the processes of the multimodal content generation method according to the exemplary embodiment of the present disclosure may include the following steps 101 to 102:

[0057] Step 101: Collect initial materials from multiple data sources based on the corresponding topic tags on the user side.

[0058] In an embodiment of the present disclosure, when generating multimodal content, it can be executed based on the collected initial materials. The initial materials can be texts or pictures that describe events, concepts, methods, etc., and can be collected based on the corresponding topic tags on the user side. On the user side, the topic tags can represent the user's topic preferences for multimodal content. The topic tags can be associated with the topics expressed by the initial materials, and one initial material can be associated with one or several topic tags. Depending on the multimodal content push object, the topic tags corresponding to the user side can come from a single user or from a user group, and different topic tags of a single user can also be clustered or merged in the user group.

[0059] In the embodiment of the present disclosure, when initial materials can be collected from different data sources based on topic tags, the data sources may include data content published by online platforms, such as information platforms, social platforms, audio and video platforms, etc., and may also include data content submitted by offline users; wherein, different data sources may be of different types, such as collecting initial materials associated with topic tags from information platforms and social platforms respectively, or they may be different entities of the same type, such as collecting initial materials associated with topic tags from information platform 1 and information platform 2 respectively.

[0060] In an optional method embodiment of the present disclosure, the topic tag corresponding to the user side is determined according to the user side's first interaction behavior with historical multimodal content under different topic tags.

[0061] In the disclosed embodiment, the first interactive behavior may be a user's browsing, posting, sharing, downloading, favorite, commenting, or other actions on multimodal content under different topic tags. By collecting and analyzing the first interactive behaviors occurring on the user side, the user's topic preferences for historical multimodal content can be assessed, and the corresponding topic tags on the user side can be extracted.

[0062] In an optional embodiment of the present disclosure, the subject tag may also be determined based on the object selected by the user from the tag options.

[0063] Step 102: Based on the content style corresponding to the user side, multimodal content is generated on the basis of the initial material; the multimodal content includes at least two of summary text, description image, description audio, and description video.

[0064] In the embodiments of the present disclosure, on the user side, the content style can represent the user's style preference for multimodal content. It can be extracted from the content involved in the user's browsing, publishing, sharing and other historical behaviors, or it can be determined based on the object selected by the user from the tag options. The content style can be associated with the presentation method of multimodal content, and each modality can have a corresponding content style. During the content generation process, different corresponding content styles can be output for the same input, thereby obtaining multimodal content that adapts to the diverse style requirements of the user side.

[0065] In an optional method embodiment of the present disclosure, the content style corresponding to the user side is determined according to the second interaction behavior of the user side with historical multimodal content under different content styles.

[0066] In the disclosed embodiment, the second interactive behavior can be a user's browsing, publishing, sharing, downloading, favorite, commenting, and other behaviors on multimodal content with different content styles. By collecting and analyzing the second interactive behaviors that occurred on the user side, the user's style preference for historical multimodal content can be assessed, and the corresponding content style on the user side can be extracted.

[0067] In the embodiments of the present disclosure, multimodal content may include a combination of any two or more of summary text, image descriptions, audio descriptions, and video descriptions. Depending on the selection of multimodal content, the corresponding modal content generation method can be flexibly called, and the process of generating different modal content can be arranged according to needs. This embodiment of the present disclosure does not impose specific restrictions on this. Multimodal content can be generated using a content generation model for the corresponding modality.

[0068] In an optional method embodiment of the present disclosure, the aforementioned multimodal content includes summary text, and the aforementioned step 102 may include the following steps A1 to A2.

[0069] Step A1: Based on the content style corresponding to the user side, determine the target text generation model in the multi-style text generation model.

[0070] In the disclosed embodiments, the summary text can summarize and refine the content of the initial material and can have different content styles. Different content styles can correspond to different text generation models. The text generation model can correspond to one or more content styles. When the text generation model corresponds to multiple content styles, the output content style can be specified using prompt words.

[0071] For example, the content style of the summary text may include tone, language, symbols, etc. The tone may include restoring the original material, or may include personalized styles such as humor, exaggeration, and seriousness. It may also include specific genres and styles, such as documentary, epistolary, and prose. The language may be Chinese, English, Japanese, etc., and may be a single language or a mixed language. The symbols may be emoticons, and the content, type, and number of the symbols may be set. Based on the content style of the summary text, a target text generation model that can generate summary text in that content style is selected from the multi-style text generation models. The target text generation model may correspond only to that content style, or may correspond to several content styles including that content style.

[0072] In the embodiment of the present disclosure, the performance and effect of the text generation model can also be experimentally evaluated on the user side, and the evaluation results can be referenced in the decision-making process of selecting the target text generation model. For example, multiple text data can be generated simultaneously by text generation models of different content styles, or different text generation models of the same content style, for user-side experimental evaluation. The feedback of the text data on the user side, such as exposure, clicks, user portraits, etc., is analyzed to determine the user side's preference scores for different text generation models, and then the target text generation model is selected with reference to the preference scores. On this basis, model addition and elimination rules based on preference scores can also be established. The text generation model can be eliminated when its preference score is lower than or is lower than a set threshold several times in a row, or its weight can be increased when its preference score is higher than or is higher than a set threshold several times in a row. The embodiment of the present disclosure does not impose specific restrictions on this.

[0073] Step A2: Based on the initial material, generate summary text using the target text generation model.

[0074] In the embodiment of the present disclosure, based on determining the target text generation model, the initial material can be input into the target text generation model so that the target text generation model extracts a summary of the initial material and outputs a summary text in the content style corresponding to the user side.

[0075] In an optional method embodiment of the present disclosure, the aforementioned multimodal content includes a description image, and the aforementioned step 102 may include the following steps B1 and B2.

[0076] Step B1: Based on the content style corresponding to the user side, determine the target image generation model in the multi-style image generation model.

[0077] In the disclosed embodiments, the description image represents the content of the initial material using image content. This image content may have different content styles, and different content styles may correspond to different image generation models. The image generation model may correspond to one or more content styles. When the image generation model corresponds to multiple content styles, a prompt word may be used to specify the output content style.

[0078] For example, the content style of the described image may include genre, painting style, color, etc. Genre may include cover, illustration, etc. Painting style may include watercolor, oil painting, animation, movie, comics, etc. Color may include black and white, color, etc. Based on the content style of the described image, a target image generation model that can generate a description image with the content style is selected from the multi-style image generation models.

[0079] The target text generation model may correspond only to the content style, or may correspond to several content styles including the content style.

[0080] In the embodiment of the present disclosure, the performance and effect of the image generation model can also be experimentally evaluated on the user side, and the evaluation results can be referenced in the decision-making process of selecting the target image generation model. For example, multiple sets of image data can be generated simultaneously by image generation models of different content styles, or different image generation models of the same content style, for user-side experimental evaluation. The image data can be analyzed based on the feedback on the user side, such as exposure, clicks, user portraits, etc., to determine the user side's preference scores for different image generation models, and then the target image generation model can be selected with reference to the preference scores. On this basis, model addition and elimination rules based on preference scores can also be established. The image generation model can be eliminated when its preference score is lower than or is lower than a set threshold several times in a row, or its weight can be increased when its preference score is higher than or is higher than a set threshold several times in a row. The embodiment of the present disclosure does not impose specific restrictions on this.

[0081] Step B2: Based on the initial material, a description image is generated using the target image generation model.

[0082] In the embodiment of the present disclosure, based on the determination of the target image generation model, the initial material can be input into the target image generation model so that the target image generation model can perform an image representation of the initial material and output a description image in the content style corresponding to the user side.

[0083] In an optional method embodiment of the present disclosure, the aforementioned multimodal content includes descriptive audio, and the aforementioned step 102 may include the following steps C1 to C2.

[0084] Step C1: Based on the content style corresponding to the user side, determine the target audio generation model in the multi-style audio generation model.

[0085] In the disclosed embodiments, the descriptive audio is a spoken representation of the content of the initial material and can have different content styles, each of which corresponds to a different audio generation model. The audio generation model can correspond to one or more content styles, and when the audio generation model corresponds to multiple content styles, the output content style can be specified using a prompt word.

[0086] For example, the content style of the audio description may include tone, timbre, etc. Tone may include personalized styles such as humor, exaggeration, and seriousness, while timbre may include male, female, or child voices, or the timbre of a specific character. Based on the content style of the audio description, a target audio generation model that can generate audio descriptions in that content style is selected from multiple audio generation models. The target audio generation model may correspond only to that content style, or may correspond to several content styles including that content style.

[0087] In the embodiment of the present disclosure, the performance and effect of the audio generation model can also be experimentally evaluated on the user side, and the evaluation results can be referenced in the decision-making process of selecting the target audio generation model. For example, multiple audio data can be generated simultaneously by audio generation models of different content styles, or different audio generation models of the same content style, for user-side experimental evaluation. The feedback of the audio data on the user side, such as exposure, clicks, playback, user portraits, etc., can be analyzed to determine the user side's preference scores for different audio generation models, and then the target audio generation model can be selected with reference to the preference scores. On this basis, model addition and elimination rules based on preference scores can also be established. The audio generation model can be eliminated when its preference score is lower than or is lower than a set threshold several times in a row, or its weight can be increased when its preference score is higher than or is higher than a set threshold several times in a row. The embodiment of the present disclosure does not impose specific restrictions on this.

[0088] Step C2: Based on the initial material, generate description audio using the target audio generation model.

[0089] In the embodiment of the present disclosure, based on the determination of the target audio generation model, the initial material can be input into the target audio generation model so that the target audio generation model performs audio conversion on the initial material and outputs description audio in the content style corresponding to the user side.

[0090] In an optional method embodiment of the present disclosure, the aforementioned multimodal content includes descriptive audio, and the aforementioned step 102 may include the following steps D1 to D3.

[0091] Step D1: Based on the content style corresponding to the user side, determine a target text generation model in the multi-style text generation model, and determine a target audio generation model in the multi-style audio generation model.

[0092] In the disclosed embodiment, the description audio may be a phonetic representation of the content of the summary text after extracting the summary text based on the initial material. Determining the target text generation model can refer to the description of step A1 above, and determining the target audio generation model can refer to the description of step C1 above. To avoid repetition, these details are not repeated here.

[0093] Step D2: Based on the initial material, generate summary text using the target text generation model.

[0094] In the embodiment of the present disclosure, step D2 may correspond to the relevant description of the aforementioned step A2, and will not be repeated here to avoid repetition.

[0095] Step D3: Based on the summary text, generate description audio using the target audio generation model.

[0096] In an embodiment of the present disclosure, based on the determination of the target audio generation model, the summary text can be input into the target audio generation model so that the target audio generation model performs audio conversion on the summary text and outputs description audio in the content style corresponding to the user side.

[0097] In an optional method embodiment of the present disclosure, the aforementioned multimodal content includes a description video, and the aforementioned step 102 may include the following steps E1 to E2.

[0098] Step E1: Based on the content style corresponding to the user side, determine the target video generation model in the multi-style video generation model.

[0099] In the disclosed embodiments, the description video can be a dynamic video representation of the content of the initial material, and can have different content styles. Different content styles can correspond to different video generation models. The video generation model can correspond to one or more content styles. When the video generation model corresponds to multiple content styles, the output content style can be specified using prompt words.

[0100] For example, the content style of a description video may include scenes, painting styles, etc. Scenes may include dynamic scenes and static scenes, and painting styles may include movies, animations, etc. Based on the content style of the description video, a target video generation model that can generate a description video of that content style is selected from multiple video generation models. The target video generation model may correspond only to that content style, or may correspond to several content styles including that content style.

[0101] In the embodiment of the present disclosure, the performance and effect of the video generation model can also be experimentally evaluated on the user side, and the evaluation results can be referenced in the decision-making process of selecting the target video generation model. For example, multiple video data can be generated simultaneously by video generation models of different content styles, or different video generation models of the same content style, for user-side experimental evaluation. The feedback of the video data on the user side, such as exposure, clicks, playback, user portraits, etc., is analyzed to determine the user side's preference scores for different video generation models, and then the target video generation model is selected with reference to the preference scores. On this basis, model addition and elimination rules based on preference scores can also be established. The video generation model can be eliminated when its preference score is lower than or is lower than a set threshold several times in a row, or its weight can be increased when its preference score is higher than or is higher than a set threshold several times in a row. The embodiment of the present disclosure does not impose specific restrictions on this.

[0102] Step E2: Based on the initial material, generate a storyboard using the target video generation model, and obtain a description video based on the storyboard synthesis.

[0103] In the embodiment of the present disclosure, on the basis of determining the target video generation model, the initial material can be input into the target video generation model so that the target video generation model generates a storyboard of the corresponding content style based on the initial material, and then synthesizes the video data based on the storyboard to obtain a description video.

[0104] For example, Figure 2 A schematic diagram of the user interface for a multimodal content generation method according to an exemplary embodiment of the present disclosure is shown. Taking a news data source and a podcast as an example, an initial item with the data identifier "JSN92TI0529AQIE" is collected based on the user-side topic tag "Sports." The collection time for this initial item is 20xx-04-09 15:01:58. The podcast identifier is 1214389768, and the program identifier is 3075331364.

[0105] The initial material may include a news title, publisher, and news content. The news title is "Club A has lost over £200 million over two seasons and will be forced to sell players this summer," the publisher is "Sports Live," and the news content is displayed at the storage address "yyimgs / 20xx0409150157894txt30365675," which can be retrieved and displayed using the "View Details" option.

[0106] Podcasts include multimodal content such as summary text, descriptive images, and descriptive audio. They are generated by calling a content generation model corresponding to the content style. The summary text is extracted from the initial material by the target text generation model. When expanded, it is shown as follows:

[0107] "Club A has suffered after-tax losses of more than 200 million pounds in the past two seasons, and its financial situation is worrying. In order to meet the requirements of the Financial Fair Play Act, Club A will be forced to sell key players this summer to raise funds. Although the team has made significant progress, losses caused by huge investments are inevitable. Club A hopes to increase revenue by improving the team's performance, but in the short term it still needs to balance its finances through transfers."

[0108] As shown above, the summary text can be displayed in the expanded state by the "fold" option, and the full text can be displayed by switching to the collapsed state to display a certain number of characters. In the collapsed state, an "expand" option (not shown) is provided to switch to the expanded state.

[0109] The description audio is obtained by converting the initial material by the target audio generation model. As the podcast sound, it includes female and male voice versions, which can be played and viewed through the playback option.

[0110] The description image is obtained by converting the initial material into a target image generation model, serving as the podcast cover, including the characters and brief text information.

[0111] like Figure 2 As shown, the interface also includes the status and detailed information of specific execution tasks. For example, when collapsed, it shows that summary generation is complete, with an end time of "20xx-04-09 15:02:03"; audio generation is complete, with an end time of "20xx-04-09 15:02:12"; program generation is complete, with an end time of "20xx-04-09 15:02:15"; and program push is complete, with an end time of "20xx-04-09 15:02:35". You can view more detailed information about each task or view information about more tasks by clicking the "Expand" option, which is not detailed here.

[0112] In addition, the interface also includes operation options for the generated multimodal content, such as "View Log", which can view the process record of multimodal content generation; "Program Deletion", which can delete the generated multimodal content and delete the podcast program generation task; "Program Generation", which can generate podcast programs based on the generated multimodal content; "Reset Data", which can delete the generated multimodal content; "Text to Speech Editing", which can edit the description audio.

[0113] like Figure 3 The second process of the multimodal content generation method according to the exemplary embodiment of the present disclosure may include the following steps 301 to 306:

[0114] Step 301: Collect initial materials from multiple data sources based on the corresponding topic tags on the user side.

[0115] In the embodiment of the present disclosure, step 301 may correspond to the relevant description of the aforementioned step 101, and will not be described again here to avoid repetition.

[0116] In the embodiments of the present disclosure, on the basis of collecting initial materials from multiple data sources, the initial materials can also be pre-processed, such as deduplication, label verification, format conversion, etc., and evaluated and selected according to the data source, such as scoring based on the constructed data source indicator system, and eliminating or retaining initial materials whose data sources meet the requirements based on the score. The data source indicator system can be designed according to the data source type, scale, public credibility, etc., and can also be designed according to user feedback.

[0117] In the disclosed embodiment, the collection of initial materials from multiple data sources can be accomplished by deploying a distributed web crawler cluster to collect structured data from the data sources in real time. A dynamic feature library can also be established in the multiple data sources to identify hot topics. When a hot topic is related to a topic tag on the user side, the collection of initial materials from multiple data sources is triggered, thereby reducing the delay in collecting initial materials and improving the accuracy of multimodal content generation.

[0118] Step 302: Based on the content style corresponding to the user side, multimodal content is generated on the basis of the initial material; the multimodal content includes at least two of summary text, description image, description audio, and description video.

[0119] In the embodiment of the present disclosure, step 302 may correspond to the relevant description of the aforementioned step 102, and will not be described again here to avoid repetition.

[0120] In an optional method embodiment of the present disclosure, the above-mentioned step 302 further includes the following step 303.

[0121] Step 303: Dynamically select a corresponding distribution channel for distributing the multimodal content based on data traffic; the distribution channel includes a push channel and a pull channel.

[0122] In the disclosed embodiments, multiple distribution channels can be used to distribute multimodal content. For example, a push channel can be used to push content to the user side based on real-time transmission; a pull channel can also be used to establish a content mirror node, which is then pulled by the user side. The push channel can be transmitted using an API (Application Programming Interface) and can support HTTPS (Hypertext Transfer Protocol Secure) / WebSocket dual-protocol execution; the pull channel can be implemented through a microservice architecture. By setting up multiple distribution channels, the business needs of the user side can be more flexibly met, and the data transmission efficiency can be improved.

[0123] like Figure 2 As shown, the operation options also include "program push", which can actively push multimodal content to the user side; or, the user side can actively pull it.

[0124] In an optional method embodiment of the present disclosure, the aforementioned step 302 also includes the following steps 304 to 305.

[0125] Step 304: Obtain a third interaction behavior of the user on the multimodal content.

[0126] In the embodiments of the present disclosure, the third interactive behavior can be interactive feedback from the user to the multimodal content. For example, feedback on the exposure, clicks, sharing, comments, favorites, and completion of the aforementioned podcast program. The third interactive behavior can be obtained through tracking technology or other means, and the embodiments of the present disclosure do not impose specific restrictions on this.

[0127] Step 305: Update at least one of the topic tag and the content style based on the third interactive behavior.

[0128] In the disclosed embodiment, data analysis based on the third interactive behavior can obtain the user's latest interest preferences for multimodal content. On this basis, topic tags, content style, etc. can be updated. Among them, topic tags can correspond to updates on the data source side, and content style can correspond to updates on the model side. Through the updates of topic tags and content style, the selection of data sources and content generation models in the data collection and content generation links can be adjusted to make the output of multimodal content more consistent with the user's long-term and short-term dynamic interest preferences.

[0129] In an optional method embodiment of the present disclosure, the above-mentioned step 302 further includes the following step 306.

[0130] Step 306: Perform a quality review on the multimodal content; the quality review includes at least one of content compliance detection, quality score detection, and dynamic extraction detection.

[0131] In the disclosed embodiments, the generated multimodal content can also be quality audited to promptly identify and address potential problems, thereby improving the quality of the multimodal content. The quality audit can include multi-level audits, such as content compliance testing, which can include multiple audit dimensions, such as character correctness, text formatting, semantic context, and voice or video coherence. A quality scoring model can also be used for quality scoring, either a single quality scoring model or a comprehensive scoring using multiple quality scoring architectures. Dynamic sampling detection can also be used to perform manual or machine sampling with a dynamically adjusted sampling rate.

[0132] In the embodiments of the present disclosure, one or more combinations of the aforementioned multi-channel distribution, topic tag update, content style update, and quality review schemes may be selected for implementation based on the generation of multimodal content. For example, multi-channel distribution may be performed based on quality review; topic update and / or content style update may be performed based on multi-channel distribution; multi-channel distribution may be performed based on quality review, and topic update and / or content style update may be performed, etc. Those skilled in the art may make selections based on actual needs, and the embodiments of the present disclosure do not impose specific limitations on this.

[0133] like Figure 4 The schematic diagram of the implementation architecture of the multimodal content generation method of the exemplary embodiment of the present disclosure can include a data source, a server, and a user side. The server includes a data acquisition module, a multimodal content generation engine, an intelligent review system, and a distributed publishing system. Taking the multimodal content of a podcast as an example, the implementation process can be as follows:

[0134] The server collects initial materials from the data source through the data collection module, processes them, and pushes them to the multimodal content generation engine; the multimodal content generation engine calls the corresponding content generation model to generate summary text, description images, and description audio based on the initial materials; the server uses the summary text as the introduction of the podcast program, the description image as the cover of the podcast program, and the description audio as the sound of the podcast program to generate the podcast program into the library; the server conducts a quality review of the podcast program through the intelligent review system, and puts the podcast program into the podcast pool after the quality review is passed; the server publishes the podcast programs in the podcast pool to the user side through the distributed publishing system, which can be achieved through multiple distribution channels, such as the user side pulling the podcast program through the open API (OpenAPI), or the server side pushing the podcast program to the user side through the distributed publishing system.

[0135] Exemplarily, the multimodal generation engine can be implemented using the DeepSeek-R3 model, and the summary text can be extracted based on the Attention mechanism, and the summary compression ratio can be adjusted within the range of 30-70%; the image description can be cross-modally converted to generate a 1024×768 resolution image; the audio description can support TTS conversion of 12 timbre libraries, with a MOS (Mean Opinion Score) score of 4.2 for speech naturalness; the video description supports H.265 encoding, and can perform automated video rendering based on storyboards.

[0136] According to the multimodal content generation method of the embodiment of the present disclosure, initial materials can be collected from multiple data sources based on the corresponding topic tags on the user side, and based on the corresponding content style on the user side, multimodal content corresponding to the content style can be generated on the basis of the initial materials; wherein, the multimodal content can include at least two of summary text, description image, description audio, and description video. In the generation of multimodal content, this solution generates content of different modalities based on consistent initial materials, thereby reducing the risk of semantic consistency between multiple modalities and facilitating the unified and coordinated generation process. It can automatically generate content of different modalities simultaneously or sequentially according to business needs, with low cost and high production efficiency, without excessive human intervention, and with stable quality of generated content.

[0137] Exemplary devices

[0138] After introducing the multimodal content generation method according to the exemplary embodiment of the present disclosure, Figure 5 A multimodal content generating apparatus according to an exemplary embodiment of the present disclosure will be described.

[0139] It should be noted that other specific details of the various functional modules of the multimodal content generation device of the embodiment of the present disclosure have been described in detail in the embodiment of the multimodal content generation method described above and will not be repeated here.

[0140] Figure 5 A multimodal content generation device 500 according to an exemplary embodiment of the present disclosure is shown. The device may include: a data collection module 501 for collecting initial materials from multiple data sources based on topic tags corresponding to the user side; a content generation module 502 for generating multimodal content based on the initial materials based on the content style corresponding to the user side; the multimodal content includes at least two of summary text, descriptive images, descriptive audio, and descriptive video.

[0141] In an optional device embodiment of the present disclosure, the multimodal content includes summary text, and the content generation module 502 is specifically used to determine the target text generation model in the multi-style text generation model based on the content style corresponding to the user side; based on the initial material, generate the summary text with the target text generation model.

[0142] In an optional device embodiment of the present disclosure, the multimodal content includes a description image, and the content generation module 502 is specifically used to determine a target image generation model in a multi-style image generation model based on the content style corresponding to the user side; and generate a description image based on the target image generation model based on the initial material.

[0143] In an optional device embodiment of the present disclosure, the multimodal content includes descriptive audio, and the content generation module 502 is specifically used to determine the target audio generation model in the multi-style audio generation model based on the content style corresponding to the user side; and generate descriptive audio based on the target audio generation model on the basis of the initial material.

[0144] In an optional device embodiment of the present disclosure, the multimodal content includes descriptive audio, and the content generation module 502 is specifically used to determine the target text generation model in the multi-style text generation model and the target audio generation model in the multi-style audio generation model based on the corresponding content style on the user side; generate summary text based on the initial material using the target text generation model; and generate descriptive audio based on the target audio generation model.

[0145] In an optional device embodiment of the present disclosure, the multimodal content includes a descriptive video, and the content generation module 502 is specifically used to determine a target video generation model in a multi-style video generation model based on the corresponding content style on the user side; generate a storyboard with the target video generation model based on the initial material, and obtain a descriptive video based on the synthesis of the storyboard.

[0146] In an optional device embodiment of the present disclosure, the topic tag corresponding to the user side is determined according to the user side's first interaction behavior with historical multimodal content under different topic tags.

[0147] In an optional device embodiment of the present disclosure, the content style corresponding to the user side is determined according to the second interaction behavior of the user side with historical multimodal content under different content styles.

[0148] In an optional device embodiment of the present disclosure, the device also includes a content distribution module for dynamically selecting a corresponding distribution channel for distribution of multimodal content based on data traffic; the distribution channel includes a push channel and a pull channel.

[0149] In an optional device embodiment of the present disclosure, the device also includes a data update module for obtaining a third interactive behavior of the user side on the multimodal content; and updating at least one of the topic tag and the content style based on the third interactive behavior.

[0150] In an optional device embodiment of the present disclosure, the device also includes a quality review module for performing quality review on multimodal content; the quality review includes at least one of content compliance detection, quality score detection, and dynamic extraction detection.

[0151] According to the multimodal content generation device of the embodiment of the present disclosure, initial materials can be collected from multiple data sources based on the corresponding topic tags on the user side, and based on the corresponding content style on the user side, multimodal content corresponding to the content style can be generated on the basis of the initial materials; wherein, the multimodal content can include at least two of summary text, description image, description audio, and description video. In the generation of multimodal content, this solution generates content of different modalities based on consistent initial materials, thereby reducing the risk of semantic consistency between multiple modalities and facilitating the unified and coordinated generation process. It can automatically generate content of different modalities simultaneously or sequentially according to business needs, with low cost and high production efficiency, without excessive human intervention, and with stable quality of generated content.

[0152] It should be noted that although several modules or units of the multimodal content generation device are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided to be embodied by multiple modules or units.

[0153] Exemplary Storage Media

[0154] The storage medium according to the exemplary embodiment of the present disclosure will be described below.

[0155] In this exemplary embodiment, referring to Figure 6 As shown, a program product 600 for implementing the above method according to an exemplary embodiment of the present disclosure is described. For example, a portable compact disk read-only memory (CD-ROM) may be used and includes program code, and can be run on a device such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0156] The program product 600 can be implemented in any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0157] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0158] The program code contained on the readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RE, etc., or any suitable combination of the foregoing.

[0159] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (FAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0160] Exemplary electronic devices

[0161] refer to Figure 7 An electronic device according to an exemplary embodiment of the present disclosure will be described.

[0162] Figure 7 The electronic device 700 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0163] like Figure 7 As shown, electronic device 700 is implemented as a general-purpose computing device. Components of electronic device 700 may include, but are not limited to, at least one processing unit 710, at least one storage unit 720, a bus 730 connecting various system components (including storage unit 720 and processing unit 710), and a display unit 740.

[0164] The storage unit stores program codes, which can be executed by the processing unit 710, so that the processing unit 710 performs the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above. For example, the processing unit 710 can perform the following steps: Figure 1 The method steps shown, etc.

[0165] The storage unit 720 may include a volatile storage unit, such as a random access memory unit (RAM) 721 and / or a cache memory unit 722 , and may further include a read-only memory unit (ROM) 723 .

[0166] The storage unit 720 may also include a program / utility 724 having a set (at least one) of program modules 725, such program modules 725 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0167] The bus 730 may include a data bus, an address bus, and a control bus.

[0168] The electronic device 700 can also communicate with one or more external devices 800 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), and such communication can be performed via an input / output (I / O) interface 750. The electronic device 700 also includes a display unit 740, which is connected to the input / output (I / O) interface 750 for display. In addition, the electronic device 700 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 760. As shown, the network adapter 760 communicates with other modules of the electronic device 700 via a bus 730. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0169] It should be noted that although several modules or submodules of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in a single unit / module. Conversely, the features and functions of a single unit / module described above can be further divided and embodied by multiple units / modules.

[0170] Furthermore, although the operations of the disclosed method are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0171] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division into various aspects does not mean that the features in these aspects cannot be combined to benefit. Such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the appended claims.

Claims

1. A multimodal content generation method, characterized in that: The method comprises: Collect initial materials from multiple data sources based on the corresponding topic tags on the user side; Based on the content style corresponding to the user side, multimodal content is generated on the basis of the initial material; the multimodal content includes at least two of summary text, description image, description audio, and description video.

2. The method according to claim 1, characterized in that The multimodal content includes summary text, and the generating of the multimodal content based on the initial material based on the content style corresponding to the user side includes: Determining a target text generation model from among the multi-style text generation models based on the content style corresponding to the user side; Based on the initial material, the summary text is generated using the target text generation model.

3. The method according to claim 1, characterized in that The multimodal content includes a description image, and the generating of the multimodal content based on the initial material based on the content style corresponding to the user side includes: Determining a target image generation model in a multi-style image generation model based on the content style corresponding to the user side; Based on the initial material, the description image is generated using the target image generation model.

4. The method according to claim 1, wherein The multimodal content includes a description audio, and the generating of the multimodal content based on the initial material based on the content style corresponding to the user side includes: Determining a target audio generation model from multiple audio generation models based on the content style corresponding to the user side; Based on the initial material, the description audio is generated using the target audio generation model.

5. The method according to claim 1, wherein The multimodal content includes a description audio, and the generating of the multimodal content based on the initial material based on the content style corresponding to the user side includes: Determining a target text generation model from among the multiple-style text generation models and a target audio generation model from among the multiple-style audio generation models based on the content style corresponding to the user side; Based on the initial material, generating the summary text using the target text generation model; Based on the summary text, the description audio is generated using the target audio generation model.

6. The method according to claim 1, characterized in that The multimodal content includes a description video, and the generating of the multimodal content based on the initial material based on the content style corresponding to the user side includes: Determining a target video generation model from multiple-style video generation models based on the content style corresponding to the user side; Based on the initial material, a storyboard is generated using the target video generation model, and the description video is synthesized based on the storyboard.

7. The method according to any one of claims 1 to 6, characterized in that The topic tag corresponding to the user side is determined according to the first interaction behavior of the user side on the historical multimodal content under different topic tags.

8. A multimodal content generation device, characterized in that: The device comprises: The data collection module is used to collect initial materials from multiple data sources based on the corresponding topic tags on the user side; A content generation module is used to generate multimodal content based on the initial material based on the content style corresponding to the user side; the multimodal content includes at least two of summary text, description image, description audio, and description video.

9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multimodal content generation method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the multimodal content generation method according to any one of claims 1 to 7 by executing the executable instructions.

Citation Information

Patent Citations

  • Multi-modal data generation method and device, computer equipment and storage medium

    CN119537511A

  • Multi-modal content generation method and related device

    CN119597943A