Multimedia content generation method and device, electronic equipment and storage medium

By extracting reference features from users' historical generated data, the problem of insufficient consistency and coherence of narrative elements in the generation of multimedia content by generative models is solved, thus achieving high-quality multimedia content generation and improving user experience.

CN121353451APending Publication Date: 2026-01-16DOUYIN VISION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511400393.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing generative models struggle to maintain good consistency and coherence in narrative elements across a series of generated results for the same book or similar books when generating multimedia content, especially lacking consideration for users' historical generation data.

Method used

By extracting reference features of narrative elements from historically generated data authorized by users, and combining these narrative elements with current features, prompts for multimedia content are generated to ensure the consistency and coherence of narrative elements during the generation process.

Benefits of technology

It achieves good consistency and coherence in narrative elements of multimedia content, and improves the coherence and adaptability of user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353451A_ABST
    Figure CN121353451A_ABST
Patent Text Reader

Abstract

The invention relates to a multimedia content generation method and device, electronic equipment and a computer readable storage medium, and relates to the technical field of computers. The multimedia content generation method can comprise the following steps: in response to a content generation request for a narrative text, determining narrative elements in the narrative text; determining a reference feature for the narrative element, wherein the reference feature is extracted from historical generation data which is authorized by a user and is associated with the user; based on the narrative elements and the reference features, determining prompt information used for generating the multimedia content; and generating the multimedia content according to the prompt information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and in particular, to a multimedia content generation method, a multimedia content generation apparatus, an electronic device, a computer readable storage medium, and a computer program product. BACKGROUND

[0002] With the rapid development of artificial intelligence generation technology, text-to-image (hereinafter referred to as "text-to-image") technology has become an important application field. At present, text-to-image technology has been widely integrated into various applications to optimize user experience. Specifically, users can convert input text into corresponding images through generative models (such as adversarial networks, diffusion models, etc.). Taking narrative text (such as e-books, web novels, etc.) as an example, with the help of text-to-image technology, matching visual images can be automatically generated for text content, so that users can intuitively obtain the corresponding pictures while reading the text content, thereby obtaining a more rich and immersive reading experience. SUMMARY

[0003] In view of this, the embodiments of the present disclosure provide a multimedia content generation method, a multimedia content generation apparatus, an electronic device, a computer readable storage medium, and a computer program product, thereby allowing high-quality multimedia content to be generated for narrative text by considering reference features extracted from historical generation data associated with a user in the multimedia content generation process, so that the multimedia content can maintain good consistency and coherence in related narrative elements.

[0004] According to a first aspect of the present disclosure, a multimedia content generation method is provided, comprising: determining a narrative element in a narrative text in response to a content generation request for the narrative text; determining a reference feature for the narrative element, the reference feature being extracted from historical generation data associated with a user and authorized by the user; determining prompt information for generating the multimedia content based on the narrative element and the reference feature; and generating the multimedia content according to the prompt information.

[0005] According to a second aspect of the present disclosure, a multimedia content generation apparatus is provided, comprising: a first determination module configured to determine a narrative element in a narrative text in response to a content generation request for the narrative text; a second determination module configured to determine a reference feature for the narrative element, the reference feature being extracted from historical generation data associated with a user and authorized by the user; a third determination module configured to determine prompt information for generating the multimedia content based on the narrative element and the reference feature; and a generation module configured to generate the multimedia content according to the prompt information.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to perform the generating method according to some embodiments of the present disclosure based on instructions stored in the at least one memory.

[0007] According to a fourth aspect of the present disclosure, a computer readable storage medium is provided, having stored thereon computer instructions which, when executed by a processor, implement the generating method according to some embodiments of the present disclosure.

[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, which, when running on a computer, causes the computer to implement the generating method according to some embodiments of the present disclosure.

[0009] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of the exemplary embodiments of the present disclosure with reference to the following drawings. BRIEF DESCRIPTION OF DRAWINGS

[0010] Embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be understood that the drawings described below only relate to some embodiments of the present disclosure, and do not limit the present disclosure. In the drawings:

[0011] Figure 1 A flowchart of a method of generating multimedia content according to some embodiments of the present disclosure is shown;

[0012] Figure 2 A flowchart of a method of generating multimedia content according to some embodiments of the present disclosure is shown;

[0013] Figure 3 A flowchart of a method of determining reference features according to some embodiments of the present disclosure is shown;

[0014] Figure 4 A flowchart of a method of generating multimedia content according to some embodiments of the present disclosure is shown;

[0015] Figure 5 A flowchart of a method of generating multimedia content according to some embodiments of the present disclosure is shown;

[0016] Figure 6 A flowchart of a processing procedure of an intelligent agent according to some embodiments of the present disclosure is shown;

[0017] Figure 7 And 8 A schematic diagram of a display interface according to some embodiments of the present disclosure is shown;

[0018] Figure 9A block diagram illustrating a multimedia content generation apparatus according to some embodiments of the present disclosure is shown;

[0019] Figure 10 A block diagram illustrating an electronic device according to some embodiments of the present disclosure is shown;

[0020] Figure 11 A block diagram illustrating an electronic device according to some other embodiments of the present disclosure is shown.

[0021] It should be understood that the sizes of the respective portions shown in the drawings are not necessarily drawn to scale. Identical or similar reference numerals are used to indicate identical or similar components throughout the specification. Therefore, once a component is defined in one drawing, it can not be discussed further in subsequent drawings. DETAILED DESCRIPTION

[0022] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein.

[0023] It should be understood that the respective steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specified, the relative arrangement of the components and steps set forth in these embodiments, numerical expressions, and numerical values should be interpreted as being merely exemplary, and do not limit the scope of the present disclosure.

[0024] The term “comprising” and variations thereof as used in the present disclosure mean an open-ended term that includes at least the recited elements / features, but does not exclude other elements / features, i.e., “including but not limited to”. The term “based on” means “based at least in part on”.

[0025] It should be noted that the terms “first”, “second”, and the like as mentioned in the present disclosure are merely used to distinguish different devices, modules, or units, and are not intended to limit the functions of the devices, modules, or units. Unless otherwise specified, the terms “first”, “second”, and the like are not intended to imply a given order or a mutual dependency in any other manner.

[0026] It should be noted that the adjectives “one”, “more than one” mentioned in the present disclosure are illustrative rather than limiting, and those skilled in the art should understand that “one” or “more than one” should be understood as “one or more” unless the context clearly indicates otherwise.

[0027] The user information and data involved in this disclosure (including but not limited to users' historical generated data) are all information and data authorized by users or fully authorized by all parties. The collection, use and processing of such information and data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0028] It should be understood that, in this disclosure, the corresponding controls may refer to human-computer interaction tools used in human-computer interfaces to implement input, output, or operation functions. For example, the corresponding controls may include labels, buttons, text boxes, sliders, menus, switches, etc. User triggering operations on the corresponding controls may include, for example, single-click, double-click, long-press, swipe, voice control, etc.

[0029] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0030] With the rapid development of artificial intelligence generative technology, text-to-image (hereinafter referred to as "text-to-image") technology has become an important application area. Currently, text-to-image technology has been widely integrated into various applications to optimize user experience. Specifically, users can use generative models (such as adversarial networks, diffusion models, etc.) to transform input text into corresponding images. Taking narrative text (such as e-books, online novels, etc.) as an example, text-to-image technology can automatically generate matching visual images for the text content, allowing users to intuitively obtain images corresponding to the content while reading, thus gaining a richer and more immersive reading experience.

[0031] However, in the process of generating multimedia content (e.g., images and / or videos) using generative models, a series of generated results for the same book or similar books may not maintain good consistency and coherence in terms of related narrative elements (e.g., character elements). For example, the physical characteristics and behavioral features of the same character may contradict each other in different generated results, disrupting the overall coherence of the content. Furthermore, the lack of consideration and adoption of the user's historical generation data in the generation process makes it difficult for multimedia content generated by a series of processes initiated by the same user to maintain good consistency and coherence in terms of related narrative elements. Therefore, how to use generative models to generate high-quality multimedia content that maintains good consistency and coherence in terms of related narrative elements is a technical problem that urgently needs to be solved by those skilled in the art. It should be noted that the information described in this section is intended to aid in understanding the background of this disclosure and may contain content not yet known to those skilled in the art; it should not be considered as a limitation of the prior art.

[0032] Based on this, this disclosure provides a method for generating multimedia content, an apparatus for generating multimedia content, an electronic device, a computer-readable storage medium, and a computer program product, thereby allowing the generation of high-quality multimedia content from narrative text, ensuring good consistency and coherence in related narrative elements. Therefore, the multimedia content generation method of this disclosure can effectively overcome one or more of the aforementioned deficiencies.

[0033] Figure 1 A flowchart illustrating a method for generating multimedia content according to some embodiments of the present disclosure is shown. Figure 1 As shown, the multimedia content generation method of this embodiment may include steps S101 to S104. The method of this embodiment can be executed by an electronic device, such as a computer and / or a terminal device. It should be understood that the method of this embodiment can be executed at least partially on the client side and also partially on the server side.

[0034] In step S101, in response to a content generation request for narrative text, narrative elements in the narrative text are determined; in step S102, reference features for the narrative elements are determined, the reference features being extracted from historical generation data authorized by the user and associated with the user; in step S103, based on the narrative elements and the reference features, prompt information for generating the multimedia content is determined; in step S104, the multimedia content is generated according to the prompt information.

[0035] The narrative text disclosed herein can be any form of electronic text. It should be noted that the sources and forms of such narrative texts are diverse: they can be content existing directly in electronic text form, or they can be electronic text converted from other types of multimedia content (such as audio content or printed text). Specifically, for audio content, known speech-to-text functions can be used to convert it into electronic text; for printed text, known character recognition functions can be used to convert it into electronic text.

[0036] In some embodiments, the narrative text may be text content within a plot text. Taking plot text as an example, the text content may be structured into one or more text blocks, each text block corresponding to at least one text paragraph or sentence, and can be independently selected by the user.

[0037] In some embodiments, the narrative text may be text content from a social platform (or social software). For example, status texts posted by users on social software platforms may contain daily narratives, emotional expressions, etc.; in addition, dialogue content between users within social software can also constitute the narrative text. The relevant text content may be structured into one or more text blocks, each text block corresponding to at least one text paragraph or sentence, and can be independently selected by the user.

[0038] In some embodiments, the narrative text may be text content on a news media platform. For example, the text content may be news reports published by news clients, opinion pieces expressed in news commentary articles, and brief news items pushed by information platforms. Related text content may be structured into one or more text blocks, each text block corresponding to at least one text paragraph or sentence, and can be independently selected by the user.

[0039] It should be understood that this disclosure does not limit how a content generation request for narrative text is generated. In some embodiments of this disclosure, the content generation request for narrative text may be initiated or launched in response to a related generation operation for the narrative text. In other embodiments of this disclosure, in response to a user's selection operation on the narrative text, operation controls for the narrative text may be displayed, and the content generation request for the narrative text may be initiated or launched in response to the triggering of the operation controls. Furthermore, this disclosure does not limit how the narrative text is obtained. In some embodiments of this disclosure, the corresponding narrative text may be obtained from an internal storage device. In other embodiments of this disclosure, the corresponding narrative text may be obtained from an external storage device.

[0040] In this disclosure, the determination or extraction of narrative elements in the narrative text can be performed automatically, for example, by relying on one or more machine learning models. In some embodiments, the determination or extraction of narrative elements in the narrative text can be achieved using a large language model, a basic model, or a parameter fine-tuning model. In some embodiments, the determination of narrative elements in the narrative text can also be directly achieved through an intelligent agent. Here, the intelligent agent can also be referred to as a robot, a digital human, or a virtual agent of a machine learning model. The intelligent agent can be implemented based on one or more machine learning models, such as a large language model or a basic model.

[0041] The narrative elements involved in this disclosure may include, but are not limited to, one or more of the following: character elements, object elements, scene elements, and style elements in a narrative text. In some embodiments, the character element in a narrative text should be broadly understood, encompassing various narrative subjects that drive the plot forward, including but not limited to human characters, as well as animal characters, fictional creatures, anthropomorphic characters, and other forms. In some embodiments, the object element in a narrative text should be broadly understood, encompassing various objects appearing in the narrative text, including but not limited to stationary and moving objects. In some embodiments, the scene element in a narrative text should be broadly understood, encompassing various spatial environments involved in the narrative text, such as the location where the narrative takes place, including both real and virtual scenes. In some embodiments, the style element in a narrative text should be broadly understood, encompassing various style types of the narrative text, such as common styles like fantasy, romance, science fiction, and suspense.

[0042] In this disclosure, the reference features for the narrative elements can be extracted or derived from user-authorized historical generation data associated with the user. Therefore, the reference features for the narrative elements can be configured to be considered and adopted as generation reference information during the generation of multimedia content, taking into account the user's historical generation characteristics. This ensures that for multiple rounds of multimedia content generation initiated by the same user, the resulting multimedia content maintains good consistency and coherence across the relevant dimensions corresponding to the narrative elements.

[0043] The historical generation data involved in this disclosure may include, but is not limited to, one or more of the following: user-approved historical multimedia content, historical narrative text used to generate the historical multimedia content, and historical user input information generated in response to the historical multimedia content. User-approved historical multimedia content can be understood as historical multimedia content on which the user has performed affirmative actions with a clear intention of approval. These affirmative actions may include, but are not limited to, the user's adoption, saving, sharing, or instruction confirmation of the historical multimedia content. Furthermore, the reference features of the narrative elements involved in this disclosure should be interpreted broadly; they may constitute reference features in text form (e.g., text features or text templates) or reference features in multimedia content form (e.g., multimedia content features or multimedia content templates).

[0044] Therefore, the determination of reference features for the narrative elements involved in this disclosure may include, but is not limited to, one or more of the following: determining multimedia content reference features for the narrative elements extracted from the historical multimedia content; determining text reference features for the narrative elements extracted from the historical narrative text; and determining user input reference features for the narrative elements extracted from the historical user input information. It is evident that the historical generation data involved in this disclosure covers multiple sources, including text content, multimedia content, and user input, enabling a more comprehensive adaptation to the needs for narrative element features in different application scenarios and effectively avoiding the problems of feature partiality and isolation caused by a single source. Thus, this disclosure can efficiently and comprehensively obtain reference features of narrative elements, providing support for the subsequent generation of high-quality, highly adaptable multimedia content. As an example, for some clearly quantifiable features, such as a person's gender, they can be extracted from the textual descriptions of "male" and "female" in historical narrative texts and stored as corresponding reference features in text form; while for some unquantifiable features, such as a human face or a complex sword, they can be stored as images or parts of images. When generating multimedia content later, they can be applied in a "padding image" manner to ensure that the generated multimedia content can maintain consistency and coherence.

[0045] Furthermore, this disclosure does not limit how the reference features for narrative elements are determined. In some embodiments of this disclosure, reference features for the narrative elements can be extracted in advance or offline. In this case, relevant narrative elements and reference features for the narrative elements have been extracted in advance from historical generation data associated with the user, and these extracted narrative elements and their reference features can be stored in a corresponding storage device in advance. Thus, in response to the determined narrative elements, suitable reference features can be determined, for example, automatically from the previously extracted reference features. In other embodiments of this disclosure, reference features for the narrative elements can be extracted and determined in real time or online. In this case, in response to the determined narrative elements, reference features for the narrative elements can be extracted and determined in real time or online from historical generation data associated with the user.

[0046] In some embodiments, the method for generating multimedia content may include a reference feature extraction step, which may be implemented, for example, in advance or offline. Specifically, the reference feature extraction step may include: in response to an approval operation for historical multimedia content, extracting reference features for the narrative elements from at least one of the historical multimedia content, historical narrative text used to generate the historical multimedia content, and historical user input information for the generation of the historical multimedia content. These extracted narrative elements and their reference features may be stored directly or processed (e.g., aggregated) in a corresponding storage device for use in subsequent generation processes.

[0047] In some embodiments, in response to an approval operation for historical multimedia content, extracting the reference features for the narrative element from the historical multimedia content may include: determining the presentation portion corresponding to the narrative element from the historical multimedia content; and extracting the reference features for the narrative element based on the presentation portion. As an example, the historical multimedia content may, for instance, involve a historical image generated based on a piece of text content. In this historical image, multiple characters may be presented; an exemplary scenario is "a male character and a female character walking along a riverbank," and the characters in the image possess identifiable visual features (such as clear facial contours, body shapes, etc.). Determining the presentation portion corresponding to the narrative element may, for instance, involve separately determining the visual presentation area directly corresponding to the "male character" (i.e., the image portion corresponding to the male character) and the visual presentation area directly corresponding to the "female character" (i.e., the image portion corresponding to the female character) in the historical image. For this purpose, image cropping techniques may be used, for example, to extract the image portions corresponding to the two characters from the historical image. After obtaining the image portions corresponding to the characters, the corresponding reference features can be extracted. For example, facial features can be extracted from the "male character image portion" and the "female character image portion" respectively, and these extracted facial features can be designated as reference features for the male character and the female character, respectively. This "pad image"-based implementation method can provide a unified feature reference for the subsequent generation of new multimedia content through pre-determined character elements and corresponding reference features, thereby effectively ensuring that the subsequently generated multimedia content maintains good consistency and coherence in terms of narrative elements.

[0048] In this disclosure, the series of steps related to determining the reference features of narrative elements can be performed automatically, for example, by relying on one or more machine learning models. In some embodiments, the determination or extraction of narrative elements in the narrative text can be achieved using a large language model, a base model, or a parameter fine-tuning model. In some embodiments, the determination of narrative elements in the narrative text can also be achieved directly by an intelligent agent. Here, the intelligent agent can also be referred to as a robot, a digital human, or a virtual agent of a machine learning model. The intelligent agent can be implemented based on one or more machine learning models, such as a large language model or a base model.

[0049] The reference features for the narrative elements disclosed herein may include, but are not limited to, one or more of the following: reference features for the character elements, reference features for the object elements, reference features for the scene elements, and reference features for the style elements. In some embodiments, the reference features for the character elements may include, but are not limited to, one or more of the following: facial features, skin tone, hairstyle, height, expression, and character template. This allows these accumulated reference features to be applied at least partially to the character element when it is included in subsequent raw images or videos. In some embodiments, the reference features for the object elements may include, but are not limited to, one or more of the following: shape, color, size, and detail. This allows these accumulated reference features to be applied at least partially to the object element when it is included in subsequent raw images or videos. In some embodiments, the reference features for the scene elements may include, but are not limited to, one or more of the following: size, structure, decoration, and atmosphere. This allows these accumulated reference features to be applied at least partially to the scene element when it is included in subsequent raw images or videos. In some embodiments, the reference features for the style element may include, but are not limited to, at least one of art style features, storyboard features, atmosphere features, and decorative features. This allows these precipitated reference features to be applied at least partially to the style element when it is included in subsequent raw images or videos.

[0050] It should be understood that the reference features for the aforementioned narrative elements should be interpreted broadly, and can also constitute a feature template composed of multiple reference features. More precisely, the feature template is not limited to the combination of reference features under a single narrative element (e.g., combining only "facial features + expression features" from the character element to form a character template), but can also involve the combination of reference features across narrative elements, such as combining "hairstyle features of the character element + atmosphere features of the scene element + art style features of the style element" to form a feature template adapted to a specific plot or chapter.

[0051] In this disclosure, the determination of prompts for generating the multimedia content based on the narrative elements and the reference features can be automated, for example, by relying on one or more machine learning models. In some embodiments, prompts for generating multimedia content can be determined based on the narrative elements using a large language model, a base model, or a parameter fine-tuning model. In some embodiments, the determination of prompts can be directly implemented by an intelligent agent. This intelligent agent can also be referred to as a robot, a digital human, or a virtual agent of a machine learning model. The intelligent agent can be implemented based on one or more machine learning models, such as a large language model or a base model.

[0052] In this disclosure, the prompts used to generate multimedia content can be understood as a comprehensive set of instructions, represented in text and / or multimedia content form, that can guide the multimedia content generation model to generate multimedia content. These instructions may include at least a description of the screen content of the multimedia content (e.g., including the narrative elements involved and the corresponding reference features).

[0053] In this disclosure, the generation of multimedia content based on the prompt information can be automated, for example, by relying on one or more machine learning models. In some embodiments, the prompt information can be converted into corresponding multimedia content using a generative model (e.g., adversarial networks, diffusion models, etc.). In some embodiments, the generation of multimedia content based on the prompt information can also be directly implemented by an intelligent agent. Here, the intelligent agent can also be referred to as a robot, a digital human, or a virtual agent of a machine learning model. The intelligent agent can be implemented based on one or more machine learning models, such as based on a large language model or a basic model.

[0054] The multimedia content generation method disclosed herein allows for the generation of high-quality multimedia content for narrative text by considering reference features extracted from historical generation data associated with users during the multimedia content generation process, thereby maintaining good consistency and coherence in terms of relevant narrative elements.

[0055] Figure 2 A flowchart illustrating a method for generating multimedia content according to some embodiments of the present disclosure is shown. Figure 2 As shown, the multimedia content generation method of this embodiment may include steps S201 to S204. The method of this embodiment can be executed by an electronic device, such as a computer and / or a terminal device. It should be understood that the method of this embodiment can be executed at least partially on the client side and also partially on the server side.

[0056] In step S201, in response to a content generation request for narrative text, narrative elements in the narrative text and current features of the narrative elements are determined; in step S202, reference features for the narrative elements are determined, the reference features being extracted from historical generation data authorized by the user and associated with the user; in step S203, based on the narrative elements, the reference features, and the current features, prompt information for generating the multimedia content is determined; in step S204, the multimedia content is generated according to the prompt information.

[0057] It should be noted that the generation method of this embodiment is the same as... Figure 1 The related parts of the generation method introduced can be found in the detailed content above, and will not be repeated here.

[0058] The current features for the narrative elements involved in this disclosure may include, but are not limited to, one or more of the following: current features for character elements, current features for item elements, current features for scene elements, and current features for style elements. In some embodiments, the current features for the character elements may include, but are not limited to, one or more of the following: facial features, skin color, hairstyle, height, expression, and character template. In some embodiments, the current features for the item elements may include, but are not limited to, one or more of the following: shape, color, size, and detail. In some embodiments, the current features for the scene elements may include, but are not limited to, one or more of the following: size, structure, decoration, and atmosphere. In some embodiments, the current features for the style elements may include, but are not limited to, at least one of the following: art style, storyboard, atmosphere, and decoration. It should be understood that the current features for the narrative elements should be interpreted broadly, and may also constitute a feature template composed of multiple features.

[0059] In this disclosure, in response to a content generation request for narrative text, not only are narrative elements in the narrative text determined, but also current features for those narrative elements are determined. The determined current features can then be combined with reference features as supplementary information to provide a more adaptive feature description. By considering relevant narrative elements, reference features for those narrative elements, and current features when determining the prompts used to generate the multimedia content, high-quality multimedia content can be generated for the narrative text, maintaining good consistency and coherence in relevant narrative elements while providing a richer and more adaptive content presentation.

[0060] Figure 3 A flowchart illustrating method steps for determining reference features according to some embodiments of the present disclosure is shown. Figure 3 As shown, the method steps for determining the reference feature may include steps S1021 to S1022.

[0061] In step S1021, multiple reference features for the narrative element are extracted from the historical generated data; in step S1022, the reference feature for the narrative element is determined from the multiple reference features.

[0062] In some embodiments, the historical generated data contains multiple historical multimedia contents and / or related historical narrative texts and / or historical user input information for the same narrative element. Typically, the same narrative element may also exhibit different characteristics as the narrative text unfolds. For example, for the same character, different visual representations may occur as the narrative text progresses, and therefore multiple historical images and corresponding historical text paragraphs may exist. Therefore, determining the reference features for the narrative element may also involve selecting a subset of suitable reference features from multiple reference features based on additional conditions.

[0063] In some embodiments, step S1022 may include determining the reference features for the narrative elements from the plurality of reference features based on the plot of the narrative text. This allows for the selection of a subset of reference features that are compatible with the plot from the plurality of reference features. This approach ensures that the selected reference features are highly compatible with the narrative's "plot context," avoiding a disconnect between features and the plot. For example, if the same character in the narrative text presents different states as the plot develops (e.g., "daily state" or "combat state"), the corresponding reference features, such as character templates, will also differ. In this case, based on the current plot (e.g., "combat plot"), a suitable "combat state character template" can be selected as the reference feature.

[0064] Alternatively or additionally, step S1022 may include determining the reference features for the narrative element from the plurality of reference features based on the chapter to which the narrative text belongs. This allows for the selection of a subset of reference features that are compatible with the chapter from among the multiple reference features. This approach utilizes chapter attributes to filter features, improving the compatibility between features and chapters. For example, if the same character in the narrative text presents different states as the chapter progresses, the corresponding reference features, such as character templates, will also differ. In this case, a suitable character template can be selected as the reference feature based on the current chapter.

[0065] Alternatively or additionally, step S1022 may include determining the reference features for the narrative element from the plurality of reference features based on the number of approval operations of the plurality of reference features. This allows for the selection of a subset of reference features with high approval rates from the plurality of reference features. This approach, based on user behavior, prioritizes user-preferred reference features, improving user satisfaction with subsequent generated results. These approval operations may include, but are not limited to, the user's adoption, saving, sharing, and confirmation of the historical multimedia content.

[0066] Alternatively or additionally, step S1022 may include determining reference features for the narrative element from the plurality of reference features based on the current features for the narrative element determined from the narrative text. This allows for the selection of a subset of reference features that match the current features from the plurality of reference features. This approach ensures that the reference features are adapted to the current feature description of the narrative element, avoiding a disconnect between the reference features and the actual state of the narrative element.

[0067] Alternatively or additionally, step S1022 may include determining the reference features for the narrative element from the plurality of reference features based on the user's input information. This allows for the selection of a subset of reference features that match the user's input from a plurality of reference features. This approach directly responds to the user's proactive needs, ensuring that the selected reference features meet the user's current expectations and improving user satisfaction with the subsequently generated results. User input may include, but is not limited to, information entered by the user through specific interactive areas (such as feature selection pop-ups, text input boxes, option selection areas, etc.). Typically, the user's most recent proactive input better reflects their current needs; therefore, the most recent proactive input can be prioritized as the basis for adaptation during the selection process.

[0068] Figure 4 A flowchart illustrating a method for generating multimedia content according to some embodiments of the present disclosure is shown. Figure 4 As shown, the multimedia content generation method of this embodiment may include steps S301 to S306. The method of this embodiment can be executed by an electronic device, such as a computer and / or a terminal device. It should be understood that the method of this embodiment can be executed at least partially on the client side and also partially on the server side.

[0069] In step S301, in response to a content generation request for narrative text, narrative elements in the narrative text are determined; in step S302, reference features for the narrative elements are determined, the reference features being extracted from historical generation data authorized by the user and associated with the user; in step S303, based on the narrative elements and the reference features, prompt information for generating the multimedia content is determined; in step S304, the multimedia content is generated according to the prompt information; in step S305, after the multimedia content is generated, the multimedia content is displayed; in step S306, in response to an approval operation for the multimedia content, the generated multimedia content and the narrative text are identified as historical generation data associated with the user.

[0070] It should be noted that the generation method of this embodiment is the same as... Figures 1 to 3 The related parts of the generation method introduced in this article can be found in the detailed content above, and will not be repeated here.

[0071] In this embodiment of the present disclosure, by granting approval to the currently generated and displayed multimedia content (e.g., user acceptance, saving, sharing, or confirmation of multimedia content), the current multimedia content, along with related narrative text and / or user input, can be promptly added to the historical generated data associated with the user. This allows the reference feature extraction step to be based on this latest data. Consequently, subsequent feature extraction not only better aligns with the user's latest needs, avoiding biases caused by outdated historical data, but also continuously enriches the user-associated historical generated data, providing real-time and effective data support for generating more accurate and user-expected multimedia content.

[0072] Figure 5 A flowchart illustrating a method for generating multimedia content according to some embodiments of the present disclosure is shown. Figure 5 As shown, the multimedia content generation method of this embodiment may include steps S401 to S406. The method of this embodiment can be executed by an electronic device, such as a computer and / or a terminal device. It should be understood that the method of this embodiment can be executed at least partially on the client side and also partially on the server side.

[0073] In step S401, a trigger control is displayed; in step S402, in response to a content generation request for the narrative text, narrative elements in the narrative text are determined; in step S403, reference features for the narrative elements are determined, the reference features being extracted from user-authorized historical generation data associated with the user; in step S404, in response to an activation operation for the trigger control, a prompt message for generating the multimedia content is determined based on the narrative elements and the reference features; in step S405, in response to a deactivation operation for the trigger control, the prompt message for generating the multimedia content is determined based on the narrative elements but no longer based on the reference features; in step S406, the multimedia content is generated according to the prompt message.

[0074] It should be noted that the generation method of this embodiment is the same as... Figures 1 to 4 The related parts of the generation method introduced in this article can be found in the detailed content above, and will not be repeated here.

[0075] In this embodiment of the present disclosure, by displaying a trigger control in the display interface, users can actively choose whether to activate the reference feature-based generation process. Specifically, in response to an activation operation on the trigger control, such as clicking, checking, or inputting an instruction, the prompt information for generating the multimedia content can be determined based on the narrative elements and the reference features. Conversely, in response to a deactivation operation on the trigger control, such as clicking, unchecking, or inputting an instruction, the prompt information for generating the multimedia content is determined based on the narrative elements but no longer on the reference features. Thus, on the one hand, users are aware of the existing functions of the application product; on the other hand, users can flexibly choose whether to activate these existing functions in a simple way, thereby improving the user experience.

[0076] Figure 6 A schematic diagram illustrating the processing flow of an intelligent agent according to some embodiments of the present disclosure is shown. For example... Figure 6 As shown, the reference feature extraction step is performed offline or online by a first intelligent agent. Specifically, the first intelligent agent can extract relevant narrative elements (e.g., character elements, and / or item elements, and / or scene elements, and / or style elements) and reference features for the narrative elements (e.g., character element reference features, and / or item element reference features, and / or scene element reference features, and / or style element reference features) from historical generated data associated with the user (including historical multimedia content, and / or historical narrative text, and / or historical user input information). These extracted narrative elements and their reference features can be stored directly or processed (e.g., aggregated) in a corresponding storage device for use in subsequent generation processes.

[0077] like Figure 6 As shown, after acquiring narrative text (e.g., a paragraph from a novel), a first intelligent agent or other intelligent agent extracts narrative elements and optional current features for those narrative elements from the narrative text. After extracting the narrative elements, the first intelligent agent or other intelligent agent can further determine reference features for those narrative elements based on the extracted narrative elements, more specifically, retrieve reference features adapted to the relevant narrative elements from a corresponding storage device. After determining the reference features for those narrative elements, the first intelligent agent or other intelligent agent uses the narrative elements, reference features, and optional current features to generate prompts for multimedia content. After determining the prompts for generating multimedia content, the first intelligent agent or other intelligent agent generates multimedia content based on those prompts.

[0078] Figure 7 and 8 A schematic diagram of a display interface 10 according to some embodiments of the present disclosure is shown. It should be understood that... Figure 7and 8 The displayed page can be the same as or a different page in the display interface 10. Taking narrative text or an e-book as an example... Figure 7 The displayed page may be the first display page or the narrative text page. Figure 8 The displayed page shown can be a second display page or a page displaying the generated results. It should be noted that the settings of display interface 10 are merely exemplary and not restrictive.

[0079] like Figure 7 As shown, the narrative text page of the display interface 10 displays multiple text blocks, each corresponding to at least one text paragraph or sentence, and each block can be selected independently by the user. The user can select at least one text block using appropriate operations, such as long-pressing, clicking, or dragging. In the illustrated embodiment, the selected text block is exemplarily marked with a border. It should be understood that the selected text block can be highlighted in any way that distinguishes it from other unselected text blocks.

[0080] In response to a user's selection of the narrative text, an operation control 11 (such as an operation control with an AI identifier) ​​for the narrative text can be displayed on the display interface 10. The user can trigger, for example, by clicking or long-pressing the corresponding operation control 11, thereby prompting the first machine learning model to acquire the narrative text and triggering a content generation request for the narrative text. In the illustrated embodiment, other operation controls for the narrative text can also be displayed on the display interface 10. Triggering these other operation controls can also activate other functions, which will not be elaborated further here.

[0081] like Figure 8 As shown, the generation result display page of the display interface 10 can display not only the generated multimedia content, but also the user input area 12, trigger control 13, and approval operation control 14 related to the content generation. Users can input various generation requirements in the user input area 12. The trigger control 13 can be configured to allow users to actively choose whether to activate the reference feature-based generation process. The approval operation control 14 can be configured to perform approval operations such as saving and sharing the generated multimedia content.

[0082] Figure 9 A block diagram of a multimedia content generation apparatus 6 according to some embodiments of the present disclosure is shown. The generation apparatus 6 is configured to perform any of the embodiments of the methods described above. Figure 9As shown, the generation device 6 includes: a first determining module 61, configured to determine narrative elements in the narrative text in response to a content generation request for the narrative text; a second determining module 62, configured to determine reference features for the narrative elements, the reference features being extracted from historical generation data authorized by the user and associated with the user; a third determining module 63, configured to determine prompt information for generating the multimedia content based on the narrative elements and the reference features; and a generation module 64, configured to generate the multimedia content according to the prompt information.

[0083] In some embodiments, the first determining module may be configured to: in response to a content generation request for narrative text, determine narrative elements in the narrative text and current features of the narrative elements.

[0084] In some embodiments, the third determining module may be configured to: determine prompt information for generating the multimedia content based on the narrative elements, the reference features, and the current features.

[0085] In some embodiments, the second determining module may be configured to: determine multimedia content reference features extracted from the historical multimedia content for the narrative elements; and / or determine text reference features extracted from the historical narrative text for the narrative elements; and / or determine user input reference features extracted from the historical user input information for the narrative elements.

[0086] In some embodiments, the second determining module may be configured to: determine character elements and reference features for the character elements; and / or determine item elements and reference features for the item elements; and / or determine scene elements and reference features for the scene elements; and determine style elements and reference features for the style elements.

[0087] In some embodiments, the generation apparatus 6 may include a reference feature extraction module (not shown), which may be configured to: in response to an approval operation for historical multimedia content, extract the reference features for the narrative elements from at least one of the historical multimedia content, historical narrative text used to generate the historical multimedia content, and historical user input information generated for the historical multimedia content.

[0088] In some embodiments, the reference feature extraction module may be configured to: determine the presentation portion corresponding to the narrative element from the historical multimedia content; and extract the reference feature for the narrative element based on the presentation portion.

[0089] In some embodiments, the second determining module may be configured to: extract a plurality of reference features for the narrative element from the historical generated data; and determine the reference feature for the narrative element from the plurality of reference features.

[0090] In some embodiments, the second determining module may be configured to: determine the reference feature for the narrative element from the plurality of reference features based on the plot of the narrative text; and / or determine the reference feature for the narrative element from the plurality of reference features based on the chapter to which the narrative text belongs; and / or determine the reference feature for the narrative element from the plurality of reference features based on the number of approval operations of the plurality of reference features; and / or determine the reference feature for the narrative element from the plurality of reference features based on the current feature for the narrative element determined from the narrative text; and / or determine the reference feature for the narrative element from the plurality of reference features based on the user's input information.

[0091] In some embodiments, the generation device 6 may include a display module (not shown), which may be configured to display the multimedia content after the multimedia content is generated. The reference feature extraction module may be configured to, in response to an approval operation on the multimedia content, determine the generated multimedia content and the narrative text as historical generation data associated with the user.

[0092] In some embodiments, the display module may be configured to display a trigger control. The third determining module is configured to: in response to an activation operation of the trigger control, determine the prompt information for generating the multimedia content based on the narrative elements and the reference features; and in response to a deactivation operation of the trigger control, determine the prompt information for generating the multimedia content based on the narrative elements but no longer based on the reference features.

[0093] The multimedia content generation apparatus provided in this disclosure allows for the generation of high-quality multimedia content for narrative text by taking into account reference features extracted from historical generation data associated with the user during the multimedia content generation process, thereby maintaining good consistency and coherence in terms of relevant narrative elements.

[0094] Figure 10 A block diagram of an electronic device 7 according to some embodiments of the present disclosure is shown. Figure 10As shown, the electronic device 7 includes: at least one memory 71; and at least one processor 72 coupled to the at least one memory 71, the at least one processor 72 being configured to execute any of the embodiments of the above methods based on instructions stored in the at least one memory 71.

[0095] Memory 71 is used to store one or more computer-readable instructions. Memory 71 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 111 may, for example, store operating systems, application programs, bootloaders, databases, and other programs, as well as various application programs and various data.

[0096] The processor 72 is configured to execute computer-readable instructions to implement the generation method described in any of the foregoing embodiments. Specific implementations of each step of the method can be found in the above embodiments; repeated details will not be elaborated upon here.

[0097] The aforementioned electronic device 7 of this disclosure, by taking into account reference features extracted from historical generation data associated with the user during the generation of multimedia content, allows for the generation of high-quality multimedia content for narrative text, enabling it to maintain good consistency and coherence in terms of relevant narrative elements.

[0098] The processor 72 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be based on x86 or ARM architectures, etc.

[0099] The processor 72 and the memory 71 can communicate with each other directly or indirectly. For example, the processor 72 and the memory 71 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 72 and the memory 71 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0100] It should be noted that Figure 10 The components of the electronic device 7 shown are merely exemplary and not limiting. The electronic device 7 may have other components depending on the specific application requirements. The processor 72 can control other components in the electronic device to perform desired functions.

[0101] Electronic device 7 can be implemented by software, firmware and / or hardware, and can be integrated into a device with the relevant application installed.

[0102] Figure 11 A block diagram of an electronic device 8 according to other embodiments of the present disclosure is shown.

[0103] Figure 11 The electronic device 8 shown can be a computer system with a dedicated hardware structure, which can perform corresponding functions when the relevant application is installed.

[0104] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet PCs, PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.

[0105] like Figure 11 As shown, the Central Processing Unit (CPU) 81 performs various processes based on programs stored in the Read-Only Memory (ROM) 82 or programs loaded from the Storage Section 88 into the Random Access Memory (RAM) 83. The RAM 83 stores data required as needed when the CPU 81 performs various processes. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. The ROM 82, RAM 83, and Storage Section 88 can be various forms of computer-readable storage media. It should be noted that although the ROM 82, RAM 83, and Storage Section 88 are shown separately in the figure, one or more of them can be combined or located in the same or different memories or storage modules.

[0106] CPU 81, ROM 82, and RAM 83 are interconnected via bus 84. Input / output interface 85 is also connected to bus 84.

[0107] The following components are connected to the input / output interface 85: input section 86, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 87, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 88, including hard disk, magnetic tape, etc.; and communication section 89, including network interface cards such as LAN cards, modems, etc. The communication section 89 allows communication processing to be performed via a network such as the Internet. It is readily understood that although some parts of the electronic device 8 shown in the figure communicate via bus 84, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.

[0108] As needed, drive 810 is also connected to input / output interface 85. Removable media 811, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 810 as needed, so that computer programs read from them can be installed into storage section 88 as needed.

[0109] When the above series of processes are implemented with the help of software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 811.

[0110] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to implement the generation method described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 89, or installed from storage section 88, or installed from ROM 82. When the computer program is executed by CPU 81, the methods of embodiments of this disclosure are performed.

[0111] The above method, by taking into account reference features extracted from historical generation data associated with users during the generation of multimedia content, allows for the generation of high-quality multimedia content for narrative texts, ensuring good consistency and coherence in related narrative elements.

[0112] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0113] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.

[0114] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the generation method described in any of the foregoing embodiments.

[0115] The above method, by taking into account reference features extracted from historical generation data associated with users during the generation of multimedia content, allows for the generation of high-quality multimedia content for narrative texts, ensuring good consistency and coherence in related narrative elements.

[0116] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0117] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0118] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the generation method described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.

[0119] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via an Internet service provider).

[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0121] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0122] While specific embodiments of this disclosure have been described in detail with the aid of examples, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. A method for generating multimedia content, comprising: determining a narrative element in a narrative text in response to a content generation request for the narrative text; determining reference features for the narrative element, the reference features being extracted from historical generation data associated with a user and authorized by the user; determining prompt information for generating the multimedia content based on the narrative element and the reference features; generating the multimedia content according to the prompt information.

2. The generation method of claim 1, wherein, The historical generation data comprises at least one of the following: historical multimedia content approved by the user, historical narrative text for generating the historical multimedia content, historical user input information for generation of the historical multimedia content.

3. The generation method of claim 2, wherein, Determining the reference features for the narrative element comprises at least one of the following: determining multimedia content reference features for the narrative element extracted from the historical multimedia content; determining text reference features for the narrative element extracted from the historical narrative text; determining user input reference features for the narrative element extracted from the historical user input information.

4. The generation method according to one of claims 1 to 3, wherein, The narrative element in the narrative text comprises at least one of the following: a character element, an article element, a scene element, and a style element.

5. The generation method of claim 4, wherein, Determining the reference features for the narrative element comprises at least one of the following: determining the character element and reference features for the character element; determining the article element and reference features for the article element; determining the scene element and reference features for the scene element; determining the style element and reference features for the style element. 6.The method of claim 5, wherein: The reference features for the character element comprise at least one of the following: facial features, skin color features, hairstyle features, height features, expression features, and character templates; The reference features for the article element comprise at least one of the following: shape features, color features, size features, and detail features; The reference features for the scene element comprise at least one of the following: size features, structure features, decoration features, and atmosphere features; The reference features for the style element comprise at least one of the following: painting style features, shot features, atmosphere features, and decoration features. 7.The method of any one of claims 1 to 3, wherein: Determining the narrative element in the narrative text in response to the content generation request for the narrative text comprises: determining the narrative element in the narrative text and current features for the narrative element in response to the content generation request for the narrative text; Determining the prompt information for generating the multimedia content based on the narrative element and the reference features comprises: determining the prompt information for generating the multimedia content based on the narrative element, the reference features, and the current features. 8.The method of claim 1, further comprising: In response to an approval operation for historical multimedia content, the reference feature for the narrative element is extracted from at least one of the historical multimedia content, historical narrative text used to generate the historical multimedia content, and historical user input information for generation of the historical multimedia content.

9. The generation method of claim 8, wherein, In response to an approval operation for historical multimedia content, the reference feature for the narrative element is extracted from at least one of the historical multimedia content, historical narrative text used to generate the historical multimedia content, and historical user input information for generation of the historical multimedia content. The presentation part corresponding to the narrative element is determined from the historical multimedia content. The reference feature for the narrative element is determined based on the presentation part.

10. The generation method of claim 1, wherein, Determining the reference feature for the narrative element includes: The plurality of reference features for the narrative element are extracted from the historical generation data. The reference feature for the narrative element is determined from the plurality of reference features.

11. The generation method of claim 10, wherein, Determining the reference feature for the narrative element from the plurality of reference features includes at least one of: The reference feature for the narrative element is determined from the plurality of reference features according to a plot of the narrative text; The reference feature for the narrative element is determined from the plurality of reference features according to a chapter to which the narrative text belongs; The reference feature for the narrative element is determined from the plurality of reference features according to a number of approval operations of the plurality of reference features; The reference feature for the narrative element is determined from the plurality of reference features according to a current feature for the narrative element determined from the narrative text; The reference feature for the narrative element is determined from the plurality of reference features according to the input information of the user.

12. The generation method of any one of claims 1 to 4, further comprising: displaying the multimedia content after the multimedia content is generated; in response to an approval operation for the multimedia content, determining the generated multimedia content and the narrative text as the historical generation data associated with the user.

13. The generation method of any one of claims 1 to 4, further comprising: displaying a trigger control; in response to an activation operation for the trigger control, determining the hint information for generating the multimedia content based on the narrative element and the reference feature; in response to a deactivation operation for the trigger control, determining the hint information for generating the multimedia content based on the narrative element but not based on the reference feature.

14. A multimedia content generation apparatus, comprising: a first determination module configured to determine a narrative element in a narrative text in response to a content generation request for the narrative text; a second determination module configured to determine a reference feature for the narrative element, the reference feature being extracted from historical generation data associated with a user and authorized by the user; a third determination module configured to determine hint information for generating the multimedia content based on the narrative element and the reference feature. A generating module configured to generate the multimedia content according to the prompt information.

15. An electronic device, comprising: a memory; and a processor coupled to the memory, the processor configured to perform the generating method according to any one of claims 1 to 13 based on instructions stored in the memory.

16. A computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the generating method according to any one of claims 1 to 13.

17. A computer program product which, when running on a computer, causes the computer to implement the generating method according to any one of claims 1 to 13.