Method and apparatus for generating interactive multimedia content, electronic device, and storage medium

By receiving user description information and generating content text, the complex problem of the interactive multimedia content generation process in the prior art is solved, and efficient and low-threshold content generation and increase diversity are achieved.

WO2025113271A1PCT designated stage expired Publication Date: 2025-06-05BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/133084
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-20
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In the prior art, the generation process of interactive multimedia content has a high threshold, and it is difficult to achieve efficient user-generated content, especially in content production of different genres.

Method used

By receiving the description information input by the user, content text is generated based on the description information and the target genre, including the description text of the plot content and the visually transformed description text, thereby generating interactive multimedia content, and displaying content text and interactive multimedia content.

Benefits of technology

It improves the efficiency of interactive multimedia content generation, lowers the creative threshold, enables ordinary users to participate in story creation and development, and increases the diversity and feasibility of content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133084_05062025_PF_FP_ABST
    Figure CN2024133084_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for generating interactive multimedia content, an electronic device, and a storage medium, and relates to the technical field of computers. The method of the present disclosure comprises: receiving description information inputted by a user; on the basis of the description information and a target genre, generating a content text, the content text comprising a description text of plot content and a description text of a visual transformation corresponding to the target genre; on the basis of the content text, generating interactive multimedia content; and displaying the interactive multimedia content and the content text, the interactive multimedia content presenting a visual transformation effect corresponding to the description text of the visual transformation.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, electronic device and storage medium for generating interactive multimedia content

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on the application with CN application number 202311597109.3 and application date November 27, 2023, and claims its priority. The disclosed content of the CN application is hereby introduced as a whole into this application. Technical Field

[0003] The present disclosure relates to the field of computer technology, and in particular to a method, device, electronic device, and storage medium for generating interactive multimedia content. Background Art

[0004] Interactive multimedia content is a type of content that integrates multiple elements such as images, sounds, and text, and can provide an interactive interface for its viewers (or users), such as games, interactive movies, and interactive TV series.

[0005] Interactive multimedia content is generally produced by professional designers and developers. It requires a series of processes such as plot creation, drawing by artists, coding by developers or organizing actors for filming, and post-processing to be completed. In addition, the production process of interactive multimedia content of different genres requires different professionals and production processes. Summary of the Invention

[0006] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0007] According to some embodiments of the present disclosure, a method for generating interactive multimedia content is provided, comprising: receiving descriptive information input by a user; generating content text based on the descriptive information and a target genre, wherein the content text includes descriptive text of plot content and descriptive text of visual transformation corresponding to the target genre; generating interactive multimedia content based on the content text; and displaying the interactive multimedia content and the content text, wherein the interactive multimedia content presents a visual transformation effect corresponding to the descriptive text of the visual transformation.

[0008] According to other embodiments of the present disclosure, a device for generating interactive multimedia content is provided, comprising: a receiving module configured to receive descriptive information input by a user; a content generation module configured to generate content text based on the descriptive information and a target genre, wherein the content text includes descriptive text of the plot content and descriptive text of the visual transformation corresponding to the target genre; a multimedia generation module configured to generate interactive multimedia content based on the content text; and a display module configured to display the interactive multimedia content and the content text, wherein the interactive multimedia content presents a visual transformation effect corresponding to the descriptive text of the visual transformation.

[0009] According to some embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the method for generating interactive multimedia content of any embodiment of the present disclosure based on instructions stored in the memory.

[0010] According to some embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for generating interactive multimedia content according to any embodiment of the present disclosure is performed.

[0011] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The drawings described herein are used to provide a further understanding of the present disclosure. Each of the drawings, together with the following detailed description, is included in this specification and forms a part of the specification to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation of the present disclosure. In the drawings:

[0013] FIG1 is a schematic flow chart showing a method for generating interactive multimedia content according to some embodiments of the present disclosure;

[0014] FIG2 is a schematic diagram illustrating an authoring interface according to some embodiments of the present disclosure;

[0015] FIG3 is a schematic flow chart showing a method for generating interactive multimedia content according to other embodiments of the present disclosure;

[0016] FIG4 is a schematic diagram showing an authoring interface according to some other embodiments of the present disclosure;

[0017] FIG5 is a schematic diagram showing an authoring interface according to yet other embodiments of the present disclosure;

[0018] FIG6 is a schematic diagram showing a screen of interactive multimedia content according to some embodiments of the present disclosure;

[0019] FIG7 shows a schematic diagram of an authoring interface according to yet other embodiments of the present disclosure;

[0020] FIG8 is a schematic structural diagram of an apparatus for generating interactive multimedia content according to some embodiments of the present disclosure;

[0021] FIG9 is a schematic structural diagram of an electronic device according to some embodiments of the present disclosure;

[0022] FIG10 is a schematic diagram showing the structure of a computer system according to some embodiments of the present disclosure.

[0023] It should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not necessarily drawn to scale. The same or similar reference numerals are used throughout the drawings to indicate the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. DETAILED DESCRIPTION

[0024] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. However, it is obvious that the embodiments described are only some embodiments of the present disclosure, rather than all embodiments. The following description of the embodiments is actually only illustrative and is in no way intended to limit the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein.

[0025] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions and numerical values ​​of the parts and steps set forth in these embodiments should be interpreted as being merely exemplary and do not limit the scope of the present disclosure.

[0026] As used in this disclosure, the term "include" and its variations are intended to be open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to." Furthermore, the term "comprise" and its variations are intended to be open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to." Therefore, "include" and "include" are synonymous. The term "based on" means "based, at least in part, on."

[0027] Reference throughout this specification to "one embodiment," "some embodiments," or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. For example, the term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Furthermore, the appearances of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may.

[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, concepts such as "first" and "second" are not intended to imply that the objects described in such a manner must be in a given order in time, space, ranking, or any other manner.

[0029] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0030] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0031] The following detailed description of the embodiments of the present disclosure is provided in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0032] Currently, interactive multimedia content is professionally produced content (PGC). Its creation process is challenging and often involves collaboration among professionals from multiple fields. It is difficult to produce interactive multimedia content using user-generated content (UGC), resulting in relatively low efficiency. Because different genres of interactive multimedia content have different presentation formats, it is even more difficult to efficiently create diverse interactive multimedia content using UGC.

[0033] In order to improve the efficiency of generating interactive multimedia content of different genres, the present disclosure provides a method, device, electronic device and storage medium for generating interactive multimedia content. The following describes the method for generating interactive multimedia content of the present disclosure with reference to Figures 1 to 7.

[0034] Figure 1 is a flow chart of some embodiments of the method for generating interactive multimedia content disclosed herein. As shown in Figure 1 , the method of this embodiment includes steps S102 to S108.

[0035] In step S102, description information input by the user is received.

[0036] For example, an interactive multimedia content creation interface is displayed, and an input area is displayed in the creation interface to receive descriptive information input by the user. The descriptive information may include at least one of text, voice, and image. If the descriptive information is voice, voice recognition may be performed to convert the voice into text. If the descriptive information is an image, image recognition may be performed to convert the image into text.

[0037] The descriptive information input by the user includes key information that can be used to describe the interactive multimedia content to be generated. For example, the descriptive information may include at least one of the following: plot description information, character (or role) description information, scene description information, and soundtrack description information. Different types of descriptive information can be input through different interfaces. For example, the descriptive information may describe the story background of the interactive multimedia content to be generated, the protagonist and related characters, the general context and plot of the story, and the final ending.

[0038] In step S104, a content text is generated according to the description information and the target genre.

[0039] The target genre (or target media type) is the genre (or media type) of the interactive multimedia content finally generated. In some embodiments, the target genre includes at least one of: a game, an interactive TV series, and an interactive movie, without being limited to the examples.

[0040] For example, a generative model is used to generate content text based on descriptive information and a target genre. The generative model is used to output target content based on input information. The input information includes the processing basis of the generative model during the generation process, such as which information is referenced to perform the generation process, the requirements for the target output content, etc. Generative models include, for example, models that generate content based on text or images, and the output of the generative model may include text, images, or a combination of the two. Of course, the input or output of the generative model may also be data in other modalities, such as audio, video, or a combination of multiple types of data. The generative model can be a single-modality model, such as a model that generates text from text (referred to as a "text-to-text model") or a model that generates images from images (referred to as a "image-to-image model"); or a cross-modal model, i.e., a model whose input and output belong to different modalities, such as a model that generates images from text (referred to as a "text-to-image model"). Alternatively, the input of the generative model may include multiple modalities, and the output may also include multiple modalities.

[0041] The content text is used to describe the plot, presentation format, etc. of the interactive multimedia content to be generated. In some embodiments, the content text includes at least one of a description text of the plot content and a description text of the visual transformation corresponding to the target genre.

[0042] For example, the description text of the plot content includes the plot information expanded based on the description information in the user input using a generative model. Compared with the description information input by the user, the story content is richer, with more details, and is closer to the plot information in the real script.

[0043] In some embodiments, visual transformations include at least one of camera movements, background switching, and character pose changes. For example, games have relatively fixed screens, character poses are relatively simple, and camera movements are rare or absent. Interactive television series have a slower narrative pace, so shots may last longer, camera changes are less frequent, and the variety of styles is limited. Similarly, character poses and backgrounds last longer. Interactive movies have a faster narrative pace, camera changes are more frequent, and the variety of styles is more diverse, resulting in more dynamic characters and backgrounds. Therefore, different descriptions of visual transformations can be generated for different target genres.

[0044] In step S106, interactive multimedia content is generated according to the content text.

[0045] For example, a generative model can be used to generate interactive multimedia content based on content text. This generative model can be a text-to-image model or a text-to-video model. Because the generated content already includes descriptions of the plot and visual transformations, corresponding images, characters, and sounds can be generated based on this text, creating the desired visual effects.

[0046] In step S108 , the interactive multimedia content and the content text are displayed.

[0047] The interactive multimedia content and the content text can be displayed simultaneously, for example, in different areas of the interface, and displayed in a corresponding (or associated) manner. For example, based on the progress of the interactive multimedia content display, the text can be scrolled. The interactive multimedia content and the content text can also be displayed separately in different interfaces. For example, the interactive multimedia content is displayed first, and in response to a triggering operation for displaying the content text, the content text is displayed in a floating layer, a new window, or a new page.

[0048] In some embodiments, the interactive multimedia content may be divided into multimedia content of one or more plot nodes, and the content text includes the content text of each plot node. The multimedia content of each plot node is associated with the content text of the plot node and displayed.

[0049] Steps S104 and S106 can be implemented on the client or server. In the case of server-side implementation, the client sends the description information to the server and receives and displays the content text and interactive multimedia content returned by the server. The display-related steps in this disclosure can all be implemented using a graphical user interface (GUI).

[0050] The method of the above embodiment can automatically generate content text based on the description information and target genre input by the user, and automatically generate interactive multimedia content based on the content text. A description text with a visual transformation that matches the target genre can also be generated, so that the generated interactive multimedia content presents a visual transformation effect corresponding to the target genre. Through the method of the above embodiment, the user only needs to simply input the description information to generate the expanded content text, and further generate interactive multimedia content of different genres, thereby improving the efficiency of creating interactive multimedia content and saving resources. Ordinary users can also participate in the creation and development of stories to better interact with the characters and create their own unique stories, which improves the feasibility of creating interactive multimedia content and can convert more creative ideas into interactive multimedia products.

[0051] In some embodiments, the description information is parsed to determine the genre (or media type) corresponding to the description information as the target genre; or, a user selection operation of one or more genres among multiple genres is received, and the one or more genres are determined as the target genre.

[0052] For example, the description information may include keywords corresponding to the genre of the interactive multimedia content that the user wishes to create, such as games, movies, TV series, etc. These keywords can be identified to determine the target genre. For another example, the content of the description information of different genres may be different. Games contain more interactive plots and simpler plots, while interactive TV series have longer plots and more plots. Therefore, a machine learning model can be used to parse and understand the description information, determine the genre corresponding to the description information, and use it as the target genre. The machine learning model can be an AI (Artificial Intelligence) model, a large model, etc. Specifically, it can be a neural network, a deep learning model, a generative model, etc.

[0053] After the description information is parsed to determine the genre corresponding to the description information, the determined genre may be displayed, and in response to a confirmation operation by the user, the confirmed genre may be used as a target genre.

[0054] Users can operate in the client application (APP) or web page to create interactive multimedia content. For example, in response to the opening of the application, the homepage is displayed, and creation controls and guidance information can be set on the homepage. The guidance information, for example, is "Create your interactive story". In response to the triggering of the creation control, the creation interface is entered. The homepage or creation interface can include selection controls for multiple genres. The user can select one or more genres and determine one or more genres as target genres. If the user selects multiple genres, multiple interactive multimedia contents can be generated respectively.

[0055] In some embodiments, the description information may include at least one of description information of the plot, description information of the characters (or roles), description information of the scenes, and description information of the soundtrack.

[0056] For example, plot description information includes: plot summary information and at least one item of plot information for each of one or more plot nodes. The plot summary information can be used to describe the story content of the interactive multimedia content to be generated. For example, the plot summary information can include a brief description of at least one of the story's overall background, beginning, process, and ending. The story's process and ending can have various scenarios and can be described separately.

[0057] The authoring interface can be distinguished into different panels. For example, as shown in FIG2 , the authoring interface includes a story panel, a character panel, and the like. In response to a user triggering an operation on the story panel, the story panel can be displayed, and plot description information can be entered in the story panel. For example, a story content input area can be displayed in the story panel to receive plot summary information entered by the user. For example, the story content input area can be used to enter the story's global background and / or beginning. An ending (or success condition) input area can also be provided for entering the story's ending (or success condition).

[0058] The overall plot of the interactive multimedia content to be generated can be divided into one or more plot nodes (also referred to as plots or chapters). In some embodiments, one or more plot nodes and plot information for each plot node can be generated based on the plot summary information. A text generation model can be used to split the plot summary information to generate plot information for each plot node, and then a descriptive text of the plot content of each plot node can be generated based on the plot information of each plot node. That is, the generated descriptive text of the plot content includes the descriptive text of the plot content of each plot node in the one or more plot nodes.

[0059] The plot description information input by the user may also include plot information for each of one or more plot nodes. The story panel may distinguish plot nodes (also referred to as episodes or chapters) and receive user input separately. The plot information for each plot node is used to describe the story content of each plot node. The story content of each plot node may include a brief description of at least one of the node's background, beginning, process, and ending.

[0060] As shown in Figure 2, a story panel is displayed, and a plot adding control 201 is displayed in the story panel for adding one or more plot nodes, and the plot information of each plot node input by the user is received through the input area of ​​each plot node. The input area of ​​each plot node can include a story content description area 202 of the plot node, an ending description area 203 of the plot node, and can also include at least one of a speaking role selection area 204, a music selection (upload) area 205, and a picture selection (upload) area 206. Corresponding guidance information can be set in different areas to guide user input. The music or pictures selected or uploaded by the user can be used as background music or background (scene) pictures to generate multimedia content corresponding to the plot node. If these specific information are not included in the description information, it can be automatically generated by a generative model.

[0061] Other embodiments of the method for generating interactive multimedia content disclosed herein are described below with reference to FIG. 3 .

[0062] Figure 3 is a flow chart of another embodiment of the method for generating interactive multimedia content disclosed herein. As shown in Figure 3 , the method of this embodiment includes steps S302 to S310.

[0063] In step S302, description information input by the user is received.

[0064] In step S304 , the description information is parsed using a text generation model to generate a description text of the plot content of each of the one or more plot nodes.

[0065] If the description information only includes plot summary information, a text generation model can be used to first generate plot information for each of one or more plot nodes based on the plot summary information, and then generate a description text of the plot content of each plot node based on the plot information of each plot node.

[0066] If the description information only includes the plot information of each plot node, a text generation model may be used to generate a description text of the plot content of each plot node based on the plot information of each plot node.

[0067] If the description information includes both plot summary information and plot information of each plot node, a text generation model may be used to combine the plot summary information and the plot information of each plot node to generate a description text of the plot content of each plot node.

[0068] In some embodiments, a text generation model is used to generate a descriptive text for each plot node based on the descriptive information and the target genre. The descriptive text for each plot node can be expanded to varying degrees based on the target genre. For example, if the target genre is a game, the descriptive text for each plot node can be relatively simple. If the target genre is an interactive TV series, the descriptive text for each plot node can be more complex.

[0069] In some embodiments, a text generation model is used to generate descriptive text for each plot node based on the descriptive information and plot generation prompt information. The plot generation prompt information is used to assist in generating descriptive text for each plot node. For example, the plot generation prompt information may include prompt information related to the plot type and target genre. The plot type may be determined based on the descriptive information. For example, the prompt information for an adventure game may include examples of adventure plots, while the prompt information for an action game may include examples of character combat.

[0070] In some embodiments, when the plot description information includes plot information for multiple plot nodes, the user-entered description information may further include the logical relationships between the multiple plot nodes. The logical relationships between the multiple plot nodes are used to determine the order of the plot nodes, the conditions for moving from one plot node to the next, and the like.

[0071] Interactive operation nodes can be a special type of plot node, including plots that require players (or viewing users or operating users) to interact during the development of the plot. Different interactive operation nodes can guide the plot to different branches. In some embodiments, the description information also includes: description information of the interactive operation node and / or the logical relationship between the plot node and the interactive operation node.

[0072] If the user-entered description information does not include the logical relationships between multiple plot nodes, description information of interactive operation nodes, or the logical relationships between plot nodes and interactive operation nodes, this information can be automatically generated by the generative model based on the user-entered description information. When ultimately generating interactive multimedia content, the multimedia content for each plot node and each interactive operation node can be integrated based on the logical relationships between multiple plot nodes and the logical relationships between plot nodes and interactive operation nodes.

[0073] In some embodiments, a setting interface is displayed to receive a user's setting of a logical relationship between a plot node and an interactive operation node, and the set logical relationship between the plot node and the interactive operation node is displayed.

[0074] The settings interface may display initial settings information automatically generated based on the description information entered by the user. For example, the initial settings information may include each plot node, each interactive operation node, and the logical relationships between the plot nodes and the interactive operation nodes. The user may adjust the initial settings information. The user may also directly set each plot node, each interactive operation node, and the logical relationships between the plot nodes and the interactive operation nodes in the settings interface.

[0075] As shown in Figure 4, the story setting interface is displayed in the form of a flowchart, which makes it easier for users to clearly and accurately set the logical relationship between plot nodes and interactive operation nodes, but is not limited to the examples given. The flowchart includes multiple plot nodes 401 and multiple interactive operation nodes 402, and the lines between the nodes represent the logical relationship between them. Each node can be displayed with a node (plot or chapter) name or introduction. The flowchart can be automatically generated by a generative model based on the descriptive information input by the user, and the user can adjust or confirm it. The flowchart can also be generated based on the plot information of one or more plot nodes input by the user (for example, as shown in Figure 2) and / or the descriptive information of one or more interactive operation nodes. Or the flowchart can be directly configured by the user.

[0076] As shown in Figure 4, interactive operation nodes are used to configure plots that can be interacted with by players (or viewing users or operating users). Different plot branches can be reached through different interactive operation nodes. For plot nodes or interactive operation nodes, not only can they be displayed in the form of a flow chart, but also, in response to the user's configuration or selection of each plot node or interactive operation node, the generated descriptive text of the plot content of any plot node and / or the descriptive text of any interactive operation node, as well as the generated multimedia content of any plot node and / or the multimedia content of any interactive operation node can be displayed.

[0077] In some embodiments, a text generation model is used to parse the description information to generate a description text for each interactive operation node, wherein the description text for each interactive operation node includes a description text of the interactive operation options displayed to the operating user.

[0078] As shown in FIG4 , for the interactive operation node “Fork in the Road”, the content interface 403 of the interactive operation node can be displayed based on the user’s configuration or selection. The content interface can be displayed in the form of a window or a floating layer corresponding to the node. The content interface can include a description text of the interactive operation node. For example, in the plot of the interactive operation node, the astronauts come to a fork in the road, and astronaut x says, “This looks like a fork in the road. Which direction should we go next?” etc. Furthermore, the description text of each interactive operation option can also be displayed, for example, A. Explore the mountains, B. Explore the basin… This allows users to see the description text corresponding to each node very intuitively and accurately, and can adjust it.

[0079] The description text of each interactive operation node can be part of the description text of the plot content. Using a text generation model to parse the description information can not only generate the description text of the plot content of each plot node, but also generate the description text of each interactive operation node.

[0080] The design of characters (or roles) in interactive multimedia content is a crucial component, and their speeches or dialogues are also crucial components of the plot. For example, the generated plot description text for each plot node can include speeches by different characters. These speeches need to be determined based on the plot development, as well as the characters' characteristics, relationships, and appearances.

[0081] The character description text can be generated by a generative model based on the description information. The character description text may include at least one of the following: name, personality, identity, relationship between characters, speaking style, and voice characteristics, but is not limited to the examples given. In some embodiments, the description information input by the user includes character setting information for one or more characters. The character setting information for each character may also include at least one of the following: name, personality, identity, relationship between characters, speaking style, voice characteristics, and image, but is not limited to the examples given. The character setting information is more concise than the character description text. For example, the description information is parsed using a text generation model to generate the description text for each character.

[0082] The creation interface may include a character panel, receive a user's trigger operation on the character panel, display the character panel, and enter the setting information of one or more characters in the character panel. As shown in Figure 5, the character panel is displayed, and a role adding control 501 is displayed in the character panel, which is used to add one or more characters (or roles), and the setting information of the characters input by the user is received through the input area of ​​each character. The input area of ​​each character may include a name setting area 502 of the character, a basic setting area 503 of the character, and may also include at least one of a line style setting area 504, a character music setting area 505, a character avatar setting area 506, and a character standing picture setting area 507. Corresponding guidance information can be set in different areas to guide user input. The input area of ​​each character in Figure 5 can display the initial description text generated based on the description information, and the user can adjust it.

[0083] In some embodiments, a text generation model is used to generate a description text for each character based on the description information and character generation prompt information. The character generation prompt information can correspond to different character types. For example, the character generation prompt information includes language examples of different types of characters.

[0084] The description information entered by the user may also include at least one of scene setting information and visual style setting information. These information can also assist in generating a description of the plot content. Scene setting information can, for example, describe the background scenery. Visual style setting information can, for example, be anime style or Chinese style.

[0085] Through the method of the above embodiment, based on the description information input by the user, a description text of the plot content can be automatically generated, which improves the generation efficiency of interactive multimedia content, lowers the creation threshold, and enables ordinary users to freely create interactive multimedia content.

[0086] In step S306, a text generation model is used to determine the visual transformation information between different pictures in each plot node according to the description information and the target genre, and a description text of the visual transformation corresponding to each plot node is generated.

[0087] In some embodiments, the description information includes description information of visual transformations and / or description information of visual transformations corresponding to each plot node. For example, the description information of visual transformations may include general description information such as using a close-up shot when a character speaks or acts, using a zoom-in approach when entering a plot or chapter, using body movements when the character speaks, and background changes over time. The description information of visual transformations corresponding to each plot node can be more specific. For example, if a character is speaking in a plot node, the description information of visual transformations corresponding to the plot node may include, in combination with the plot information, the character talking while running, close-ups of the character, and background changes as the character runs.

[0088] If the description information includes description information of visual transformation and / or description information of visual transformation corresponding to each plot node, a text generation model can be used to expand the description information and the target genre to determine the information of visual transformation between different pictures in each plot node.

[0089] Ordinary users may not be adept at describing visual transformations. Therefore, descriptions often do not include visual transformation information and / or descriptions of the visual transformations corresponding to each plot node. A text generation model can be used to directly determine the visual transformation information between different frames at each plot node based on the description information and the target genre.

[0090] In some embodiments, a text generation model is used to determine the visual transformation between different frames in each plot node based on the description information, the target genre, and visual transformation hint information. This model then generates a text description of the visual transformation corresponding to each plot node. The visual transformation hint information is used to assist in generating the text description of the visual transformation. For example, the visual transformation hint information may include hint information corresponding to the target genre, such as examples of shot usage in an interactive movie.

[0091] The descriptive text for the plot content and the visual transformation for each plot node can be generated simultaneously using a text generation model. The descriptive text for the plot content and the visual transformation can be contained within the same text; they are not completely separate. For example, the descriptive text for the visual transformation can be interspersed with the plot content.

[0092] The information on visual transformation between different pictures may specifically include visual effect information of each picture.

[0093] In some embodiments, the visual transformation includes the movement of the lens. The information for determining the visual transformation between different frames in each plot node based on the descriptive information and the target genre includes: determining whether to use different types of lens effects in different frames in each plot node based on the descriptive information and the target genre; and for plot nodes that are determined to use different types of lens effects, determining the target lens effect used in each frame in the plot node and the target object corresponding to the target lens effect based on the descriptive information, as the information for the movement of the lens between different frames in the plot node.

[0094] For simple games, camera changes may not be necessary at every plot point (episode or chapter), whereas for interactive TV series or movies, camera changes are required. Alternatively, camera changes may be required at some plot points but not at others. For plot points that do not require camera changes, the camera operation information for those that do not require camera changes can be such that the camera effects remain unchanged.

[0095] Lens effects can be categorized as close-up, full-view, medium-view, and long-view based on visual distance, and as push, pull, pan, rise, and fall based on the lens's motion. For each image (or frame), the target lens effect applied can be independent, such as a close-up. It can also be linked to other images, such as a gradually rising lens effect across multiple consecutive images. The target object of a target lens effect can be a specific person, object, or panoramic view.

[0096] In some embodiments, determining the target lens effect used in each frame of the plot node and the target object corresponding to the target lens effect based on the descriptive information includes: generating plot information for each plot node or descriptive text of the plot content of each plot node based on the descriptive information; and for plot nodes that use different types of lens effects, determining the target lens effect used in each frame of the plot node and the target object corresponding to the target lens effect based on the plot information of the plot node or the descriptive text of the plot content of the plot node. The plot information of each plot node or the descriptive text of the plot content of each plot node includes the plot expanded using a text generation model based on the descriptive information. Because the plot affects the lens effect, for example, when characters speak, close-ups are more likely to be used. Therefore, the target lens effect and target object determined based on the plot information of each plot node or the descriptive text of the plot content of each plot node are more accurate and have richer effects.

[0097] In some embodiments, the visual transformation includes a change in a character's posture. Determining the information of the visual transformation between different frames in each plot node based on the description information and the target genre includes: determining whether to change the character's posture in each plot node based on the description information and the target genre; and for a plot node that determines the character's posture, determining the target posture of each of one or more characters in each frame of the plot node based on the description information, as information on the change in posture of each character between different frames in the plot node.

[0098] The posture of a character (or role) can include at least one of posture, action, and expression. For simple games, most plot nodes may not require a character's posture change, or the changes in the character's posture are very simple. However, for interactive TV series or interactive movies, the changes in character posture are more complex. The description information input by the user may also include information such as the character's posture, action, and expression. If it does not include this information, a generative model can be used to directly generate it. Therefore, combining the description information and the target genre can determine the target posture of the character in each scene at each plot node.

[0099] In some embodiments, based on the descriptive information, plot information for each plot node or a descriptive text of the plot content of each plot node is generated. For plot nodes in which a character's posture changes, a target posture for each character in that plot node is determined based on the plot information of that plot node or the descriptive text of the plot content of that plot node. The plot information of each plot node or the descriptive text of the plot content of each plot node includes a plot expanded using a text generation model based on the descriptive information. Because the plot affects the character's posture, for example, a character may display a successful expression and action when completing a task. Therefore, the target posture determined based on the plot information of each plot node or the descriptive text of the plot content of each plot node is more accurate and produces a richer effect.

[0100] The description information input by the user also includes description information of the character's posture. The description information of the character's posture may be relatively simple, and a richer description text can be generated through the generative model.

[0101] In some embodiments, the visual transformation includes switching of the background. The information for determining the visual transformation between different pictures in each plot node based on the description information and the target genre includes: determining whether to switch the background in each plot node based on the description information and the target genre; and for the plot node where the background switching is determined, determining the target background for each picture in the plot node based on the description information as the background switching information between different pictures in each plot node.

[0102] For some simple games, the background may not need to be switched at each plot node (episode or chapter), while for interactive TV series or interactive movies, the background switching is more frequent. Alternatively, the background may need to be switched at some plot nodes but not at others.

[0103] For each picture (or each frame), the target background used in the picture can be independent or associated with other pictures. For example, the background is continuously changed according to the movement of the character in multiple consecutive pictures.

[0104] In some embodiments, determining the target background used in each frame of the plot node based on the descriptive information includes: generating plot information for each plot node or descriptive text of the plot content of each plot node based on the descriptive information; and for plot nodes that switch backgrounds, determining the target background for each frame of the plot node based on the plot information of the plot node or the descriptive text of the plot content of the plot node. The plot information of each plot node or the descriptive text of the plot content of each plot node includes the plot expanded using a text generation model based on the descriptive information. Because the plot affects the background, determining the target background based on the plot information of each plot node or the descriptive text of the plot content of each plot node is more accurate and has a richer effect.

[0105] The description information input by the user also includes description information of the background switching. The description information of the background switching may be relatively simple, and a richer description text can be generated through the generative model.

[0106] In some embodiments, a text generation model is used to generate text describing the visual transformation of each plot node based on the description information and visual transformation prompt information. The visual transformation prompt information is used to assist in generating the text describing the visual transformation of each plot node. For example, the visual transformation prompt information may include prompt information related to the target genre. For example, for interactive movies or interactive TV series, close-up shots are more likely to be used when characters have long dialogues.

[0107] Through the method of the above embodiment, corresponding description text can be generated for different types of visual transformations based on the description information and the target genre, so that the generated interactive multimedia content can match the target genre more accurately and improve the generation efficiency of the interactive multimedia content.

[0108] In step S308, interactive multimedia content is generated according to the content text.

[0109] Interactive multimedia content can be generated based on the content text using a text-based graph and / or text-based video model. The generated interactive multimedia media may include: complete interactive multimedia content, multimedia content for each plot node, multimedia content for each interactive operation node, multimedia content for each character, multimedia content for each scene, etc. Complete interactive multimedia content can be generated based on the multimedia content for each plot node, multimedia content for each interactive operation node, multimedia content for each character, multimedia content for each scene, as well as the logical relationships between the plot nodes and interactive operation nodes. Transition effects can be added between different nodes.

[0110] In some embodiments, multimedia content for each plot node is generated based on the descriptive text of the plot content of each plot node and the descriptive text of the visual transformation corresponding to each plot node, wherein the multimedia content of each plot node presents a visual effect corresponding to the descriptive text of the visual transformation corresponding to the plot node.

[0111] The description text of the plot content and the description text of the visual transformation can be a whole paragraph of text interspersed together, described in natural language, and can be directly used to generate multimedia content.

[0112] In some embodiments, when the descriptive text of the visual transformation corresponding to each plot node includes the target lens effect used in each picture in each plot node and the target object corresponding to the target lens effect, each picture of each plot node is generated based on the descriptive text of the plot content of each plot node and the target lens effect used in each picture in each plot node and the target object corresponding to the target lens effect, wherein the target object in each picture is displayed using the corresponding target lens effect.

[0113] In some embodiments, when the descriptive text of the visual transformation corresponding to each plot node includes the target posture of each of the one or more characters in each screen of each plot node, each screen of each plot node is generated based on the descriptive text of the plot content of each plot node and the target posture of each of the one or more characters in each screen of each plot node, wherein each character in each screen is displayed using the corresponding target posture.

[0114] In some embodiments, when the descriptive text of the visual transformation corresponding to each plot node includes the target background of each screen in each plot node, each screen of each plot node is generated based on the descriptive text of the plot content of each plot node and the target background of each screen in each plot node, wherein the background of each screen is the corresponding target background.

[0115] The above embodiments can be combined arbitrarily. Based on the description text of the plot content and the description text of the visual transformation, each frame (image) can be generated to form a video, or a video containing multiple frames can be directly generated.

[0116] In some embodiments, multimedia content for each interactive operation node is generated based on the description text of each interactive operation node. An interactive operation node can serve as a special plot node and can also include corresponding visual transformations. The method for generating multimedia content for an interactive operation node can be referenced with the method for generating multimedia content for a plot node and will not be further described here.

[0117] When generating multimedia content for each plot node, character descriptions, scene descriptions, and visual style descriptions are also used. These descriptions can be generated based on user input. Character descriptions can be used to generate character portraits and audio, while scene descriptions and visual style descriptions can be used to generate backgrounds, on-screen objects, and other elements, ultimately determining the overall style of the interactive multimedia content.

[0118] In some embodiments, at least one of a character portrait and audio is generated for each character based on the character's description text. A character portrait can be an image that reflects the character's overall appearance. Character portraits include: static portraits of characters from the front, portraits of characters in various poses, animated images or videos with added actions (such as opening and closing the mouth, swaying of a skirt or hair, etc.). For interactive movies / interactive TV series, the background scenes and the characters therein do not have fixed poses or screen layouts.

[0119] Interactive multimedia content may also include background music, such as music to add to the atmosphere, music for character appearances, and special effects music.

[0120] In some embodiments, a text generation model is used to parse the description information to generate a description text of the background music corresponding to each plot node. For example, the description text of the background music may include the music style, duration, etc.

[0121] In some embodiments, the background music for each plot node pair is generated based on the description text of the background music corresponding to each plot node.

[0122] The background music can be directly generated by using a generative model, or background music that matches the description text of the background music can be selected from a database.

[0123] When generating interactive multimedia content, the content text may include a combination of various types of description texts in the aforementioned embodiments, and the interactive multimedia content may be directly generated based on the content text.

[0124] For example, FIG6 shows a scene in the generated interactive movie, which includes characters, background, and interactive operation options that can be interacted by the viewing user. The scene can be a close-up shot and can be accompanied by background music that creates a tense atmosphere.

[0125] In step S310 , interactive multimedia content and content text are displayed.

[0126] Multiple plot nodes can be distinguished, and the content text of each plot node can be associated with the corresponding interactive multimedia content for display.

[0127] In some embodiments, the plot description text of each plot node, the description text of the visual transition between different pictures in each plot node, and the multimedia content of each plot node are displayed in association.

[0128] In some embodiments, the description text of each character is displayed in association with at least one of the character's portrait and audio; the description text of each interactive operation node is displayed in association with the multimedia content of each interactive operation node.

[0129] The character description text and character portrait, sound and audio, etc. can be displayed separately. The description text and multimedia content of each plot node or each interactive operation node can be displayed separately or in conjunction with the flowchart of Figure 4.

[0130] The content text is displayed in association with the interactive multimedia content, making it easy for users to adjust it. If the user is not satisfied with the generated content text or the interactive multimedia content, the interactive multimedia content can be regenerated by adjusting the content text.

[0131] As shown in FIG7 , a preview interface is shown. For a plot node 701 of "Exploring the Mountains," the content interface 702 of the plot node may display the content text 703 of the plot node. The content text of the plot node includes description text of the plot content and description text of the visual effects, etc. Multimedia content may also be displayed, including scene images 704, character images 705, background music 706, and corresponding videos. As shown in FIG7 , the logical relationships between the various plot nodes may also be displayed.

[0132] In some embodiments, for each plot node, a user adjustment of the description text of the visual transformation corresponding to the plot node is received; multimedia content of the plot node is generated based on the adjusted description text of the visual transformation corresponding to the plot node; and the adjusted description text of the visual transformation corresponding to the plot node and the adjusted multimedia content of the plot node are displayed.

[0133] It can also receive user's adjustment of the plot content description text of the plot node, the description text of the interactive operation node, the description text of the character, the description text of the scene, the description text of the picture style, etc., and regenerate the multimedia content.

[0134] The disclosed method can generate interactive multimedia content of different genres based on the user's (author's) descriptive information. The user can input the story background, character settings, basic summary, etc. The generative model can expand the complete script based on the descriptive information and generate the pictures and characters under the corresponding content to generate the scenes and characters corresponding to each plot node in the script. Scene background images, character portraits, character postures, background music, etc. can all be generated based on the descriptive information entered by the user, and multiple options can be generated for the user to choose from during the generation.

[0135] Generative models can use user-entered descriptive information to determine the current background and characters in the scene, and can also determine the current camera movement based on the text content. Characters in interactive multimedia content are played by AI models, and players (either viewers or operators) can interact with them. Users (authors) can also set different plot branches corresponding to different player options. Players can also interact with NPCs (non-player characters) by inputting content rather than just relying on user-provided options.

[0136] The method disclosed herein can improve the generation efficiency and diversity of interactive multimedia content.

[0137] The present disclosure also provides a device for generating interactive multimedia content, which will be described below in conjunction with FIG. 8 .

[0138] FIG8 is a structural diagram of some embodiments of the interactive multimedia content generation device disclosed herein. As shown in FIG8 , the device of this embodiment includes: a receiving module 810 , a content generation module 820 , a multimedia generation module 830 , and a display module 840 .

[0139] The receiving module 810 is configured to receive description information input by a user.

[0140] The content generation module 820 is configured to generate content text according to the description information and the target genre, wherein the content text includes a description text of the plot content and a description text of the visual transformation corresponding to the target genre.

[0141] The multimedia generation module 830 is configured to generate interactive multimedia content according to the content text.

[0142] The display module 840 is configured to display the interactive multimedia content and the content text, wherein the interactive multimedia content presents a visual transformation effect corresponding to the description text of the visual transformation.

[0143] In some embodiments, the interactive multimedia content generation device 80 also includes a genre determination module 850, which is configured to parse the description information and determine the genre corresponding to the description information as the target genre; or the receiving module 810 is also configured to receive the user's selection operation of one or more genres among multiple genres, and the genre determination module 850 is configured to determine one or more genres as the target genre.

[0144] In some embodiments, the descriptive text of the plot content includes the descriptive text of the plot content of each plot node in one or more plot nodes, and the descriptive text of the visual transformation includes the descriptive text of the visual transformation corresponding to each plot node. The content generation module 820 is configured to use a text generation model to parse the description information and generate a descriptive text of the plot content of each plot node; use the text generation model to determine the information of the visual transformation between different pictures in each plot node based on the description information and the target genre, and generate a descriptive text of the visual transformation corresponding to each plot node.

[0145] In some embodiments, the visual transformation includes the operation of the lens, and the content generation module 820 is configured to determine whether to use different types of lens effects in different frames in each plot node based on the description information and the target genre; for the plot nodes that are determined to use different types of lens effects, the target lens effect used in each frame in the plot node and the target object corresponding to the target lens effect are determined based on the description information as information on the operation of the lens between different frames in the plot node.

[0146] In some embodiments, the visual transformation includes a change in the character's posture, and the content generation module 820 is configured to determine whether to change the character's posture in each plot node based on the description information and the target genre; for the plot node where the character's posture is determined to be changed, the target posture of each of one or more characters in each screen in the plot node is determined based on the description information, as the change information of the posture of each character between different screens in the plot node.

[0147] In some embodiments, the visual transformation includes switching of the background. The content generation module 820 is configured to determine whether to switch the background in each plot node based on the description information and the target genre; for the plot node where the background switching is determined, the target background of each screen in the plot node is determined based on the description information, and the target background of each screen in the plot node is used as the switching information between the backgrounds of different screens in each plot node.

[0148] In some embodiments, the multimedia generation module 830 is configured to generate multimedia content for each plot node based on the descriptive text of the plot content of each plot node and the descriptive text of the visual transformation corresponding to each plot node, wherein the multimedia content of each plot node presents a visual effect corresponding to the descriptive text of the visual transformation corresponding to the plot node.

[0149] In some embodiments, the descriptive text of the visual transformation corresponding to each plot node includes the target lens effect used in each picture in each plot node and the target object corresponding to the target lens effect. The multimedia generation module 830 is configured to generate each picture of each plot node based on the descriptive text of the plot content of each plot node and the target lens effect used in each picture in each plot node and the target object corresponding to the target lens effect, wherein the target object in each picture is displayed using the corresponding target lens effect.

[0150] In some embodiments, the descriptive text of the visual transformation corresponding to each plot node includes the target posture of each of the one or more characters in each screen of each plot node. The multimedia generation module 830 is configured to generate each screen of each plot node based on the descriptive text of the plot content of each plot node and the target posture of each of the one or more characters in each screen of each plot node, wherein each character in each screen is displayed using the corresponding target posture.

[0151] In some embodiments, the descriptive text of the visual transformation corresponding to each plot node includes the target background of each screen in each plot node. The multimedia generation module 830 is configured to generate each screen of each plot node based on the descriptive text of the plot content of each plot node and the target background of each screen in each plot node, wherein the background of each screen is the corresponding target background.

[0152] In some embodiments, the display module 840 is configured to display the plot description text of each plot node, the description text of the visual transition between different pictures in each plot node, and the multimedia content of each plot node in an associated manner.

[0153] In some embodiments, for each plot node, the receiving module 810 is further configured to receive user adjustments to the descriptive text of the visual transformation corresponding to the plot node; the multimedia generation module 830 is further configured to generate multimedia content of the plot node based on the adjusted descriptive text of the visual transformation corresponding to the plot node; and the display module 840 is further configured to display the adjusted descriptive text of the visual transformation corresponding to the plot node and the adjusted multimedia content of the plot node.

[0154] In some embodiments, the content text also includes description text for each of one or more characters and description text for each interactive operation node in one or more interactive operation nodes. The content generation module 820 is configured to use a text generation model to parse the description information to generate description text for each character; use a text generation model to parse the description information to generate description text for each interactive operation node, wherein the description text of each interactive operation node includes description text of the interactive operation options displayed to the operating user.

[0155] In some embodiments, the description text of each character includes descriptive information of at least one of the image and sound characteristics of each character. The content generation module 820 is configured to generate at least one of the portrait and sound audio of each character based on the description text of each character; and generate multimedia content for each interactive operation node based on the description text of each interactive operation node.

[0156] In some embodiments, the display module 840 is configured to display the description text of each character in association with at least one of the portrait and audio of each character; and to display the description text of each interactive operation node in association with the multimedia content of each interactive operation node.

[0157] In some embodiments, the content text also includes descriptive text of the background music corresponding to each of one or more plot nodes. The content generation module 820 is configured to use a text generation model to parse the description information and generate descriptive text of the background music corresponding to each plot node.

[0158] In some embodiments, the multimedia generation module 830 is configured to generate background music for each plot node pair based on the description text of the background music corresponding to each plot node.

[0159] In some embodiments, the description information includes at least one of plot summary information and plot information of each of the one or more plot nodes.

[0160] In some embodiments, the description information further includes at least one of: a logical relationship between plot nodes, character setting information, scene setting information, and picture style setting information.

[0161] In some embodiments, the target genre includes at least one of: a game, an interactive TV series, and an interactive movie.

[0162] It should be noted that the above-mentioned units are merely logical modules divided according to the specific functions they implement, and are not intended to limit specific implementation methods. For example, they can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, the above-mentioned units can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, the above-mentioned units are shown with dotted lines in the accompanying drawings to indicate that these units may not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.

[0163] In addition, although not shown, the device may also include a memory that can store various information generated by the device and the various units contained in the device during operation, programs and data used for operation, data to be sent by the communication unit, etc. The memory can be volatile memory and / or non-volatile memory. For example, the memory can include but is not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Of course, the memory can also be located outside the device. Optionally, although not shown, the device may also include a communication unit that can be used to communicate with other devices. In one example, the communication unit can be implemented in an appropriate manner known in the art, for example, including communication components such as an antenna array and / or a radio frequency link, various types of interfaces, communication units, etc. This will not be described in detail here. In addition, the device may also include other components not shown, such as a radio frequency link, a baseband processing unit, a network interface, a processor, a controller, etc. This will not be described in detail here.

[0164] Some embodiments of the present disclosure also provide an electronic device. Figure 9 shows a block diagram of some embodiments of the electronic device of the present disclosure. For example, in some embodiments, the electronic device 9 can be various types of devices, for example, including but not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. For example, the electronic device 9 may include a display panel for displaying data and / or execution results utilized in the scheme of the present disclosure. For example, the display panel can be of various shapes, such as a rectangular panel, an elliptical panel, or a polygonal panel. In addition, the display panel can be not only a flat panel, but also a curved panel or even a spherical panel.

[0165] As shown in FIG9 , the electronic device 9 of this embodiment includes a memory 91 and a processor 92 coupled to the memory 91. It should be noted that the components of the electronic device 9 shown in FIG9 are merely exemplary and non-limiting. The electronic device 9 may also include other components as required by actual applications. The processor 92 may control the other components in the electronic device 9 to perform desired functions.

[0166] In some embodiments, the memory 91 is configured to store one or more computer-readable instructions. When the processor 92 is configured to execute the computer-readable instructions, the computer-readable instructions, when executed by the processor 92, implement a method according to any of the above-described embodiments. The specific implementation and related explanations of each step of the method can be found in the above-described embodiments, and any repetitive details are omitted here.

[0167] For example, the processor 92 and the memory 91 may communicate with each other directly or indirectly. For example, the processor 92 and the memory 91 may communicate with each other via a network. The network may include a wireless network, a wired network, and / or any combination of wireless networks and wired networks. The processor 92 and the memory 91 may also communicate with each other via a system bus, which is not limited in this disclosure.

[0168] For example, the processor 92 can be embodied as various appropriate processors, processing devices, etc., such as a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) can be an X86 or ARM architecture, etc. For example, the memory 91 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The memory 91 can include, for example, a system memory, which stores, for example, an operating system, an application, a boot loader (Boot Loader), a database, and other programs. Various applications and various data can also be stored in the storage medium.

[0169] In addition, according to some embodiments of the present disclosure, when various operations / processes according to the present disclosure are implemented through software and / or firmware, the programs constituting the software can be installed from a storage medium or a network to a computer system having a dedicated hardware structure, such as the computer system (or electronic device) 1000 shown in Figure 10. When the various programs are installed, the computer system can perform various functions, including functions such as those described above. Figure 10 is a block diagram showing an example structure of a computer system that can be used in embodiments of the present disclosure.

[0170] In Figure 10, a central processing unit (CPU) 1001 performs various processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage part 1008 to a random access memory (RAM) 1003. In the RAM 1003, data required when the CPU 1001 performs various processes, etc., is also stored as needed. The central processing unit is merely exemplary and may also be other types of processors, such as the various processors described above. The ROM 1002, RAM 1003, and storage part 1008 may be various forms of computer-readable storage media, as described below. It should be noted that although ROM 1002, RAM 1003, and storage device 1008 are shown separately in Figure 10, one or more of them may be combined or located in the same or different memory or storage modules.

[0171] The CPU 1001, the ROM 1002, and the RAM 1003 are connected to one another via a bus 1004. An input / output interface 1005 is also connected to the bus 1004.

[0172] The following components are connected to the input / output interface 1005: an input portion 1006, such as a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output portion 1007, including a display, such as a cathode ray tube (CRT), liquid crystal display (LCD), speaker, vibrator, etc.; a storage portion 1008, including a hard disk, magnetic tape, etc.; and a communication portion 1009, including a network interface card, such as a LAN card, modem, etc. The communication portion 1009 allows communication processing to be performed via a network, such as the Internet. It will be readily understood that although FIG10 shows that the various devices or modules in the computer system 1000 communicate via the bus 1004, they may also communicate via a network or other means, where the network may include a wireless network, a wired network, and / or any combination of wireless and wired networks.

[0173] A drive 1010 is also connected to the input / output interface 1005 as needed. A removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1010 as needed so that a computer program read therefrom is installed in the storage section 1008 as needed.

[0174] In the case of realizing the above-described series of processing by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 1011 .

[0175] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 1009, or installed from the storage part 1008, or installed from the ROM 1002. When the computer program is executed by the CPU 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0176] It should be noted that in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0177] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0178] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform any of the methods of the above embodiments. For example, the instructions may be embodied as computer program codes.

[0179] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure can be written in one or more programming languages ​​or combinations thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In situations involving a remote computer, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).

[0180] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0181] The modules, components, or units described in the embodiments of the present disclosure may be implemented in software or hardware. The names of the modules, components, or units do not necessarily limit the modules, components, or units themselves.

[0182] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0183] According to some embodiments of the present disclosure, a method for generating interactive multimedia content is provided, comprising: receiving descriptive information input by a user; generating content text based on the descriptive information and a target genre, wherein the content text includes descriptive text of the plot content and descriptive text of a visual transformation corresponding to the target genre; generating interactive multimedia content based on the content text; and displaying the interactive multimedia content and the content text, wherein the interactive multimedia content presents a visual transformation effect corresponding to the descriptive text of the visual transformation.

[0184] In some embodiments, the generation method further includes: parsing the description information to determine the genre corresponding to the description information as the target genre; or receiving a user selection operation for one or more genres among multiple genres, and determining the one or more genres as the target genre.

[0185] In some embodiments, the descriptive text of the plot content includes the descriptive text of the plot content of each plot node in one or more plot nodes, and the descriptive text of the visual transformation includes the descriptive text of the visual transformation corresponding to each plot node. Generating the content text according to the description information and the target genre includes: using a text generation model to parse the description information to generate the descriptive text of the plot content of each plot node; using the text generation model to determine the information of the visual transformation between different pictures in each plot node according to the description information and the target genre, and generating the descriptive text of the visual transformation corresponding to each plot node.

[0186] In some embodiments, the visual transformation includes the movement of the lens. The information for determining the visual transformation between different frames in each plot node based on the descriptive information and the target genre includes: determining whether to use different types of lens effects in different frames in each plot node based on the descriptive information and the target genre; for plot nodes that are determined to use different types of lens effects, determining the target lens effect used in each frame in the plot node and the target object corresponding to the target lens effect based on the descriptive information, as the information for the movement of the lens between different frames in the plot node.

[0187] In some embodiments, the visual transformation includes a change in a character's posture. Determining the information of the visual transformation between different frames in each plot node based on the description information and the target genre includes: determining whether to change the character's posture in each plot node based on the description information and the target genre; and for the plot node in which the character's posture is determined to be changed, determining the target posture of each of one or more characters in each frame in the plot node based on the description information, as the change information of the posture of each character between different frames in the plot node.

[0188] In some embodiments, the visual transformation includes switching of the background. The information for determining the visual transformation between different pictures in each plot node based on the description information and the target genre includes: determining whether to switch the background in each plot node based on the description information and the target genre; and for the plot node where the background switching is determined, determining the target background for each picture in the plot node based on the description information as the background switching information between different pictures in each plot node.

[0189] In some embodiments, generating interactive multimedia content based on content text includes: generating multimedia content for each plot node based on the description text of the plot content of each plot node and the description text of the visual transformation corresponding to each plot node, wherein the multimedia content of each plot node presents a visual effect corresponding to the description text of the visual transformation corresponding to the plot node.

[0190] In some embodiments, the descriptive text of the visual transformation corresponding to each plot node includes the target lens effect used in each picture in each plot node and the target object corresponding to the target lens effect. Generating the multimedia content of each plot node based on the descriptive text of the plot content of each plot node and the descriptive text of the visual transformation corresponding to each plot node includes: generating each picture of each plot node based on the descriptive text of the plot content of each plot node and the target lens effect used in each picture in each plot node and the target object corresponding to the target lens effect, wherein the target object in each picture is displayed using the corresponding target lens effect.

[0191] In some embodiments, the descriptive text of the visual transformation corresponding to each plot node includes the target posture of each of the one or more characters in each screen of each plot node. Generating the multimedia content of each plot node based on the descriptive text of the plot content of each plot node and the descriptive text of the visual transformation corresponding to each plot node includes: generating each screen of each plot node based on the descriptive text of the plot content of each plot node and the target posture of each of the one or more characters in each screen of each plot node, wherein each character in each screen is displayed using the corresponding target posture.

[0192] In some embodiments, the descriptive text of the visual transformation corresponding to each plot node includes the target background of each screen in each plot node. Generating the multimedia content of each plot node based on the descriptive text of the plot content of each plot node and the descriptive text of the visual transformation corresponding to each plot node includes: generating each screen of each plot node based on the descriptive text of the plot content of each plot node and the target background of each screen in each plot node, wherein the background of each screen is the corresponding target background.

[0193] In some embodiments, displaying the interactive multimedia content and content text includes: displaying the plot description text of each plot node, the description text of the visual transitions between different screens in each plot node, and the multimedia content of each plot node in an associated manner.

[0194] In some embodiments, the generation method further includes, for each plot node: receiving a user's adjustment of the description text of the visual transformation corresponding to the plot node; generating multimedia content of the plot node based on the adjusted description text of the visual transformation corresponding to the plot node; and displaying the adjusted description text of the visual transformation corresponding to the plot node and the adjusted multimedia content of the plot node.

[0195] In some embodiments, the content text also includes description text for each of one or more characters and description text for each interactive operation node in one or more interactive operation nodes. Generating content text based on the description information and the target genre includes: parsing the description information using a text generation model to generate description text for each character; parsing the description information using a text generation model to generate description text for each interactive operation node, wherein the description text of each interactive operation node includes description text for the interactive operation options displayed to the operating user.

[0196] In some embodiments, the description text of each character includes descriptive information of at least one of the image and sound characteristics of each character. Generating interactive multimedia content based on the content text includes: generating at least one of the character's portrait and sound audio based on the description text of each character; generating multimedia content for each interactive operation node based on the description text of each interactive operation node.

[0197] In some embodiments, displaying interactive multimedia content and content text includes: displaying the description text of each character and at least one of the character's portrait and sound audio in association with each character; and displaying the description text of each interactive operation node and the multimedia content of each interactive operation node in association with each character.

[0198] In some embodiments, the content text also includes descriptive text of the background music corresponding to each of one or more plot nodes. Based on the descriptive information and the target genre, generating the content text includes: using a text generation model to parse the descriptive information and generate descriptive text of the background music corresponding to each plot node.

[0199] In some embodiments, generating interactive multimedia content based on the content text includes generating background music for each plot node pair based on a description text of the background music corresponding to each plot node.

[0200] In some embodiments, the description information includes at least one of plot summary information and plot information of each of the one or more plot nodes.

[0201] In some embodiments, the description information further includes at least one of: a logical relationship between plot nodes, character setting information, scene setting information, and picture style setting information.

[0202] In some embodiments, the target genre includes at least one of: a game, an interactive TV series, and an interactive movie.

[0203] According to other embodiments of the present disclosure, a device for generating interactive multimedia content is provided, comprising: a receiving module configured to receive descriptive information input by a user; a content generating module configured to generate content text based on the descriptive information and a target genre, wherein the content text includes descriptive text of the plot content and descriptive text of the visual transformation corresponding to the target genre; a multimedia generating module configured to generate interactive multimedia content based on the content text; and a display module configured to display the interactive multimedia content and the content text, wherein the interactive multimedia content presents a visual transformation effect corresponding to the descriptive text of the visual transformation.

[0204] According to some further embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the method for generating interactive multimedia content of any of the aforementioned embodiments based on instructions stored in the memory.

[0205] According to some further embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for generating interactive multimedia content according to any of the aforementioned embodiments is implemented.

[0206] According to still other embodiments of the present disclosure, a computer program is provided, comprising: instructions, which, when executed by a processor, cause the processor to execute the method for generating interactive multimedia content according to any embodiment of the present disclosure.

[0207] According to some further embodiments of the present disclosure, a computer program product is provided, comprising instructions, which, when executed by a processor, implement the method for generating interactive multimedia content according to any embodiment of the present disclosure.

[0208] The above descriptions are merely some embodiments of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present disclosure.

[0209] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the present invention may be practiced without these specific details. In other cases, well-known methods, structures, and techniques are not presented in detail in order not to obscure the understanding of the description.

[0210] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0211] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A method for generating interactive multimedia content, comprising: Receive description information input by the user; Generate a content text according to the description information and the target genre, wherein the content text includes a description text of the plot content and a description text of the visual transformation corresponding to the target genre; generating interactive multimedia content according to the content text; The interactive multimedia content and the content text are displayed, wherein the interactive multimedia content presents a visual transformation effect corresponding to the description text of the visual transformation.

2. The generation method according to claim 1, further comprising: Parsing the description information to determine a genre corresponding to the description information as the target genre; or A selection operation of one or more genres among a plurality of genres is received from a user, and the one or more genres are determined as the target genre.

3. The generation method according to claim 1 or 2, wherein: The description text of the plot content includes the description text of the plot content of each plot node in one or more plot nodes, the description text of the visual transformation includes the description text of the visual transformation corresponding to each plot node, and the generating of the content text according to the description information and the target genre includes: The description information is parsed using a text generation model to generate a description text of the plot content of each plot node; The text generation model is used to determine the information of the visual transformation between different pictures in each plot node according to the description information and the target genre, and a description text of the visual transformation corresponding to each plot node is generated.

4. The generation method according to claim 3, wherein: The visual transformation includes the movement of the lens, and the information for determining the visual transformation between different pictures in each plot node according to the description information and the target genre includes: Determining whether to use different types of lens effects in different pictures in each plot node according to the description information and the target genre; For a plot node that determines to use different types of lens effects, a target lens effect used in each picture in the plot node and a target object corresponding to the target lens effect are determined according to the description information. As information about the operation of shots between different pictures in the plot node.

5. The generation method according to claim 3 or 4, wherein: The visual transformation includes a change in a character's posture, and the information for determining the visual transformation between different pictures in each plot node according to the description information and the target genre includes: Determining whether to change the posture of the character in each plot node according to the description information and the target genre; For a plot node for determining a changed character's posture, a target posture of each of one or more characters in each frame of the plot node is determined based on the description information as posture change information of each character between different frames in the plot node.

6. The generation method according to any one of claims 3 to 5, wherein: The visual transformation includes switching of backgrounds, and the information for determining the visual transformation between different pictures in each plot node according to the description information and the target genre includes: Determining whether to switch the background in each plot node according to the description information and the target genre; For the plot node for determining the background switching, the target background of each screen in the plot node is determined according to the description information, as the background switching information between different screens in each plot node.

7. The generation method according to any one of claims 3 to 6, wherein: Generating interactive multimedia content according to the content text includes: The multimedia content of each plot node is generated based on the description text of the plot content of each plot node and the description text of the visual transformation corresponding to each plot node, wherein the multimedia content of each plot node presents a visual effect corresponding to the description text of the visual transformation corresponding to the plot node.

8. The generation method according to claim 7, wherein: The description text of the visual transformation corresponding to each plot node includes the target lens effect used in each picture in each plot node and the target object corresponding to the target lens effect. The generating of the multimedia content of each plot node according to the description text of the plot content of each plot node and the description text of the visual transformation corresponding to each plot node includes: Each picture of each plot node is generated based on the description text of the plot content of each plot node, the target lens effect used in each picture of each plot node, and the target object corresponding to the target lens effect, wherein the target object in each picture is displayed using the corresponding target lens effect.

9. The generation method according to claim 7 or 8, wherein: The description text of the visual transformation corresponding to each plot node includes the target posture of each character in one or more characters in each screen of each plot node, and the generating of the multimedia content of each plot node according to the description text of the plot content of each plot node and the description text of the visual transformation corresponding to each plot node includes: Each picture of each plot node is generated according to the description text of the plot content of each plot node and the target posture of each of one or more characters in each picture of each plot node, wherein each character in each picture is displayed with a corresponding target posture.

10. The generation method according to any one of claims 7 to 9, wherein: The description text of the visual transformation corresponding to each plot node includes the target background of each screen in each plot node, and the generating of the multimedia content of each plot node according to the description text of the plot content of each plot node and the description text of the visual transformation corresponding to each plot node includes: Each picture of each plot node is generated according to the description text of the plot content of each plot node and the target background of each picture in each plot node, wherein the background of each picture is the corresponding target background.

11. The generation method according to any one of claims 7 to 10, wherein: Displaying the interactive multimedia content and the content text includes: The plot description text of each plot node, the description text of the visual transformation between different pictures in each plot node and the multimedia content of each plot node are displayed in association.

12. The generation method according to any one of claims 3 to 11, further comprising, for each plot node: receiving, from the user, an adjustment of a description text of a visual transformation corresponding to the plot node; Generating multimedia content of the plot node according to the adjusted description text of the visual transformation corresponding to the plot node; The adjusted visually transformed description text corresponding to the plot node and the adjusted multimedia content of the plot node are displayed.

13. The generation method according to any one of claims 1 to 12, wherein: The content text also includes a description text of each of the one or more characters and a description text of each of the one or more interactive operation nodes. Generating the content text according to the description information and the target genre includes: Using a text generation model to parse the description information to generate a description text for each character; The description information is parsed using a text generation model to generate a description text for each interactive operation node, wherein the description text for each interactive operation node includes a description text for interactive operation options displayed to an operating user.

14. The generation method according to claim 13, wherein: The description text of each character includes description information of at least one of the image and voice characteristics of each character, and generating interactive multimedia content according to the content text includes: Generate at least one of a portrait and a sound audio of each character according to the description text of each character; The multimedia content of each interactive operation node is generated according to the description text of each interactive operation node.

15. The generation method according to claim 14, wherein: The displaying of the interactive multimedia content and the content text comprises: Displaying the description text of each character in association with at least one of the portrait and the audio of each character; The description text of each interactive operation node and the multimedia content of each interactive operation node are displayed in association.

16. The generation method according to any one of claims 1 to 15, wherein: The content text also includes a description text of the background music corresponding to each of the one or more plot nodes, and generating the content text according to the description information and the target genre includes: The description information is parsed using a text generation model to generate a description text of the background music corresponding to each plot node.

17. The generation method according to claim 16, wherein: Generating interactive multimedia content according to the content text includes: The background music for each plot node pair is generated according to the description text of the background music corresponding to each plot node.

18. The generation method according to any one of claims 1 to 17, wherein: The description information includes plot summary information and at least one item of plot information of each of the one or more plot nodes.

19. The generation method according to claim 18, wherein: The description information also includes: at least one of the logical relationship between various plot nodes, character setting information, scene setting information, and picture style setting information.

20. The generation method according to any one of claims 1 to 19, wherein: The target genre includes at least one of: games, interactive TV series, and interactive movies.

21. A device for generating interactive multimedia content, comprising: A receiving module, configured to receive description information input by a user; The content generation module is configured to generate content text according to the description information and the target genre, wherein the content text includes a description text of the plot content and a description text of the visual transformation corresponding to the target genre. Book; A multimedia generation module, configured to generate interactive multimedia content according to the content text; The display module is configured to display the interactive multimedia content and the content text, wherein the interactive multimedia content presents a visual transformation effect corresponding to the description text of the visual transformation.

22. An electronic device, comprising: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the method for generating interactive multimedia content according to any one of claims 1 to 20 based on instructions stored in the memory.

23. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method for generating interactive multimedia content according to any one of claims 1 to 20 is implemented.

24. A computer program, comprising instructions, which, when executed by a processor, enable the processor to implement the method for generating interactive multimedia content according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • Plot animation production method and device

    CN109493402A

  • Multimedia data generation method and device, readable medium and electronic equipment

    CN113778419A

  • Generation method and device of interactive multimedia content, electronic equipment and storage medium

    CN117633258A

  • Presenting interactive content

    WO2019169068A1

Cited By

  • Video generation method and system based on environmental perception, server and medium

    CN120711257A