Method and device for providing media data, equipment and medium

CN121264053APending Publication Date: 2026-01-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480003606.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-29
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

In existing interactive multimedia content, users can only interact in a predetermined way and cannot obtain personalized media data, which affects interactivity.

Method used

By receiving user input, the system calls machine learning models to generate dynamic media data, including personalized dynamic media data based on interactive multimedia content and user input selection scenarios and interactive entities.

Benefits of technology

It enhances the interactivity of interactive multimedia content, provides richer media data, and enables personalized customization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121264053A_ABST
    Figure CN121264053A_ABST
Patent Text Reader

Abstract

According to the embodiment of the invention, a method, a device, equipment and a medium for providing media data in interactive multimedia content are provided. In the method, the interactive multimedia content is presented in response to a request of a user receiving the interactive multimedia content to enter the interactive multimedia content. In response to receiving a first user input of a user in the interactive multimedia content, a scene in the interactive multimedia content is selected, the scene including an interactive entity. In response to receiving a second user input by the user for the interactive entity, dynamic media data is generated based on the interactive multimedia content and the second user input in response to the second user input. In this way, the powerful processing capacity of the machine learning model can be called, the response data of the interactive entity can be generated based on the user input, personalized customization can be achieved, and the interactivity of the user in the interactive multimedia content can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and medium for providing media data TECHNICAL FIELD

[0001] Exemplary implementations of the present disclosure generally relate to the field of computers, and particularly relate to a method, apparatus, device and computer readable storage medium for providing media data in interactive multimedia content. BACKGROUND

[0002] With the development of computers, various forms of electronic devices can greatly enrich people's daily life. For example, people can use electronic devices to interact in various ways. Interactive multimedia content can be run at an electronic device, and the electronic device can present an interactive interface supporting user interaction through a display screen. Various interactive elements supporting user interaction can be included in the interactive interface. However, the user can only perform the interaction in a predetermined manner in the interactive multimedia content, and the media data provided by the interactive multimedia content to different users is similar, and the user cannot obtain personalized media data. It is desirable that the interactive multimedia content can provide more rich media data, for example, generate and present corresponding personalized media data based on the user input received in the interactive multimedia content.

[0003] SUMMARY

[0004] In a first aspect of the present disclosure, a method for providing media data is provided. The method comprises: in response to receiving a request of a user of interactive multimedia content to enter the interactive multimedia content, presenting the interactive multimedia content. In response to receiving a first user input of the user in the interactive multimedia content, selecting a scene in the interactive multimedia content, the scene comprising an interactive entity. In response to receiving a second user input of the user for the interactive entity, generating dynamic media data based on the interactive multimedia content and the second user input to respond to the second user input.

[0005] In a second aspect of the present disclosure, an apparatus for providing media data is provided. The apparatus comprises: a presentation module configured to present the interactive multimedia content in response to receiving a request of a user of interactive multimedia content to enter the interactive multimedia content; a selection module configured to select a scene in the interactive multimedia content in response to receiving a first user input of the user in the interactive multimedia content, the scene comprising an interactive entity; and a generation module configured to generate dynamic media data based on the interactive multimedia content and a second user input of the user for the interactive entity to respond to the second user input in response to receiving the second user input.

[0006] In a third aspect of the disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the disclosure.

[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to the first aspect of the disclosure.

[0008] According to a fifth aspect of the disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the disclosure.

[0009] It is to be understood that the particulars shown herein are by way of example and for purposes of illustrative discussion of the present implementations of the present disclosure only and are presented in the cause of providing what is believed to be the most useful and readily understood description of the principles and conceptual aspects of various implementations of the present disclosure. In this regard, no requirement exists for the details of construction or design herein shown and described to detract from the conceptual understanding of the present disclosure. Further, various implementations of the present disclosure can have additional features or advantages which may BRIEF DESCRIPTION OF DRAWINGS

[0010] In the following detailed description, reference will be made to the accompanying drawings, of which:

[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIGS. 2A and 2B show schematic diagrams of example pages, according to some embodiments of the present disclosure;

[0013] FIGS. 3A to 3C show schematic diagrams of example pages, according to some other embodiments of the present disclosure;

[0014] FIG. 3D shows a schematic diagram of media data provided at different stages of performing a task in an interactive multimedia content, according to some other embodiments of the present disclosure;

[0015] FIG. 4 shows a flowchart of a process of providing media data in an interactive multimedia content, according to some embodiments of the present disclosure;

[0016] FIG. 5 shows a schematic structural block diagram of an apparatus for providing media data in an interactive multimedia content, according to some embodiments of the present disclosure; and

[0017] FIG. 6 shows a block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be understood more thoroughly and completely. It is understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0019] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meaning of "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.

[0020] In this document, unless explicitly stated, performing a step "in response to A" does not mean performing the step immediately after A, but can include one or more intermediate steps.

[0021] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0023] For example, in response to receiving the active request of the user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user, so that the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.

[0024] As an optional but not limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be the manner of pop-up window, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0026] As used herein, the term “model” can learn the relationship between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. The neural network model is an example of a model based on deep learning. In this article, “model” can also be referred to as “machine learning model”, “learning model”, “machine learning network” or “learning network”, which are used interchangeably in this article.

[0027] “Neural network” is a machine learning network based on deep learning. Neural networks can process inputs and provide corresponding outputs, which usually include input layers and output layers and one or more hidden layers between the input layers and the output layers. Neural networks used in deep learning applications usually include many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also known as processing nodes or neurons), each of which processes input from the previous layer.

[0028] Generally, machine learning can include three stages, namely training stage, testing stage and application stage (also known as inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter value is updated iteratively until the model can obtain consistent inference from the training data that meets the expected target. Through training, the model can be considered to learn the relationship between input and output (also known as input to output mapping) from the training data. The parameter value of the trained model is determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, so as to determine the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter value obtained by training to determine the corresponding model output.

[0029] A variety of interactive multimedia content has been developed. For example, a user can perform various tasks in a virtual reality environment. The interactive multimedia content can include, for example, a game environment and / or other simulation environment. The user can perform various tasks in the game environment. For example, the user can have conversational interactions with a corresponding virtual character and other virtual characters (e.g., non-player characters (NPCs)) in the game environment. Conventionally, in the interactive multimedia content, when the user’s corresponding virtual character has conversational interactions with the other virtual characters, the other virtual characters can only answer preset questions and reply with fixed responses to the preset questions. This can affect the interactivity of the user in the interactive multimedia content.

[0030] In view of this, to at least partially address the deficiencies in the prior art, embodiments of the present disclosure propose an improved scheme for providing media data in interactive multimedia content. According to the scheme, in response to receiving a request of a user of the interactive multimedia content to enter the interactive multimedia content, the interactive multimedia content is presented. In response to receiving a first user input of the user in the interactive multimedia content, a scene in the interactive multimedia content is selected, the scene including an interactable entity. In response to receiving a second user input of the user for the interactable entity, dynamic media data is generated based on the interactive multimedia content and the second user input to respond to the second user input.

[0031] In this way, the powerful processing capability of the machine learning model can be invoked to generate the response data of the interactable entity in each scene based on the user input. Thus, more rich media data can be provided in the interactive multimedia content and individualized customization can be achieved, thereby the interactivity of the user in the interactive multimedia content can be improved.

[0032] Example environment

[0033] FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, the example environment 100 can include a terminal device 110.

[0034] In this example environment 100, the terminal device 110 can run an application 120 that supports interactive multimedia content. The application 120 can be any suitable type of application, and the interactive multimedia content examples can include, but are not limited to, a game environment, a simulation environment, a simulation environment, a virtual reality environment, an augmented reality environment, and the like, and embodiments of the present disclosure are not limited in this respect. A user 140 can interact with the application 120 via the terminal device 110 and / or its attached devices.

[0035] In the environment 100 of FIG. 1, the terminal device 110 can render a page 150 of interactive multimedia content through the application 120 if the application 120 is active. The page 150 can be any suitable type of interactive interface that can support any suitable type of data for user input.

[0036] In some embodiments, the terminal device 110 communicates with the server 130 to implement provisioning of services for the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a palmtop computer, a portable gaming terminal, a VR / AR device, a Personal Communication System (PCS) device, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including an accessory or peripheral device for the foregoing, or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface to the user (such as “wearable” circuitry, etc.).

[0037] The server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content distribution network, and big data and artificial intelligence platform. The server 130 may, for example, include a computing system / server such as a mainframe, an edge computing node, a computing device in a cloud environment, etc. The server 130 may, for example, provide background services for the application 120 in the terminal device 110.

[0038] A communication connection can be established between the server 130 and the terminal device 110. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and embodiments of the present disclosure are not limited in this regard. In embodiments of the present disclosure, the server 130 and the terminal device 110 can implement signaling interaction through the communication connection therebetween.

[0039] It should be appreciated that the structure and functionality of the various elements in the environment 100 are described for illustrative purposes only and without implying any limitation on the scope of the present disclosure. Some example embodiments of the present disclosure will be described hereinafter with continuous reference to the drawings.

[0040] Example page providing media data

[0041] FIGS. 2A-3C show schematic diagrams of example pages 200A-300C (may also be referred to simply as examples 200A-300C) in accordance with some embodiments of the present disclosure. It should be appreciated that the pages shown in the drawings are merely examples and various page designs can actually exist. Various graphical elements in the pages can have different arrangements and different visual representations, one or more elements among them can be omitted or replaced, and one or more other elements can also exist. Embodiments of the present disclosure are not limited in this respect.

[0042] The pages shown in the examples 200A-300C can be presented at the terminal device 110. For ease of discussion, the examples 200A-300C will be described with reference to the environment 100 of FIG. 1. It is noted that the operations performed by the aforementioned terminal device 110 and the operations performed by the terminal device 110 described hereinafter can be specifically performed by a relevant application (e.g., the application 120) installed on the terminal device 110. In some embodiments, the operations performed on the terminal device 110 can be completed with the assistance of the server 130.

[0043] Generally speaking, the terminal device 110 can receive user input of a user in an interactive multimedia content, utilize the capability of a machine learning model (e.g., a machine learning model deployed at the server 130) to generate dynamic media data, and present the dynamic media data in the interactive multimedia content. For ease of description, the specific creation process will be described hereinafter only with the game environment as an example of the interactive multimedia content. Alternatively and / or additionally, the interactive multimedia content can also include a simulation environment, a simulation environment, a virtual reality environment, an augmented reality environment, etc.

[0044] The method for providing media data in an interactive multimedia content shown in the present disclosure can be performed, for example, in a sharing platform (e.g., the application 120) for sharing interactive multimedia content. A user of the sharing platform can create and publish an interactive multimedia content in the sharing platform so as to enable users of the sharing platform to access the interactive multimedia content. The terminal device may, for example, publish an interactive multimedia content in the sharing platform in response to receiving a publishing request for the interactive multimedia content. A user (e.g., the user 140) can browse and view the interactive multimedia content published by himself / herself and the interactive multimedia content published by other users in the sharing platform for sharing interactive multimedia content.

[0045] Further, a user of the sharing platform can access the published (published by himself / herself and / or published by other users) interactive multimedia content in the sharing platform. In embodiments of the present disclosure, the terminal device 110 presents the interactive multimedia content in response to receiving a request of a user of the interactive multimedia content to enter the interactive multimedia content. Illustratively, the terminal device 110 can present at least one promotional information associated with at least one published interactive multimedia content in the sharing platform, and determine that a request of a user of the sharing platform to enter a target interactive multimedia content corresponding to target promotional information is received in response to receiving a user operation of the user of the sharing platform to the target promotional information.

[0046] For example, the terminal device 110 can present a card associated with the interactive multimedia content in the sharing platform, the card being presented with a cover image of the interactive multimedia content. The terminal device 110 can determine that a request of a user of the sharing platform to interact with the promotional information is received in response to receiving a selection operation of the user (e.g. the user 140 or other user) to the card. The terminal device 110 further presents the interactive multimedia content in response to receiving the request of the user of the sharing platform to interact with the promotional information.

[0047] In some embodiments, the terminal device 110 can provide an interaction page for presenting the promotional information of the interactive multimedia content, and present the promotional information of the interactive multimedia content via the interaction page. The request of the user to enter the interactive multimedia content can be received, and the data in the interactive multimedia content is further presented. Illustratively, referring to FIG. 2A and FIG. 2B, example 200A shows one example of a page for presenting the promotional information of a scenario type of interactive multimedia content, and example 200B shows one example of a page for presenting the promotional information of a reasoning type of interactive multimedia content. Both example 200A and example 200B can include an operation control (e.g. control 210 in example 200A and control 220 in example 200B) indicating entering the interactive multimedia content.

[0048] The terminal device 110 can determine that the request of the user to enter the interactive multimedia content is received in response to receiving a user operation to the operation control indicating entering the interactive multimedia content (e.g. a trigger operation of the user to the operation control), and further present the corresponding interactive multimedia content. For example, in example 200B, the terminal device 110 can determine that the request of the user to enter the interactive multimedia content shown in example 200B is received in response to receiving a user operation to control 220, and further present the corresponding interactive multimedia content (e.g. the game environment of the game “Mountain Rescue”).

[0049] In some embodiments, the terminal device 110 can also provide an interaction page for presenting the interactive multimedia content, and present the interactive multimedia content, receive user input of the user in the interactive multimedia content, and present media data (including dynamic media data and static media data) via the interaction page. In some embodiments, the interactive multimedia content can include a main character (e.g., a virtual character corresponding to the user, i.e., a player character) and a task performed by the main character in the interactive multimedia content. The terminal device 110 can provide description information of the main character and the task in the interactive multimedia content, so as to guide the user to enter the interactive multimedia content with the main character. In this way, the user can be explicitly informed of the task to be completed at the beginning of the game, so as to guide the user to explore and discover in the interactive multimedia content and then complete the task.

[0050] In some embodiments, in a case where the type corresponding to the interactive multimedia content is an inference type, the interactive multimedia content can further include an event and a truth, the event indicating an event occurring in the interactive multimedia content, the truth indicating a truth associated with the event in the interactive multimedia content, and the task being to determine the truth based on the event. In some embodiments, the interactive multimedia content can further include predetermined media data. The predetermined media data can include, for example, pre-defined images, texts, videos, audios, etc. The predetermined media data can include static media data.

[0051] The static media data can include, for example, a name of the interactive multimedia content, a brief introduction, a background image, an image of a character (including an image of the main character and an image of another character (e.g., an NPC)), etc. The terminal device 110 can present the static media data in the interactive multimedia content. In this way, key information of the interactive multimedia content can be provided according to a pre-design of an author of the interactive multimedia content, so as to guide the user to perform a predetermined task in the interactive multimedia content.

[0052] In some embodiments, the interactive multimedia content can include one or more predetermined scenes. In a case where the interactive multimedia content involves an inference type, each scene can be associated with a respective node in the inference process. For example, in a game environment of “Mountain Rescue”, scenes of “Forest Cabin A” and “Camp B” can be provided. The user can interact with the game environment, and then select a scene to enter. Each scene can include a respective interactable entity, which can include a virtual character in the interactive multimedia content or a virtual object in the interactive multimedia data. In this way, the ability of the machine learning model can be invoked to automatically generate response data of the interactable entity to user interaction, so as to generate more rich interaction information.

[0053] The game player can communicate with the desired interactable entity, and thus obtain more abundant data. In the case that the type of the interactive multimedia content is a reasoning type, the virtual role can be a witness role appearing in the scene. For example, a tourist A, a forest ranger B, etc. appearing in the scene of "Cabin A in the Forest". The game player can interact with the tourist A, the forest ranger B, etc. and thus find more reasoning evidence from the testimony of the above-mentioned roles. For example, the name, occupation, age, testimony, etc. of the witness.

[0054] Alternatively and / or additionally, the virtual object can include an evidence associated with the scene. For example, a hiking pole, a water bottle, etc. appearing in the scene of "Cabin A in the Forest". The user can interact with the above-mentioned evidence and thus obtain more information about the evidence, such as the name of the evidence, the content of the evidence, the association of the evidence with other evidence, the discovery location of the evidence, and the image, etc.

[0055] In some embodiments, in the process of generating the dynamic media data, a viewing perspective of the scene desired to be viewed can be acquired, the viewing perspective corresponding to a virtual role in the interactive multimedia content. Based on the viewing perspective and the scene, perspective media data for viewing the scene is generated. Specifically, the user can be allowed to understand more abundant information of the scene from different perspectives. For example, assuming that the viewing perspective corresponds to the tourist A in the game environment, the cabin scene can be entered in the perspective of the tourist A, and relevant information of the scene is provided. For example, the route of the tourist A entering the cabin, the scenery and characters observed at the cabin, etc. In this way, the game player can be facilitated to obtain detailed information of each scene through the perspective of multiple roles, and thus find more clues, etc.

[0056] It should be understood that the dynamic media data is generated in real time based on the user input and the game content. At this time, the details of the generated dynamic media data will be different for different game players, and the main concept is the same. In this way, more abundant media data can be generated, and the situation of being able to only provide a fixed answer can be avoided.

[0057] Exemplarily, with reference to FIG. 3A, an example 300A shows an example of presenting an interactive page of an interactive multimedia content (i.e., example 300A). The example 300A can include an exit control 301. The terminal device 110 can determine that a user request of exiting the interactive multimedia content is received in response to receiving a user operation on the exit control, and the terminal device 110 can exit the interactive multimedia content (i.e., no longer present the example 300A) in turn. The example 300A includes a name of the interactive multimedia content (e.g., the text "Mountain Rescue" in the figure), a task in the interactive multimedia content (e.g., the text "Your task: Rescue the survivor trapped in the mountain" in the figure), avatars of a plurality of virtual characters in the interactive multimedia content (e.g., avatar 303, avatar 304, avatar 305, avatar 306, and avatar 307), and names, etc. The example 300A can also include a control 308. The terminal device 110 can present a clue found by the user in the interactive multimedia content in response to receiving a user operation on the control 308.

[0058] In some embodiments, the terminal device 110 can also determine that an interaction request indicating that the main character interacts with the target virtual character is received in response to receiving a user selection operation on the target virtual character. For example, the terminal device 110 can determine that the user wants to control the main character to interact with the character B in response to receiving a user selection operation on the image 304 in the example 300A. The terminal device 110 can present an example 300B as shown in FIG. 3B. The example 300B includes a return control 311. The terminal device 110 can determine to cancel the interaction of the main character with the target virtual character and return to present the example 300A in response to receiving a user operation on the return control.

[0059] In some embodiments, the terminal device 110 can present static media data in the interactive multimedia content. The static media data can include a predetermined script of a virtual character in the interactive multimedia content. For example, the example 300B can present an image of the character B and a text 312, which can be a predetermined script of the character B, for example. In some embodiments, the static media data can also include a portion of the static media data that matches the user input. Exemplarily, the terminal device 110 can present the static media data that matches the user input (e.g., present a predefined image that matches the user input, such as an image of a scene, an image of a character, an image of an evidence, etc.) to the user based on the received user input.

[0060] The example 300B can include an input panel 314 and an input box 313. The terminal device 110 may, for example, first present only the input box 313 in the example 300B, and present the input panel 314 in response to receiving a user operation on the input box 313. Alternatively or additionally, the terminal device 110 can also directly present the input box 313 and the input panel 314 together in the example 300B. The terminal device 110 can receive a user input via the input panel 314, and present the received user input in the input box 313. The terminal device 110 can select, based on the received user input, static media data matching the user input from predetermined static media data, and present the static media data to the user.

[0061] The terminal device 110 can also generate dynamic media data based on the interactive multimedia content and the user input in response to receiving the user input in the interactive multimedia content. The dynamic media data can include a video, an animated image, an audio, a text stream, and the like. The terminal device 110 can generate the dynamic media data in any suitable manner. For example, the terminal device 110 can generate the dynamic media data based on a predetermined algorithm or rule. For another example, the terminal device 110 can also generate the dynamic media data by means of a trained machine learning model. The machine learning model can be any suitable model, including but not limited to a CNN, an RNN, a transformer, and the like. In some embodiments, the machine learning model can be a language model (LM).

[0062] Specifically, the terminal device 110 can generate, in the interactive multimedia content, a prompt based on the interactive multimedia content and the user input. The terminal device 110 can provide the prompt to the machine learning model, so that the machine learning model generates a response to the prompt. The terminal device 110 can receive the response to the prompt from the machine learning model as the dynamic media data. In this way, the powerful processing capability of the machine learning model can be invoked, so that the dynamic media data is automatically generated conveniently and quickly. In this way, the response matching the user input can be generated under the constraint of the specific environment of the interactive multimedia content. In particular, the response other than the predetermined response in the interactive multimedia content can be provided, so that the question not specified in the interactive multimedia content is answered.

[0063] In some embodiments, the terminal device 110 can acquire background data of the virtual role in response to determining that the user input is a user input directed to a virtual role in the interactive multimedia content. The background data of the virtual role can indicate information such as a character setting, habits, etc. of the virtual role. For example, if the user input is a question directed to the role B, the terminal device 110 can acquire the background data of the role B. The terminal device 110 can generate response data to the user input as dynamic media data based on the background data. That is, the terminal device 110 can generate the response data to the user input from the perspective of the virtual role. The response data can be a text stream, a video, etc.

[0064] In some embodiments, the terminal device 110 can also acquire object attributes of a virtual object in response to determining that the user input is a user input directed to a virtual object in the interactive multimedia content. The virtual object can be, for example, an item in the interactive multimedia content. The object attributes of the virtual object can include, for example, a style, a shape, a material, a use, etc. of the virtual object. The terminal device 110 can generate response data to the user input as dynamic media data based on the object attributes, for example. In this way, the terminal device 110 can answer the question of the user input directed to the virtual object, and can improve the exploration of the interactive multimedia content by the user. In this way, the terminal device 110 can generate different dynamic media data based on different user inputs. The interactivity of the user can be improved.

[0065] The terminal device 110 can further present the dynamic media data in the interactive multimedia content. For example, the terminal device 110 can present the dynamic media data in an interaction page for presenting the interactive multimedia content. In some embodiments, if the user input is a user input directed to a virtual role, the terminal device 110 can present the dynamic media data with the virtual role, which includes a side role and a character role in the interactive multimedia content.

[0066] For example, in the example 300B, the terminal device 110 can receive a user input of the user directed to the role B via the input panel 314, for example. The terminal device 110 can generate a text stream to the user input based on the background data of the user B, and present the text stream with the role B in the example 300B. Alternatively or additionally, if the user input is a user input directed to a virtual object, the terminal device 110 can present the dynamic media data with a side role in the interactive multimedia content.

[0067] In some embodiments, the user can search for clues in the interactive multimedia content. The terminal device 110 may, for example, present example 300C in response to receiving a user operation on the control 302 in example 300A. Referring to FIG. 3C, in example 300C, the terminal device 110 may, for example, receive a user input of the user on the virtual object in the scenario of the cabin A in the forest, and generate a text stream for the user input based on the object attribute of the virtual object, and present the text stream with a voice-over character in example 300C.

[0068] Example 300C includes a control 320, and the terminal device 220 may, in response to receiving a user operation on the control 320, search for more clues about the cabin A in the forest. For example, the user can find relevant clues of the survivor through a dialogue with various characters in the interactive multimedia content. For example, based on the answer of the character B "I saw a hiker near the cabin A in the forest", the user can determine the clue "the hiker appeared near the cabin A in the forest", and so on.

[0069] Alternatively and / or additionally, the user can interact with the virtual object in the interactive multimedia content, and thus determine the clues of the survivor. For example, based on the discovery location of the virtual object "hiking stick" being near the cabin A in the forest upstream of the river, the user can determine the clue "the hiker appeared near the cabin A in the forest upstream of the river", and so on. The user can continue to find witnesses and evidence, and thus obtain more clues. At this time, the number of clues about the cabin A in the forest will increase. After completing the evidence search process, the user can click the control 322 to return to present example 300A.

[0070] In some embodiments, the terminal device 110 can also determine the state of the main character performing the task based on the user input (e.g., the first user input and the second user input). For example, referring back to FIG. 3A, example 300A can also include a control 309. Specifically, the user can click the control 309 to input the reasoning after finding all the clues. The terminal device 110 may, in response to receiving a user operation on the control 309, present a user page for receiving the reasoning result, and receive a user input indicating the reasoning result via the user page. The terminal device 110 can determine the state of the user's main character performing the task based on the received user input.

[0071] For example, if the task is "rescue the survivors trapped in the mountainous area", the terminal device 110 can determine the status of performing the task based on the comparison between the location and number of survivors indicated by the user input and the correct location and number of survivors. For example, if the location and number of survivors indicated by the user input are exactly the same as the correct location and number of survivors, the terminal device 110 can determine that the task has been completed. Specifically, a machine learning model can be invoked to determine whether the user has completed the task. In this way, compared with the conventional technical solution in which the user input is compared with the standard answer in text to determine whether the user has completed the task, the machine learning model can determine the status of the task in a more flexible and effective manner, and reduce the potential error risk caused by typos and unclear spoken language expressions in the user input.

[0072] In some embodiments, the terminal device 110 can further present statistical data of the user performing the task in response to determining that the status indicates that the user has completed the task. For example, the terminal device 110 can determine the ranking of the time length of the user completing the task among all records of the user, the ranking of the time length of the user completing the task in the sharing platform, and the like based on the time length of the user completing the task, the time length of the user historically completing the task, the time length of other users in the sharing platform completing the task, and the like.

[0073] Having described the various steps involved in presenting media data in interactive multimedia content, in the following, the overall flow of presenting media data in interactive multimedia content is described with reference to FIG. 3D. FIG. 3D illustrates a schematic diagram 300D of media data provided at different stages of performing a task in interactive multimedia content, according to some other embodiments of the present disclosure. As shown in FIG. 3D, a player 334 can enter the interactive multimedia content and complete a task in the interactive multimedia content. The static media data in FIG. 3D includes static media data 330-1, 330-2, 330-3, 330-4, 330-5, 330-6 (collectively referred to as 330), which are pre-generated in the interactive multimedia content. The corresponding dynamic media data 331-1, 331-2, 331-3 (collectively referred to as dynamic media data 331) can be generated based on the static media data and other information in the interactive multimedia content, with the aid of a machine learning model.

[0074] At the story introduction stage, content 332-1, e.g., a story title and a story introduction, can be presented in the interactive multimedia content based on the static media data 330-1. Further, at the plot introduction stage, the model can generate dynamic media data 331-1 based on the static media data 330-2, and in turn present content 332-2 in the interactive multimedia content (e.g., involving the static media data 330-3 and the dynamic media data 331-1).

[0075] In the evidence searching reasoning phase, the content 332-3 can be presented based on the static media data 330-4. In response to receiving the user input of the user with a role in the role list, the content 332-4 (e.g., a dialogue between the player and the NPC) can be presented; in response to receiving the user input of the user with the physical evidence, the content 332-5 (e.g., physical evidence details) can be presented. In response to receiving the reasoning of the user input, the dynamic media data 331-3 can be generated based on the static media data 330-6, and the content 332-6 is presented to provide the reasoning of the host judgment input whether correct. In the case of correct input result, the content 333 can be provided in the data statistics phase to provide various statistical data of the user during the execution of the task.

[0076] In this way, the static media data can be used to control the key information in the interactive multimedia content, and the real-time generated dynamic media data can be used to provide more rich information to the user. In this way, the machine learning model can process the situation that cannot be processed in the interactive multimedia content in a more flexible and effective way, thereby improving the authenticity of the interactive multimedia content.

[0077] In summary, according to the embodiments of the present disclosure, corresponding dynamic media data can be generated for different user inputs. In particular, the powerful processing capability of the machine learning model can be called to generate corresponding dynamic media data based on user input, personalized customization can be achieved, and the interactivity of the user in the interactive multimedia content can be improved.

[0078] Example process

[0079] The specific details of each step of the query have been described above, and a method of providing media data in interactive multimedia content is provided. FIG. 4 shows a flowchart of a process 400 of providing media data in interactive multimedia content according to some embodiments of the present disclosure. The process 400 can be implemented at the terminal device 110. The process 400 is described below with reference to FIG. 1.

[0080] At block 410, in response to receiving a request of a user of interactive multimedia content to enter the interactive multimedia content, the interactive multimedia content is presented. At block 420, in response to receiving a first user input of the user in the interactive multimedia content, a scene in the interactive multimedia content is selected, the scene including an interactive entity. At block 430, in response to receiving a second user input of the user for the interactive entity, dynamic media data is generated based on the interactive multimedia content and the second user input to respond to the second user input.

[0081] In some embodiments, the scene is defined by the interactive multimedia content, and the interactive entity includes a virtual role in the interactive multimedia content or a virtual object in the interactive multimedia data.

[0082] In some embodiments, generating the dynamic media data comprises: in response to determining that the second user input is a user input directed to a virtual character in the interactive multimedia content, obtaining background data of the virtual character; and generating response data to the second user input based on the background data as the dynamic media data.

[0083] In some embodiments, the type of the interactive multimedia content is a reasoning type, and the method further comprises: presenting the dynamic media data with a virtual character, the virtual character comprising a witness character and a sidekick character in the interactive multimedia content.

[0084] In some embodiments, the type of the interactive multimedia content is a reasoning type, and generating the dynamic media data comprises: in response to determining that the second user input is a user input directed to a virtual object in the interactive multimedia content, obtaining object attributes of the virtual object, the virtual object being an evidence object in the interactive multimedia content; and generating response data to the second user input based on the object attributes as the dynamic media data.

[0085] In some embodiments, the method further comprises: presenting the dynamic media data with a sidekick character in the interactive multimedia content.

[0086] In some embodiments, generating the dynamic media data comprises: obtaining a viewing perspective of a viewing scene, the viewing perspective corresponding to a virtual character in the interactive multimedia content; and generating perspective media data for viewing the scene based on the viewing perspective and the scene.

[0087] In some embodiments, generating the dynamic media data comprises: in the interactive multimedia content, generating a prompt word based on the interactive multimedia content and the second user input; and receiving a response to the prompt word as the dynamic media data.

[0088] In some embodiments, the interactive multimedia content comprises a main character and a task performed by the main character in the interactive multimedia content, and the method further comprises: providing description information of the main character and the task in the interactive multimedia content to guide the user to enter the interactive multimedia content with the main character.

[0089] In some embodiments, the interactive multimedia content comprises an event and a true image, the event indicating an event occurring in the interactive multimedia content, the true image indicating a truth associated with the event in the interactive multimedia content, and the task being to determine the truth based on the event.

[0090] In some embodiments, generating the dynamic media data further comprises: determining a state of the main character performing the task based on the first user input and the second user input.

[0091] In some embodiments, the method further includes presenting statistical data of the user performing the task in response to determining that the status indicates that the user has completed the task.

[0092] In some embodiments, the interactive multimedia content includes predetermined media data, and the method further includes presenting static media data in the interactive multimedia content, the static media data being a portion of the static media data that matches the first user input and the second user input.

[0093] In some embodiments, the method is performed in a sharing platform for sharing the interactive multimedia content, and the interactive multimedia content is published in the sharing platform.

[0094] Example apparatus and device

[0095] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 5 shows a schematic structural block diagram of an apparatus 500 for providing media data in interactive multimedia content, according to some embodiments of the present disclosure. The apparatus 500 can be implemented as or included in the terminal device 110. Various modules / components in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.

[0096] As shown in FIG. 5, the apparatus 500 includes a presentation module 510 configured to present the interactive multimedia content in response to receiving a request of a user to enter the interactive multimedia content; a selection module 520 configured to select a scene in the interactive multimedia content in response to receiving a first user input of the user in the interactive multimedia content, the scene including an interactable entity; and a generation module 530 configured to generate dynamic media data based on the interactive multimedia content and a second user input of the user for the interactable entity in response to receiving the second user input.

[0097] In some embodiments, the scene is defined by the interactive multimedia content, and the interactable entity includes a virtual character in the interactive multimedia content or a virtual object in the interactive multimedia data.

[0098] In some embodiments, the generation module is further configured to, in response to determining that the second user input is a user input for a virtual character in the interactive multimedia content, obtain background data of the virtual character; and generate response data for the second user input as the dynamic media data based on the background data.

[0099] In some embodiments, the type of the interactive multimedia content is a reasoning type, and the apparatus further includes a presentation module configured to present the dynamic media data with a virtual character, the virtual character including a sidekick character and a witness character in the interactive multimedia content.

[0100] In some embodiments, the type of the interactive multimedia content is a reasoning type, and the generating module is further configured to: in response to determining that the second user input is a user input for a virtual object in the interactive multimedia content, acquire an object attribute of the virtual object, the virtual object being a piece of evidence object in the interactive multimedia content; and based on the object attribute, generate the response data for the second user input as the dynamic media data.

[0101] In some embodiments, the apparatus further comprises a presenting module configured to present the dynamic media data with a side role in the interactive multimedia content.

[0102] In some embodiments, the generating module is further configured to: acquire a viewing perspective of a viewing scene, the viewing perspective corresponding to a virtual role in the interactive multimedia content; and based on the viewing perspective and the scene, generate perspective media data for viewing the scene.

[0103] In some embodiments, the generating module is further configured to: in the interactive multimedia content, generate a prompt word based on the interactive multimedia content and the second user input; and receive a response for the prompt word as the dynamic media data.

[0104] In some embodiments, the interactive multimedia content comprises a main role and a task performed by the main role in the interactive multimedia content, and the apparatus further comprises a providing module configured to provide description information of the main role and the task in the interactive multimedia content, so as to guide the user to enter the interactive multimedia content with the main role.

[0105] In some embodiments, the interactive multimedia content comprises an event and a true image, the event indicating an event occurring in the interactive multimedia content, the true image indicating a truth associated with the event in the interactive multimedia content, and the task being to determine the truth based on the event.

[0106] In some embodiments, the apparatus further comprises a state determining module configured to determine a state of the main role performing the task based on the first user input and the second user input.

[0107] In some embodiments, the apparatus further comprises a data presenting module configured to present statistical data of the user performing the task in response to determining that the state indicates that the user completes the task.

[0108] In some embodiments, the interactive multimedia content comprises predetermined media data, and the apparatus further comprises a static data presenting module configured to present static media data in the interactive multimedia content, the static media data being a portion in the static media data that matches the first user input and the second user input.

[0109] In some embodiments, the apparatus is invoked in a sharing platform for sharing interactive multimedia content, and the interactive multimedia content is posted in the sharing platform.

[0110] The units and / or modules included in the apparatus 500 can be implemented utilizing various means including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine-executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 500 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.

[0111] It should be understood that one or more steps in the above methods can be performed by an appropriate electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices can include, for example, the terminal device 110 in FIG. 1.

[0112] FIG. 6 illustrates a block diagram of an electronic device 600 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 600 illustrated in FIG. 6 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 600 illustrated in FIG. 6 can be used to implement the terminal device 110 of FIG. 1 or the apparatus 500 of FIG. 5.

[0113] As shown in FIG. 6, the electronic device 600 is in the form of a general electronic device. Components of the electronic device 600 can include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 can be a real or virtual processor and is capable of performing various processes according to programs stored in the memory 620. In a multi-processor system, multiple processing units perform computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.

[0114] The electronic device 600 typically includes a plurality of computer storage media. Such media can be any available media that is located either internally or externally to the electronic device 600, including, but not limited to, memory, removable storage, and non-removable storage. The memory 620 can be volatile (such as, for example, registers, cache, RAM), non-volatile (such as, for example, ROM, EEPROM, flash memory), or some combination of the two. The storage 630 can be removable or non-removable media, and can include machine- readable media, such as flash drives, disk drives, or any other media that can be used to store information and / or data and that can be accessed by the electronic device 600.

[0115] The electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, a disk drive for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM) can be provided. In such instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 620 can include a computer program product 625 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0116] The communication unit 640 enables communications with other electronic devices over a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with one another over a communication connection. As such, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.

[0117] The input device 650 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 640, as needed, one or more devices that enable a user to interact with the electronic device 600, or any device (e.g., a network card, a modem, etc.) that enables the electronic device 600 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0118] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0119] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0120] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0121] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0123] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for providing media data, comprising: presenting an interactive multimedia content in response to receiving a request of a user to enter the interactive multimedia content; selecting a scene in the interactive multimedia content in response to receiving a first user input of the user in the interactive multimedia content, the scene comprising an interactable entity; and generating dynamic media data based on the interactive multimedia content and a second user input of the user in response to the second user input in response to receiving the second user input for the interactable entity. 2.The method of claim 1, wherein the scene is defined by the interactive multimedia content, and the interactable entity comprises a virtual character in the interactive multimedia content or a virtual object in the interactive multimedia data. 3.The method of claim 2, wherein generating the dynamic media data comprises: obtaining background data of the virtual character in response to determining that the second user input is a user input for a virtual character in the interactive multimedia content; and generating response data for the second user input as the dynamic media data based on the background data. presenting the dynamic media data with the virtual character, the virtual character comprising a narrator character and a witness character in the interactive multimedia content.

4. The method of claim 3, wherein the type of the interactive multimedia content is a reasoning type, the method further comprising: 5.The method of claim 1, wherein the interactive multimedia content is of an inference type, and generating the dynamic media data comprises: obtaining object attributes of a virtual object in response to determining that the second user input is a user input for the virtual object in the interactive multimedia content, the virtual object being an evidence object in the interactive multimedia content; and generating response data for the second user input as the dynamic media data based on the object attributes. presenting the dynamic media data with a narrator character in the interactive multimedia content. 7.The method of claim 1, wherein generating the dynamic media data comprises:

6. The method of claim 5, wherein the method further comprises: obtaining a viewing perspective for viewing the scene, the viewing perspective corresponding to a virtual character in the interactive multimedia content; and generating perspective media data for viewing the scene based on the viewing perspective and the scene. in the interactive multimedia content, generating a hint word based on the interactive multimedia content and the second user input; and receiving a response for the hint word as the dynamic media data.

8. The method of claim 1, wherein generating dynamic media data comprises: providing description information of the main character and the task in the interactive multimedia content so as to guide the user to enter the interactive multimedia content with the main character. 10.The method of claim 9, wherein the interactive multimedia content comprises an event and a true image, the event indicating an event occurring in the interactive multimedia content, the true image indicating a true fact associated with the event in the interactive multimedia content, and the task being to determine the true fact based on the event. ​ ​ 9. The method of claim 1, wherein the interactive multimedia content includes a primary character and a task performed by the primary character in the interactive multimedia content, the method further comprising: ​ ​ 11. The method of claim 9, wherein generating the dynamic media data further comprises: determine, based on the first user input and the second user input, a status of the primary character performing the task.

12. The method of claim 11, further comprising: present, in response to determining that the status indicates that the user completed the task, statistics of the user performing the task.

13. The method of claim 1, wherein the interactive multimedia content comprises predetermined media data, and the method further comprises: present static media data in the interactive multimedia content, the static media data being a portion of the static media data that matches the first user input and the second user input. 14.The method of claim 1, wherein the method is performed in a sharing platform for sharing the interactive multimedia content, and the interactive multimedia content is published in the sharing platform. 15.An apparatus for providing media data, comprising: a presentation module configured to present the interactive multimedia content in response to receiving a request of a user of the interactive multimedia content to enter the interactive multimedia content; a selection module configured to select a scene in the interactive multimedia content in response to receiving a first user input of the user in the interactive multimedia content, the scene including an interactable entity; and a generation module configured to generate dynamic media data based on the interactive multimedia content and a second user input of the user for the interactable entity in response to receiving the second user input to respond to the second user input. 16.An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1-14. 17.A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1-14. 18.A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-14.