Display device and interaction method
Patent Information
- Application Number
- PCT/CN2025/134518
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-06-26
- Filing Date
- 2025-11-12
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025134518_01102026_PF_FP_ABST
Abstract
Description
A display device and an interaction method
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese patent applications filed on March 28, 2025, application number 202510387903.8, and on June 26, 2025, application number 202510875417.0, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of display device technology, and more particularly to a display device and an interaction method. Background Technology
[0004] In applications where users watch video media assets, to enhance the user experience and enjoyment, human-computer interaction functions are typically supported during video playback, allowing for interaction based on the video media assets being played. However, in related technologies, users can usually only passively receive interactive recommendations at preset playback times, resulting in a lack of personalized and interactive features due to the limited interaction methods. Summary of the Invention
[0005] According to some embodiments of this application, a display device is provided, including: a display, a memory, and at least one processor; the memory is configured to store computer programs or instructions; the at least one processor is connected to the display and the memory, and is configured to execute the computer program or instructions to cause the display device to: control the display to display a media asset details page associated with the currently playing media asset; wherein the media asset details page displays a digital human interactive control; in response to a trigger operation on the digital human interactive control, control the display to display digital human cards of candidate digital humans associated with the currently playing media asset; wherein the digital human card of each candidate digital human contains a digital human image corresponding to the candidate digital human and / or the character name in the currently playing media asset represented by the candidate digital human; after a first digital human's digital human card is selected, receive input user interaction voice, wherein the first digital human is one of the candidate digital humans; based on the user interaction voice, obtain a first response voice from the first digital human in response to the user interaction voice from a server, and play the first response voice; wherein the timbre in the first response voice is the timbre of the character name corresponding to the first digital human.
[0006] According to some embodiments of this application, an interactive method is provided, applied to a display device, comprising: controlling a display to show a media asset details page associated with a currently playing media asset; wherein the media asset details page displays a digital human interactive control; responding to a trigger operation on the digital human interactive control, controlling the display to show digital human cards of candidate digital humans associated with the currently playing media asset; wherein the digital human card of each candidate digital human contains a digital human image corresponding to the candidate digital human and / or the name of a character in the currently playing media asset represented by the candidate digital human; after a first digital human's digital human card is selected, receiving input user interaction voice, wherein the first digital human is one of the candidate digital humans; based on the user interaction voice, obtaining a first response voice from the first digital human in response to the user interaction voice from a server, and playing the first response voice; wherein the timbre in the first response voice is the timbre of the character name corresponding to the first digital human.
[0007] According to some embodiments of this application, an interactive device is provided, configured on a display device, including: a first control unit, configured to control the display to show a media asset details page associated with the currently playing media asset; wherein the media asset details page displays a digital human interactive control; a second control unit, configured to control the display to show digital human cards of candidate digital humans associated with the currently playing media asset in response to a trigger operation on the digital human interactive control; wherein each candidate digital human card contains a digital human image corresponding to the candidate digital human and / or the name of a character in the currently playing media asset represented by the candidate digital human; a receiving unit, configured to receive input user interaction voice after the digital human card of the first digital human is selected, wherein the first digital human is one of the candidate digital humans; and a playback unit, configured to obtain a first response voice from the first digital human in response to the user interaction voice from a server, and play the first response voice; wherein the timbre in the first response voice is the timbre of the character name corresponding to the first digital human.
[0008] According to some embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that executes the above-described interactive method when the computer program is run by a processor.
[0009] According to some embodiments of this application, a computer program product is provided, the computer program product comprising: a computer program that executes the above-described interactive method when the computer program processor is running. Attached Figure Description
[0010] Figure 1 is a schematic diagram of an operation scenario between a display device and a control device according to some embodiments of this application;
[0011] Figure 2 is a hardware configuration block diagram of a control device provided according to some embodiments of this application;
[0012] Figure 3 is a hardware configuration block diagram of a display device provided according to some embodiments of this application;
[0013] Figure 4 is a software configuration block diagram of a display device provided according to some embodiments of this application;
[0014] Figure 5 is a flowchart illustrating the interactive methods provided in some embodiments of this application;
[0015] Figure 6 is a schematic diagram of digital human controls provided in some embodiments of this application;
[0016] Figure 7 is a schematic diagram of a digital human card provided in some embodiments of this application;
[0017] Figure 8 is a flowchart illustrating the display of media asset details page according to some embodiments of this application;
[0018] Figure 9 is a schematic diagram of the process of obtaining the first response voice according to some embodiments of this application;
[0019] Figure 10 is a flowchart illustrating the method for processing user interactive voice in a voice service application provided in some embodiments of this application;
[0020] Figure 11 is a flowchart illustrating the return notification method provided in some embodiments of this application;
[0021] Figure 12 is a schematic diagram of the process of playing the prompt voice of the third digital human provided in some embodiments of this application;
[0022] Figure 13 is a timing diagram of training a digital human model provided in some embodiments of this application;
[0023] Figure 14 is an interactive timing diagram provided by some embodiments of this application;
[0024] Figure 15 is a flowchart illustrating an interactive method provided according to some embodiments of this application;
[0025] Figure 16 is a schematic diagram of an interactive card provided according to some embodiments of this application;
[0026] Figure 17 is a schematic diagram of interactive cards and character introduction cards provided according to some embodiments of this application;
[0027] Figure 18 is a schematic diagram of the interface in dialogue mode provided according to some embodiments of this application;
[0028] Figure 19 is a schematic diagram of an interface in a dialog mode according to some other embodiments of this application;
[0029] Figure 20 is a timing diagram of an interaction method provided according to some embodiments of this application;
[0030] Figure 21 is a timing diagram of an interaction method provided according to other embodiments of this application;
[0031] Figure 22 is a schematic diagram of an interactive device provided according to some embodiments of this application;
[0032] Figure 23 is a schematic diagram of an interactive device provided according to some embodiments of this application. Detailed Implementation
[0033] The solutions in the embodiments of this application will now be described with reference to the accompanying drawings. When referring to the drawings in the following description, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The drawings described below are merely some embodiments of this application; those skilled in the art can obtain other related drawings based on these drawings without creative effort. The implementation methods described in the following embodiments do not represent all implementation methods consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.
[0034] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0035] The display device provided in this application refers to any device with screen display and data processing capabilities, and can take many forms, such as a television, smart television, mobile terminal, computer, monitor, advertising screen, wearable device, virtual reality device, augmented reality device, laser projection device, monitor, electronic bulletin board, etc. Figures 1 and 2 show a specific embodiment of the display device of this application.
[0036] Figure 1 is a schematic diagram of an operation scenario between a display device and a control device according to some embodiments of this application. As shown in Figure 1, a user can operate the display device 200 through the smart device 300 or the control device 100.
[0037] In some embodiments, the control device 100 may be a remote control, stylus, gamepad, etc. Communication between the remote control and the display device includes infrared or Bluetooth communication, as well as other short-range communication methods, to control the display device 200 wirelessly or via wired means. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, or control panel input.
[0038] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) can also be used to control the display device 200. For example, an application running on the smart device can be used to control the display device 200. Audio and video content displayed on the smart device 300 can also be transmitted to the display device 200 to achieve synchronized display.
[0039] In some embodiments, the display device 200 also communicates with the server 400 via various communication methods.
[0040] Figure 2 is a hardware configuration block diagram of a control device according to some embodiments of this application. In some embodiments, as shown in Figure 2, the control device 100 includes a control component 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive operation commands input by the user and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.
[0041] In some embodiments, the control component 110 includes a processor 112, RAM 113 and ROM 114, a communication interface 130, and a communication bus. The control component 110 is used to control the operation and function of the control device 100, as well as communication and cooperation between internal components and external and internal data processing functions.
[0042] In some embodiments, under the control of the control component 110, the communication interface 130 enables communication of control signals and data signals with the display device 200. The communication interface 130 may include at least one of other near-field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC module 133.
[0043] In some embodiments, the user input / output interface 140 includes at least one of other input interfaces such as a microphone 141, a touchpad 142, a sensor 143, and a button 144.
[0044] As shown in Figure 3, in some embodiments, the display device 200 includes at least one of a tuner 210, a communication device 220, a detector 230, an external device interface 240, at least one processor 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.
[0045] In some embodiments, the user can input user commands through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the graphical user interface (GUI). Alternatively, the user can input user commands by inputting specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.
[0046] As shown in Figure 4, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the Android runtime and system library layer (referred to as the "System Runtime Library Layer"), and the kernel layer.
[0047] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the function of the display device, the type of operating system, and other factors, the number of levels and the specific level type of the operating system may be expressed in other forms.
[0048] With the development of artificial intelligence technology, display devices are no longer limited to simply displaying images and playing videos. They can now support human-computer interaction during image display or video playback to enhance the user experience and enjoyment. Taking video playback as an example, current technologies typically only allow users to passively receive interactive recommendations during preset video playback times, resulting in a lack of personalized and interactive features due to limited interaction options.
[0049] Based on this, embodiments of this application provide a display device and an interaction method. The interaction method is applied to the display device, which includes a display, a memory, and at least one processor. The memory is configured to store a computer program; the at least one processor is configured to execute the computer program to enable the display device to: control the display to show a media asset details page associated with the currently playing media asset. The media asset details page displays digital human interactive controls, providing a bridge for user interaction with the digital human. During video playback and user interaction, the user no longer passively receives interactive recommendations but can actively interact with the digital human through the digital human interactive controls on the media asset details page. This enhances the user's initiative in interacting with the digital human; consequently, it can respond to trigger operations on the digital human interaction controls, controlling the display to show the digital human cards of candidate digital humans associated with the currently playing media asset; and after the first digital human's digital human card is selected, it receives the input user interaction voice, and based on the user interaction voice, retrieves the first digital human's first response voice from the server in response to the user interaction voice, and plays the first response voice with the timbre of the character name corresponding to the first digital human. In this way, the first response voice not only matches the user interaction voice, but also combines with the media asset content of the currently playing media asset, thereby effectively improving the intelligence and fun of the interaction between the digital human and the user.
[0050] Optionally, referring to Figure 5, which is a flowchart illustrating an interaction method according to some embodiments of this application, the following description will use the application of this method to a display device as an example. This interaction method includes, but is not limited to, the following steps:
[0051] S501 controls the display to show the details page of the media asset associated with the currently playing media asset.
[0052] In some embodiments, the currently playing media asset can be digital audio and video content owned by a television station, video platform, or other content organization, i.e., video media assets, such as movies, TV series, documentaries, animations, short videos, etc. One or more media asset playback windows can be displayed in the navigation bar of the display device. When a user selects a media asset and clicks on its playback window, the selected media asset becomes the currently playing media asset. Specifically, the user can select a media asset by controlling the control device 100 to focus on it; they can also select a media asset by controlling the smart device 300 to focus on it; or they can select the first media asset and trigger its playback window by using the voice assistant on the display device, for example, by voice inputting "play the first media asset".
[0053] In some embodiments, a user can focus on a media asset by controlling the control device 100 or the smart device 300, and then control the control device 100 or the smart device 300 to click on the playback window of that media asset. After clicking on the playback window, the display device sends a request to the server to retrieve the media asset details page associated with the currently playing media asset. This request may carry the media asset identifier of the currently playing media asset. The server can retrieve the details page data of the currently playing media asset based on the media asset identifier. For example, the details page data of the playing media asset can be obtained from a third-party platform, or it can be obtained by calling the port corresponding to the currently playing media asset.
[0054] Furthermore, the server can send detail page data back to the display device. This detail page data may include, but is not limited to, a summary of the media asset's content, the media asset's cover image, and related data. Taking a TV series as an example, the detail page data could include a plot synopsis, title, cover image, episode numbers, cast list, ratings, digital character information, and related reviews. The digital character information could include details such as the character's appearance within the TV series.
[0055] In some embodiments, after receiving the details page data associated with the currently playing media asset, the display device can render the page based on the details page data. After the page rendering is complete, the display device controls the monitor to display the details page of the media asset associated with the currently playing media asset. The details page displays a digital human interactive control. The digital human interactive control is a control that allows the user to interact with the digital human associated with the currently playing media asset.
[0056] In some embodiments, the digital human interaction control may display a target image, which may be an independent digital human image or an aggregated digital human image. An independent digital human image may be an image of any character in the currently playing media asset; an aggregated digital human image may be an aggregated image of multiple characters in the currently playing media asset.
[0057] When the target image is an independent digital human image, the independent digital human image can be the digital human image corresponding to a specified digital human among the candidate digital humans, where the specified digital human can be the first candidate digital human in the display order of digital humans. Alternatively, it can be the digital human image corresponding to a digital human randomly determined from the candidate digital humans. Candidate digital humans can be a collection of all or some of the characters in the currently playing media asset.
[0058] Referring to Figure 6, which is a schematic diagram of a digital human control according to some embodiments of this application, the digital human control can be presented in the form of images, text, or a combination of images and text on the media asset details page. For example, taking an animated film as an example, the digital human interactive control can display an image of a digital human, accompanied by the text "AI Interaction." This allows for a clear presentation of the digital human interactive control and a direct demonstration of its functionality.
[0059] S502, in response to a trigger operation on the digital human interactive control, controls the display to show the digital human card of the candidate digital human associated with the currently playing media asset.
[0060] In some embodiments, users can trigger the digital human interaction control by manipulating the control device 100 or the smart device 300. Users can also trigger the digital human interaction control by far-field voice. For example, users can voice input "open digital human interaction control" to trigger the digital human interaction control.
[0061] Upon triggering the digital human interaction control, the display device responds to the trigger operation by controlling the display to show the digital human cards of the candidate digital humans associated with the currently playing media asset. If there is only one digital human, the digital human card can be displayed directly. If there are multiple digital humans, the display can be controlled to show the digital human cards of the candidate digital humans associated with the currently playing media asset in the order in which the digital humans are displayed, or in a random display order.
[0062] The digital human cards can be generated based on digital human information. This information may include digital human images of candidate digital humans associated with the currently playing media asset and the names of the characters they represent in the current media asset. Each candidate digital human's digital human card contains either the candidate digital human's corresponding digital human image or the name of the character the candidate digital human represents in the current media asset; alternatively, each candidate digital human's digital human card may contain both the candidate digital human's corresponding digital human image and the name of the character the candidate digital human represents in the current media asset. For example, if the current media asset is an animated film, and the characters in the current media asset include Qiqi, Miaomiao, and Tutu, then Qiqi's digital human card may include Qiqi's character image and name, and may also include a description or greeting for Qiqi.
[0063] Referring to Figure 7, which is a schematic diagram of a digital human card according to some embodiments of this application, the digital humans Qiqi, Miaomiao, and Tutu are used as examples for illustration. Each digital human card may include a digital human image corresponding to the candidate digital human, the name of the character in the currently playing media represented by the candidate digital human, and may also include an introduction and display number of the digital human.
[0064] Figure 7 shows the digital humans displayed in the order of Qiqi, Miaomiao, and Tutu. The introduction on Qiqi's digital human card could be, "I am Qiqi from the Panda family, and I am 7 years old." Miaomiao's digital human card could be, "I am Miaomiao from the Panda family, and I am 6 years old." Tutu's digital human card could be, "I am Tutu from the Panda family, and I am 5 years old." These introductions on the digital human cards are for illustrative purposes only and are not intended to limit the types of introductions that can be used. The introductions on the digital human cards can be set according to actual needs.
[0065] S503, after the first digital human's digital human card is selected, receives input user interaction voice.
[0066] In some embodiments, the first digital person is the selected digital person from the candidate digital persons, that is, the first digital person can be one of the candidate digital persons. The candidate digital persons can be a collection of all or some of the characters in the currently playing media asset.
[0067] Users can select the first digital person's card by operating the control device 100 or the smart device 300. For example, they can select the first digital person by focusing on the card of the first digital person using the control device 100 or the smart device 300. Alternatively, users can select the first digital person by voice input; for example, they can select the first digital person by voice inputting the name of the character corresponding to the first digital person.
[0068] In some embodiments, after the first digital human's digital human card is selected, the user can input user interaction voice, for example, the user's interaction voice could be "How's the weather today?"; the display device receives the user's input user interaction voice.
[0069] S504, based on the user's interactive voice, obtain the first response voice of the first digital human in response to the user's interactive voice from the server, and play the first response voice.
[0070] Then, the user's interactive voice and the digital human identifier of the first digital human can be sent to the server, so that the server can generate a first response voice based on the user's interactive voice and the digital human identifier, and feed the first response voice back to the display device so that the display device can play the first response voice. The timbre in the first response voice is the timbre of the character name corresponding to the first digital human.
[0071] Optionally, while playing the first reply voice message, the corresponding text can also be visualized to improve the user's viewing and interactive experience.
[0072] In the above embodiments, by displaying the media asset details page associated with the currently playing media asset, which displays digital human interaction controls, a bridge is provided for user interaction with the digital human. Therefore, during video playback, users no longer passively receive interactive recommendations but can actively interact with the digital human through the digital human interaction controls on the media asset details page, increasing the user's initiative in interacting with the digital human. Furthermore, in response to the trigger operation of the digital human interaction controls, digital human cards of candidate digital humans associated with the currently playing media asset are displayed. Since different candidate digital humans represent different characters in the currently playing media asset, subsequent voice interactions are more closely aligned with the currently playing media asset, and a foundation is provided for subsequent personalized voice interactions. In addition, after the first digital human's digital human card is selected, the system can receive input user interaction voice and, based on the user interaction voice, obtain the first response voice from the server in response to the user interaction voice, playing the first response voice with the timbre of the character name corresponding to the first digital human, enhancing the fun of interaction between the digital human and the user.
[0073] In some optional implementations, when controlling the display to show the details page of the media asset associated with the currently playing media asset, it can first be determined whether there is digital human information in the details page data of the currently playing media asset. If digital human information exists, a digital human interactive control can be generated to support users to interact with the digital human based on the digital human interactive control.
[0074] Based on this, referring to Figure 8, which is a flowchart illustrating a display of media asset details page associated with the currently playing media asset according to some embodiments of this application, specifically including the following steps:
[0075] S801 sends a request to the server to retrieve the details page for the currently playing media asset.
[0076] Optionally, one or more media playback windows can be displayed in the navigation bar of the display device. When a user selects a media asset and clicks on the playback window of that media asset, the selected media asset becomes the currently playing media asset.
[0077] Users can focus on a media asset using control device 100 or smart device 300, and then click on the playback window of that media asset. After clicking the playback window, the display device sends a request to the server to retrieve the media asset details page associated with the currently playing media asset. This request may carry the media asset identifier of the currently playing media asset. The server can then retrieve the details page data of the currently playing media asset based on its media asset identifier. For example, the details page data can be retrieved from a third-party platform, or it can be obtained by calling the port corresponding to the currently playing media asset.
[0078] S802, Receive detail page data from the server based on the detail page retrieval request.
[0079] Then, the server sends the details page data back to the display device. The display device can receive the details page data sent back by the server based on the details page retrieval request.
[0080] S803, if the details page data contains digital human information, generates a media asset details page containing digital human interactive controls based on the details page data, and controls the display to show the media asset details page.
[0081] In some embodiments, when the details page data contains digital human information, the display device can render the page based on the details page data to generate a media asset details page containing digital human interactive controls. The display device can then control the monitor to display the media asset details page. In this case, the media asset details page displays digital human interactive controls. These digital human interactive controls are controls that allow the user to interact with the digital human associated with the currently playing media asset.
[0082] In the above embodiments, when the details page data of the currently playing media asset contains digital human information, a media asset details page containing digital human interactive controls can be generated, providing a bridge for users to interact with digital humans, so that users can actively interact with digital humans and improve the initiative of users to interact with digital humans.
[0083] In some alternative implementations, when displaying the digital human cards of candidate digital humans associated with the currently playing media asset, the display can be controlled to show the digital human cards of candidate digital humans associated with the currently playing media asset in the order of digital human display.
[0084] In some embodiments, the display order of digital humans may be randomly generated by the server or by the display device; it may also be determined by the server based on at least one of the following: the popularity information of candidate digital humans, the number of historical interaction selections, and the matching degree between candidate digital humans and the currently played media; or it may be determined by the display device based on at least one of the following: the popularity information of candidate digital humans, the number of historical interaction selections, and the matching degree between candidate digital humans and the currently played media.
[0085] The popularity information of candidate digital humans can be determined based on the number of topics, discussion frequency and number of participants related to the candidate digital human on social platforms. Candidate digital humans with higher popularity can be displayed first.
[0086] The number of historical interaction selections can be understood as the number of times a user interacts with a digital human within a historical period. This historical period can be a month, half a month, etc., and can be set according to the actual situation. Candidate digital humans with higher historical interaction selection counts can be displayed first.
[0087] The matching degree between candidate digital humans and the currently playing media assets can be understood as the relevance of the candidate digital human to the episodes and plots of the currently playing media assets. The higher the relevance of the candidate digital human, the higher its matching degree with the currently playing media assets. Candidate digital humans with a high matching degree with the currently playing media assets can be displayed first.
[0088] In some embodiments, after the control display shows the digital human card of the candidate digital human associated with the currently playing media asset, the interactive prompts of the first digital human can also be obtained and played after the digital human card of the first digital human is selected.
[0089] In some possible embodiments, a user can select the first digital human's card by manipulating the control device 100 or the smart device 300. For example, the user can select the first digital human by focusing the attention on the card of the first digital human through the control device 100 or the smart device 300. Alternatively, the user can select the first digital human by voice input; for example, the user can select the first digital human by voice inputting the name of the character corresponding to the first digital human.
[0090] After the first digital human card is selected, the interactive prompts of the first digital human can be obtained and played. Taking "Qiqi" in Figure 7 as an example, the interactive prompts of the first digital human can be "Hi, I am Qiqi. I am so happy to have learned new knowledge today. I will do my own thing. Come and chat with me."
[0091] In some embodiments, when playing the introductory text of the first digital human, the voice of the character's name corresponding to the first digital human can be used for playback, and the text corresponding to the introductory text of the first digital human can be visualized.
[0092] In the above embodiments, on the one hand, displaying the digital human cards of candidate digital humans associated with the currently playing media asset according to the display order of digital humans can make the digital humans displayed first more likely to meet the user's interaction needs; on the other hand, playing the interactive guidance text of the first digital human after the first digital human's digital human card is selected can increase the user's enthusiasm for interaction.
[0093] In some alternative implementations, when playing the first response voice to the user's interactive voice, the voice can be played based on the currently selected digital persona, or it can be played based on the digital persona selected by the user.
[0094] In some possible embodiments, when the user interaction voice does not contain digital human association information, a first response voice from the first digital human in response to the user interaction voice can be obtained from the server and played. The digital human association information can be information included in the user interaction voice for selecting the digital human; for example, it could be the character name corresponding to the digital human in the user interaction voice, or the display sequence number of the digital human in the display order. When the user interaction voice does not contain digital human association information, the first response voice can be played using the currently selected first digital human.
[0095] In some other possible embodiments, if the user interaction voice contains digital human association information, and the second digital human represented by the digital human association information is different from the first digital human, then the screen focus is focused on the digital human card corresponding to the second digital human, and according to the user interaction voice, the second response voice of the second digital human in response to the user interaction voice is obtained from the server and played; wherein, the timbre in the second response voice is the timbre of the character name corresponding to the second digital human.
[0096] When the user interaction voice contains digital human association information, such as the name of the corresponding character or the display sequence number of the digital human, it indicates that the user interaction voice contains digital human association information. In this case, we can first determine whether the second digital human represented by the association information is the same as the currently selected first digital human. If the second digital human is different from the first, we can focus the screen on the digital human card corresponding to the second digital human, retrieve the second response voice from the server based on the user interaction voice, and play the second response voice; this ensures that the timbre of the second response voice is the same as the character name of the second digital human. If the second digital human is the same as the first digital human, we can retrieve the first response voice from the server based on the user interaction voice and play the first response voice.
[0097] In the above embodiments, when inputting user interaction voice, digital human association information can be carried to select a digital human to respond in a personalized way, making the process of user interaction with digital human more personalized and improving the user's initiative in interacting with digital human.
[0098] In some alternative implementations, when obtaining the first response voice from the first digital human in response to the user's interactive voice from the server, the user interaction text and digital human identifier corresponding to the user interaction voice can be sent to the server so that the server can determine the first response voice based on the user interaction text and digital human identifier.
[0099] Based on this, referring to Figure 9, which is a schematic diagram of a process for obtaining a first response voice according to some embodiments of this application, specifically including but not limited to the following steps:
[0100] S901 converts user-interacting voice into user-interacting text.
[0101] Before retrieving the relevant voice response from the server, the user's voice interaction can first be converted into text. This conversion can be achieved in several ways. For example, a speech-to-text tool can be used; alternatively, the voice interaction can be input into a speech recognition model to obtain the corresponding text. These methods are merely illustrative and do not constitute a limitation on the methods used for converting voice interaction into text.
[0102] S902, send a first query request to the server carrying user interaction text and the digital human identifier of the first digital human.
[0103] Then, a first query request carrying user interaction text and the digital human identifier of the first digital human can be sent to the server. The digital human identifier is used to distinguish different digital humans, and each digital human has a unique digital human identifier. The first query request instructs the server to generate a first response voice based on the user interaction text and the digital human identifier of the first digital human.
[0104] After receiving the first query request, the server can obtain the corresponding response text information based on the user's interaction text. Then, based on the digital human's identifier, it can locate the corresponding digital human and invoke it to generate the first response voice based on the user's interaction text. For example, the response text information can first be converted into response voice data, and then the timbre in the response voice data can be adjusted to match the timbre of the first digital human to obtain the first response voice.
[0105] S903, receive the first response voice from the server based on the first query request.
[0106] The first response voice can be determined by the server based on the digital human identifier of the first digital human and by calling the first digital human based on the user's interactive text.
[0107] The server sends the first response voice message to the display device. The display device receives the first response voice message from the server based on the first query request. This first response voice message can be determined by the server based on the digital human identifier of the first digital human and by calling the first digital human based on the user's interactive text. In this way, the timbre in the first response voice message can be the timbre of the character name corresponding to the first digital human.
[0108] In the above embodiments, when obtaining the first response voice, a first query request carrying the user interaction text and the digital human identifier of the first digital human is sent to the server, so that the server generates the first response voice based on the user interaction text and the digital human identifier of the first digital human. The timbre in the first response voice is the timbre of the character name corresponding to the first digital human, making the first response voice more consistent with the digital human image, thereby improving the fun of interacting with the user.
[0109] In some optional implementations, the display device may also integrate a voice service application and a desktop application for playing the currently playing media asset. The desktop application for playing the currently playing media asset may be an application that is enabled by default when the display device is opened, and can be used to support the display device in playing media asset data. In some embodiments, the desktop application is a pre-built application that can serve as a signal source for the display device. The voice service application is an application in the display device that provides voice interaction functionality. Optionally, in some embodiments, the voice service application can provide two voice processing functions: one is to directly output the text converted from the user's interactive voice, and the other is to output the corresponding response voice. Specifically, the voice service application can interact with the desktop application to obtain current scene information, user interactive voice, etc., and then process the user interactive voice using either of the aforementioned voice processing functions based on the current scene information.
[0110] Based on this, before receiving user-interactive voice input, the desktop application can send the user's wake-up voice and current scene information to the voice service application; wherein, the user's wake-up voice is used to wake up the voice service application. As one possible implementation, the user's wake-up voice can be preset voice data, for example, it can be "Xiao Ju Xiao Ju". When the user inputs "Xiao Ju Xiao Ju" via voice, the desktop application can send the user's input "Xiao Ju Xiao Ju" to the voice service application, thus waking up the voice service application.
[0111] The current scene information is used to characterize the current scene of the desktop application. Optionally, the desktop application can determine the current scene information based on whether the digital human interaction control has been triggered. For example, if the digital human interaction control is triggered, the generated current scene information can be a first identifier; if the digital human interaction control is not triggered, the generated current scene information can be a second identifier; the first identifier corresponds to a role-based interaction scene, and the second identifier corresponds to a non-role-based interaction scene.
[0112] After the desktop application sends the current scene information to the voice service application, the voice service application can determine whether the current scene information is a role-based interaction scene or a non-role-based interaction scene based on the identifier corresponding to the current scene information and the correspondence between the identifier and the interaction scene. For example, if the identifier corresponding to the current scene information is the first identifier, then the current scene information is determined to be a role-based interaction scene; if the identifier corresponding to the current scene information is the second identifier, then the current scene information is determined to be a non-role-based interaction scene.
[0113] In some embodiments, referring to FIG10, FIG10 is a flowchart illustrating a method for processing user interactive voice in a voice service application according to some embodiments of the present application, specifically including but not limited to the following steps:
[0114] S1001, the desktop application sends user interaction voice to the voice service application.
[0115] S1002, after the voice service application recognizes that the current scene information is a role interaction scene, it converts the user's interactive voice into user interactive text.
[0116] After the voice service application is activated, the desktop application receives the user's input voice interaction and can send the voice interaction back to the voice service application. The voice service application then first determines whether it is a role-based interaction scenario based on the current scene information. For example, if the voice service application recognizes that the current scene information includes information about the digital human's interactive controls being triggered, it determines that the current scene information is a role-based interaction scenario. That is, after recognizing the current scene information as a role-based interaction scenario, the voice service application converts the user's voice interaction into user interaction text and sends the user interaction text to the desktop application. The desktop application then sends a first query request to the server, carrying the user interaction text and the digital human's identifier, so that the server can respond with a first voice response based on the first query request. After receiving the first voice response from the server based on the first query request, the desktop application can invoke a multimedia player to play the first voice response.
[0117] For example, a voice service application converts the received user interaction voice into text, such as "How's the weather today?", and sends it to a desktop application. The desktop application then forwards it to a server. At this point, the server invokes the relevant digital human to generate a first response voice based on the user interaction text, such as "The weather is sunny today," and sends it to the desktop application. After receiving the first response voice, the desktop application can invoke a multimedia player to play the "The weather is sunny today" response voice, and the timbre in the first response voice is the timbre of the character name corresponding to the first digital human.
[0118] S1003, after the voice service application recognizes that the current scene information is a non-role interaction scene, it obtains the third response voice corresponding to the user's interaction voice and plays the third response voice.
[0119] After the desktop application sends user interaction voice to the voice service application, if the voice service application identifies the current scene as a non-role interaction scene—for example, if the voice service application identifies that the current scene includes information that the digital human interaction controls have not been triggered—then the current scene is determined to be a non-role interaction scene. That is, after the voice service application identifies the current scene as a non-role interaction scene, it can choose not to convert the user interaction voice into user interaction text, but instead directly send the user interaction voice to the server. The server then generates a third-party response voice based on the user interaction voice and sends the third-party response voice to the voice service application. The voice service application obtains the third-party response voice corresponding to the user interaction voice and plays it.
[0120] Alternatively, if the voice service application identifies the current scene as a non-role interaction scenario, it can convert the user's voice interaction into text, then send the corresponding text to the server. The server then generates a third-party response voice based on the text and sends it to the voice service application. The voice service application then retrieves and plays the third-party response voice.
[0121] For example, when a voice service application receives a user's interactive voice message, such as "How's the weather today?", it can send the user's interactive voice message to the server. The server then generates a third-party response voice message based on the user's interactive voice message, such as "The weather is sunny today." After that, the server can send the third-party response voice message to the voice service application, and the voice service application will play the third-party response voice message "The weather is sunny today".
[0122] In the above embodiments, the display device integrates a voice service application and a desktop application that plays the currently playing media. The desktop application can send the user's wake-up voice and current scene information to the voice service application, so that the voice service application can determine whether it is a role-based interaction scene based on the current scene information. This facilitates determining the method of playing the response voice corresponding to the user's interaction voice. On the one hand, this broadens the functionality of the voice service; on the other hand, it distinguishes between role-based and non-role-based interaction scenes, making the interaction method with the user more in line with the user's interaction needs, thereby improving the user interaction experience.
[0123] In some optional implementations, when the display device integrates a desktop application for playing the currently playing media, the desktop application can invoke a multimedia player to play the first response voice message. This achieves the goal of using the voice of the character corresponding to the first digital human to play the first response voice message. In this embodiment, using a desktop application to invoke a multimedia player to play the first response voice message enhances the user interaction experience and reduces development costs. Furthermore, multimedia players typically support multiple audio formats, reducing playback issues caused by format incompatibility and improving the stability and reliability of the first response voice message playback.
[0124] In some optional implementations, after a user wakes up the voice service application, the application can listen to the user's voice input in real time so as to provide timely feedback and response voice to improve the real-time performance of human-computer interaction.
[0125] Referring to Figure 11, which is a flowchart illustrating a return-to-work notification method according to some embodiments of this application, the method includes, but is not limited to, the following steps:
[0126] S1101, after the voice service application detects that the time without voice input exceeds a preset duration, it sends a return prompt message to the desktop application.
[0127] In some embodiments, if the voice service application does not detect user interaction voice for a period of time, i.e., after the voice service application detects no voice input for more than a preset duration, it can send a return prompt message to the desktop application. The preset duration can be set according to actual conditions; for example, it can be set to 10 minutes. No specific limit is made to the preset duration here. The return prompt message is an instruction message used to instruct the user to watch the currently playing media or interact with the digital human.
[0128] S1102, the desktop application obtains the prompt voice of the third digital human from the server based on the encore prompt information, and calls the multimedia player to play the prompt voice.
[0129] After receiving the encore notification, the desktop application can retrieve the prompt voice from the server based on the notification and play it using the multimedia player. The third digital human can be the currently selected digital human, any candidate digital human in the currently playing media assets, or the first digital human in the list. No specific limitations are imposed on the third digital human; it can be set according to the actual situation.
[0130] In some embodiments, the prompt voice may be a voice prompting the user to continue interacting with the digital human, such as "Dear user, please continue interacting with me"; it may also be a voice prompting the user to continue watching the subsequently played media assets, such as "Don't go away, the content is even more exciting"; or it may be a voice corresponding to an overview of the exciting content of the subsequently played media assets, such as "Next, Qiqi is about to embark on an adventure, let's go on an adventure with Qiqi." The above prompt voices are for illustrative purposes only and are not intended to limit the content of the prompt voices.
[0131] In the above embodiments, after the voice service application detects that the time without voice input exceeds a preset duration, it sends a return prompt message to the desktop application. Based on the return prompt message, the desktop application obtains the prompt voice from the third digital human from the server and calls the multimedia player to play the prompt voice, thereby attracting users to continue watching media materials or continuing to interact with the digital human, thus increasing users' enthusiasm for watching media materials and interacting with the digital human.
[0132] In some alternative implementations, during the process of the desktop application obtaining the prompt voice of the third digital human from the server based on the encore prompt information, the prompt voice of the third digital human can be generated based on the media asset identifier of the currently playing media asset, the image of the currently playing frame, and the third digital human.
[0133] Referring to Figure 12, Figure 12 is a schematic flowchart of acquiring and playing the prompting voice of a third digital human according to some embodiments of this application, specifically including but not limited to the following steps:
[0134] S1201, the desktop application sends a prompt request to the server based on the encore notification information.
[0135] In some embodiments, after the voice service application detects that the time without voice input exceeds a preset duration, it sends a replay prompt message to the desktop application. Upon receiving the replay prompt message, the desktop application sends a prompt message retrieval request to the server based on the message. This request may include the media asset identifier and the current playback frame image of the currently playing media asset. The server can then use this information to obtain the episode and plot details of the currently playing media asset, as well as the episode and plot details of the upcoming media asset. This allows the server to generate an engaging prompt message based on the episode and plot details of both the currently playing and upcoming media assets.
[0136] In some possible embodiments, the server can extract key elements from the currently playing frame image, including but not limited to themes, plots, and core character emotions. Then, the server can obtain prompt text and prompt audio related to the plot and character emotions based on the media asset identifier and the identified key image elements.
[0137] S1202, the desktop application obtains the prompt voice of the third digital human from the server.
[0138] In some embodiments, the prompt voice can be determined by the server based on the digital human identifier of the third digital human, and by invoking the third digital human according to the media asset identifier and the currently playing frame image. For example, after receiving a prompt retrieval request, the server can determine the episode and plot of the currently playing media asset, as well as the episode and plot of the media asset to be played next, based on the media asset identifier and the currently playing frame image carried in the prompt retrieval request. Based on the episode and plot of the currently playing media asset, as well as the episode and plot of the media asset to be played next, the server can determine voice content that can attract the user to continue watching the media asset or continue interacting with the digital human; and the server can convert the timbre in the voice content into the timbre of the third digital human.
[0139] The third digital person can be selected from the candidate digital persons. For example, the third digital person can be the currently selected digital person among the candidate digital persons, or a digital person specified by the user; it can also be one of the digital persons included in the currently playing frame image, or any one of the digital persons included in the episodes after the currently playing frame image.
[0140] S1203, when the desktop application plays a prompt voice when the third digital human is the same as the candidate digital human corresponding to the digital human card where the screen focus is located.
[0141] In some embodiments, before playing the prompt voice, it can be first determined whether the third digital human is the same as the candidate digital human corresponding to the digital human card where the screen focus is located. If they are the same, the desktop application can call the multimedia player to play the prompt voice.
[0142] S1204, when the desktop application is different from the candidate digital human corresponding to the digital human card where the screen focus is located, it focuses the screen focus on the digital human card corresponding to the third digital human and plays the prompt voice.
[0143] In some embodiments, if the third digital human is different from the candidate digital human corresponding to the digital human card where the screen focus is located, the desktop application needs to first focus the screen focus on the digital human card corresponding to the third digital human and then call the multimedia player to play a prompt voice. This allows the prompt voice to be played using the voice of the character name corresponding to the third digital human, making the played voice more consistent with the image of the digital human.
[0144] In the above embodiments, a prompt voice is generated based on the digital human identifier of the third digital human, the media asset identifier of the currently played media asset, and the currently played frame image. On the one hand, this makes the generated prompt voice more consistent with the current plot of the currently played media asset and the plot to be played next, which can increase the user's enthusiasm to continue watching the media asset. On the other hand, by using the voice of the character name corresponding to the third digital human to play the prompt voice, not only is it possible to switch the digital human according to the screen focus, but it also makes the played voice more consistent with the image of the switched digital human.
[0145] In some optional implementations, the server in the above embodiments can be implemented by a single server or by a server cluster. If implemented by a server cluster, the server cluster can consist of a desktop application, a voice service application, a terminal-oriented subsystem, a content subsystem, an image recognition service, and a digital human model. The terminal-oriented subsystem provides a bridge for interaction between the desktop application and the content subsystem, image recognition service, and digital human model; the content subsystem stores the digital human's introductory text and audio data; the image recognition service provides image recognition and other functions; and the digital human model provides digital human-related services. The digital human model can be pre-trained based on sample media data, sample audio data, and sample text data.
[0146] Referring to Figure 13, which is a timing diagram of a training digital human model according to some embodiments of this application, specifically including but not limited to the following steps:
[0147] S1301, the digital human model sends the sample audio and video data of the character corresponding to the digital human and the script of the TV series to the big model service.
[0148] S1302, the large model service performs voice replication and script knowledge learning for digital humans, generates trained digital humans, and assigns a unique digital human identifier to each digital human.
[0149] S1303, The large model service sends the digital human and digital human identifier to the digital human model.
[0150] S1304, Operational system performs digital human maintenance.
[0151] The digital human maintenance process includes setting a prompt for each digital human, associating each digital human with its corresponding digital human identifier, and setting a media asset identifier. The digital human identifier can be a digital human ID, and the media asset identifier can be a media asset IP.
[0152] S1305, the operating system sends the digital human and its corresponding digital human identifier, the introductory text for each digital human, and the media asset identifier to the content subsystem.
[0153] S1306, The content subsystem associates digital humans with their corresponding digital human identifiers, the introductory text for each digital human, and the media asset identifier.
[0154] S1307, the content subsystem sends the introductory text for each digital human and the corresponding digital human identifier to the digital human model.
[0155] S1308, the digital human model generates interactive guidance text corresponding to each digital human identifier based on each digital human identifier and the corresponding guidance text.
[0156] S1309, the digital human model feeds back each digital human identifier and corresponding interactive prompts to the content subsystem.
[0157] S1310, the content subsystem stores each digital human identifier and its corresponding interactive prompts.
[0158] S1311, Third-party content sends a notification of episode update to the content subsystem.
[0159] S1312, The content subsystem sends a notification to a third-party content provider that it has received a notification of a new episode.
[0160] S1313, The content subsystem sends a notification of episode updates to the large model service.
[0161] S1314, the large model service sends a request to the content subsystem to obtain newly added script content.
[0162] S1315, The content subsystem feeds back the newly added script content to the large model service.
[0163] S1316, the large model service updates the digital human based on the newly added script content.
[0164] S1317, the large model service sends the updated digital human to the digital human model.
[0165] In the above embodiments, digital humans corresponding to the characters in the media asset data can be trained based on the media asset data, so that users can interact with the digital humans corresponding to the media asset data when watching the media asset data, thereby improving the intelligence and fun of human-computer interaction.
[0166] In some optional implementations, see Figure 14, which is a timing diagram of an interaction provided according to some embodiments of this application, specifically including but not limited to the following steps:
[0167] S1401, The user triggers a request to retrieve the details page of the media asset associated with the currently playing media asset.
[0168] S1402, after receiving the request to obtain the media asset details page associated with the currently playing media asset, the desktop application sends the request carrying the media asset identifier to the terminal subsystem.
[0169] S1403, the terminal subsystem sends a digital human query request carrying media asset identifier to the content subsystem.
[0170] S1404, The content subsystem feeds back a list of digital humans that are identical to the media asset identifier to the terminal-facing subsystem.
[0171] S1405, the content subsystem feeds back the relevant data from the media asset details page and the list of digital humans to the desktop application.
[0172] S1406, the desktop application displays the media asset details page based on relevant data and a list of digital humans.
[0173] S1407, The user triggers a click command on the digital human interactive control in the media asset details page.
[0174] S1408, the desktop application responds to the digital human interaction control's triggered operation and displays the digital human card of the candidate digital human associated with the currently playing media asset.
[0175] S1409, the desktop application sends a prompt retrieval request to the terminal-facing subsystem based on the digital human identifier corresponding to the digital human card with the focus on the page.
[0176] S1410, the terminal subsystem sends a request to the content subsystem to obtain the digital human identifier corresponding to the digital human card where the page is focused and the corresponding prompt text.
[0177] S1411, the content subsystem feeds back the corresponding digital human interactive guidance text to the terminal-facing subsystem.
[0178] S1412, the terminal subsystem feeds back the corresponding digital human's interactive guidance to the desktop application.
[0179] S1413, the desktop application uses a multimedia player to play interactive prompts.
[0180] S1414, The user inputs a wake-up voice command into the desktop application.
[0181] S1415, the desktop application sends a wake-up request to the voice service application, carrying the user's wake-up voice and current scene information.
[0182] S1416, Voice service application launched.
[0183] S1417, The user inputs interactive voice input for the target digital human.
[0184] The target digital human can be either the currently selected digital human or the digital human represented by the user's interactive voice containing digital human-related information.
[0185] S1418, the voice service application recognizes the current scene information, and when the current scene information is recognized as a role interaction scene, it converts the user's input voice interaction into user interaction text.
[0186] S1419, The voice service application sends user interaction text to the desktop application.
[0187] S1420, the desktop application sends user interaction text and the identifier of the target digital human to the digital human model.
[0188] S1421, the digital human model calls the target digital human based on the target digital human's digital human identifier, so that the target digital human can respond with voice and text output based on the text.
[0189] S1422, the digital human model will send voice and text responses to the desktop application.
[0190] S1423, The desktop application displays the reply text and plays the reply audio using a player.
[0191] S1424, The voice service application continuously listens to the user's voice.
[0192] S1425, after the voice service application detects that the user has not input voice for more than a preset time, it sends a return prompt message to the desktop application.
[0193] S1426, the desktop application sends the digital human identifier, the media asset identifier of the currently playing media asset, and the currently playing frame image to the terminal subsystem.
[0194] S1427, the terminal subsystem sends a request to the image recognition service to obtain key elements in the currently playing frame image.
[0195] S1428, The image recognition service feeds back the key elements in the currently playing frame image to the terminal-oriented subsystem.
[0196] S1429, The terminal subsystem sends a prompt request to the digital human model, which includes the digital human identifier, the episode identifier, and the identified key image elements.
[0197] S1430, the digital human model will provide voice prompts and corresponding text feedback to the desktop application.
[0198] S1431, the desktop application displays the text of the prompt voice and uses a player to play the acquired prompt voice.
[0199] In some embodiments of this application, as shown in FIG15, the display device may further perform the following steps to enrich the interactive methods in media asset playback scenarios:
[0200] S1501: When playing the currently playing media asset and receiving an interactive command, take a screenshot of the currently playing media asset to obtain the media asset image.
[0201] Interactive commands can be sent to the display device via a control device such as a remote control, or a smart device such as a mobile terminal. For example, a remote control may have interactive buttons. When the remote control detects a trigger operation on an interactive button, such as pressing an interactive button, it can send an interactive command to the display device. Interactive buttons could be, for example, Artificial Intelligence (AI) buttons. Interactive commands can also be voice commands.
[0202] Specifically, the display device includes a screenshot application. In response to an interactive command, the display device launches the screenshot application to capture a screenshot of the currently playing media asset. This can be a screenshot of the entire screen, resulting in a media asset image containing the content of the currently playing media asset; or it can be a screenshot of the playback interface of the currently playing media asset, resulting in a media asset image. The size of the playback interface can be smaller than or equal to the screen size.
[0203] S1502, obtain the identified role identifier obtained by performing role recognition on the media asset image; and query the interactive role identifier associated with the media asset identifier of the currently playing media asset and the role interaction information corresponding to the interactive role identifier.
[0204] The currently playing media asset contains multiple characters, meaning at least two. For example, if the currently playing media asset is an animated film, it contains multiple cartoon characters; similarly, if the currently playing media asset is a movie, it contains multiple characters. Characters in the currently playing media asset can be interactive or regular characters. Interactive characters are characters given the ability to interact with the user. Regular characters, on the other hand, do not have this interactive function. For example, a character tag can be used to distinguish between interactive and regular characters. If the character tag value is the first tag value, the character is an interactive character; if the character tag value is the second tag value, the character is a regular character. Each media asset identifier can be associated with at least one interactive character identifier. There is a one-to-one correspondence between interactive character identifiers and interactive characters.
[0205] In some embodiments, after obtaining the media asset image, the display device can send a role recognition request to the server, the request carrying the media asset image. In response to the role recognition request, the server performs image segmentation and recognition on the media asset image, obtaining an image recognition result. Specifically, if the character in the media asset image is a virtual avatar, the image recognition result includes the character's identifier (i.e., identified role identifier); if the character in the media asset image is played by a real person, the image recognition result includes the corresponding person's identifier, which uniquely identifies a real person. The role recognition request may also carry a media asset identifier. Based on the person's identifier in the image recognition result, the server can query the character's identifier (i.e., identify role identifier) based on the media asset identifier and the person's identifier, thus converting the person's identifier to a role identifier. Then, the server can return the character's identifier (i.e., identified role identifier) from the media image to the display device.
[0206] In some embodiments, the display device can obtain the media asset identifier of the currently playing media asset, and then query the server based on the media asset identifier to obtain the associated interactive character identifier and the character interaction information corresponding to the interactive character identifier.
[0207] S1503, Match the identified role identifier with the interactive role identifier associated with the media asset identifier of the currently playing media asset to determine the first interactive role identifier.
[0208] The display device may use an interactive role identifier that is consistent with the identified role identifier as the first interactive role identifier.
[0209] Specifically, the display device uses the interactive role identifier that matches any of the identification role identifiers among the interactive role identifiers associated with the media asset identifier as a candidate interactive role identifier. Thus, if there are at least two identification role identifiers and the media asset identifier is associated with at least two interactive role identifiers, for example, A, B, and C are three role identifiers, and A, B, and C are three interactive role identifiers associated with the media asset identifier, and A and B are two identification role identifiers, then there may be more than one candidate interactive role identifier.
[0210] In some embodiments, if only one candidate interactive character identifier exists, that candidate interactive character identifier can be used as the first interactive character identifier; if more than one candidate interactive character identifier exists, one candidate interactive character identifier can be selected as the first interactive character identifier. For example, the selection can be based on the weight corresponding to the candidate interactive character identifier, with the one with the highest weight being selected as the first interactive character identifier. The weight corresponding to the candidate interactive character identifier can be determined based on at least one of the following: the popularity of the character represented by the candidate interactive character identifier, or the position of the character represented by the candidate interactive character identifier in the media asset image. For example, the higher the popularity, the greater the weight; the closer the character's position is to the center of the media asset image, the greater the weight.
[0211] S1504, the control display shows the first role option corresponding to the first interactive role identifier and the second role option corresponding to the second interactive role identifier in the interactive card, and sets the default focus on the first role option. The second interactive role identifier is the interactive role identifier other than the first interactive role identifier among the interactive role identifiers associated with the media asset identifier. The first role option is generated based on the role interaction information corresponding to the first interactive role identifier, and the second role option is generated based on the role interaction information corresponding to the second interactive role identifier. The first role option is configured to enter a dialogue with the first interactive role after being selected.
[0212] The character interaction information may include at least one of the following: character identifier, character name, character image, character feature description, and background image. When the character is played by a real person, the character name refers to the character's name within the currently playing media, not the character's real-life name. When the character is a virtual avatar, such as a cartoon character, the character name is the name of the virtual avatar. Character feature description information uses text to describe the character's characteristics. The character image is an image used to depict the character's appearance; it may, but is not, a portrait or full-body photo of the character. The background image may be an image depicting the interactive character.
[0213] In some embodiments, the display device can generate a background image for an interactive card based on the background image of the first interactive character in the character interaction information corresponding to the first interactive character identifier, and generate a first character option based on the character image pattern of the first interactive character, wherein the first character option displays the character image pattern of the first interactive character. The display device can also generate a second character option based on the character image pattern of the second interactive character represented by the second interactive character identifier, wherein the second character option displays the character image pattern of the second interactive character.
[0214] In some embodiments, the display device can obtain guided dialogue content based on a first interactive character identifier. The guided dialogue content is preset, or generated based on plot information of the current storyline in the currently playing media and the character setting information of the first interactive character. The interactive card may also contain this guided dialogue content. The display device can send the first interactive character identifier and plot information to a local or remote intelligent agent. The intelligent agent can obtain the character setting information corresponding to the first interactive character identifier, generate guided dialogue content based on the character setting information and plot information, and return the guided dialogue content to the display device. Thus, the display device can generate an interactive card based on the character interaction information corresponding to the interactive character identifier and the guided dialogue content.
[0215] In some embodiments, the display device displays a first dialogue prompt message in response to a dialogue initiation operation triggered by an interactive card. The first dialogue prompt message prompts the user to initiate a dialogue with a first interactive role represented by a first interactive role identifier. The interactive card may also display a dialogue initiation entry point, and the dialogue initiation operation may be a triggering operation targeting this entry point, such as a click operation.
[0216] Figure 16 illustrates an interactive card. Interactive card 1601 includes guiding dialogue content 1602, three role options (role option 1603, role option 1604, and role option 1605), and a dialogue initiation entry point 1606. Role option 1604 is the option for the first interactive role, while role options 1603 and 1605 are options for two different second interactive roles. The default focus is set on role option 1604, and the image in the role option is the character's image. The background of interactive card 1601 is generated using the background image from the role interaction information corresponding to the first interactive role's identifier. The background image can present the image of the interactive role; for example, image 1608 is the image of the first interactive role presented in the background image.
[0217] In this embodiment, when playing the currently playing media asset and receiving an interaction command, a screenshot of the currently playing media asset is taken to obtain a media asset image. A role identifier obtained by performing role recognition on the media asset image is acquired. The interaction role identifier associated with the media asset identifier of the currently playing media asset and the corresponding role interaction information are queried. The identified role identifier is matched with the interaction role identifier associated with the media asset identifier of the currently playing media asset to determine the first interaction role identifier. Since when a user triggers an interaction command, they usually see a character they like and want to interact with that character, it is easy to locate the character the user wants to interact with by taking a screenshot and identifying the interaction role identifier. Furthermore, the control... The display shows the first role option corresponding to the first interactive role identifier and the second role option corresponding to the second interactive role identifier in the interactive card, and sets the default focus on the first role option. The first role option is configured to enter a dialogue with the first interactive role after being selected. Thus, the first role option is selected by default and the second role option is provided. It can automatically select a role that the user is likely to want to interact with, and also provides the opportunity to select other roles. This makes it more convenient for the user to select the role to interact with, and helps to efficiently determine the role to interact with. It realizes a way for users to interact with the roles in the media asset playback scenario, and enriches the interaction methods in the media asset playback scenario.
[0218] In some embodiments, when querying the interactive role identifier associated with the media asset identifier of the currently playing media asset and the corresponding role interaction information, the display device can specifically obtain the media asset information of the currently playing media asset. The media asset information includes the media asset identifier and first indication information. If the first indication information indicates that the currently playing media asset supports role interaction, the device queries the interactive role identifier associated with the media asset identifier of the currently playing media asset and the corresponding role interaction information. The first indication information indicates whether the currently playing media asset supports role interaction. If the first indication information indicates that the currently playing media asset supports role interaction, it means that the media asset identifier of the currently playing media asset has an associated interactive role identifier.
[0219] In this embodiment, on the one hand, by recording the first indication information in the media asset information to characterize whether the currently playing media asset supports character interaction, it is convenient and quick to determine whether character interaction is supported. On the other hand, the step of querying the interactive character identifier and character interaction information only when character interaction is supported can avoid performing invalid queries and help save resources.
[0220] In some embodiments, when querying the interactive role identifier associated with the media asset identifier of the currently playing media asset and the corresponding role interaction information, the display device can specifically query the second indication information corresponding to the identified role identifier; if the second indication information indicates that the role represented by the identified role identifier supports role interaction, the device queries the interactive role identifier associated with the media asset identifier of the currently playing media asset and the corresponding role interaction information. The second indication information indicates whether the role represented by the identified role identifier supports role interaction.
[0221] In this embodiment, the interactive character identifier and interaction information are only queried when the character located from the media asset image supports character interaction. This avoids displaying interactive cards that are irrelevant to the character the user expects to interact with, thus improving the user experience. Furthermore, since not all media assets contain the first indication information, this embodiment can be used to determine whether to enter the interaction process for media assets that do not have the first indication information, broadening the application scenarios.
[0222] In some embodiments, when setting the default focus on the first role option, the display device may specifically set the label value of the focus label corresponding to the first role option to a first label value, thereby setting the default focus on the first role option; wherein, when the label value of the focus label corresponding to the first role option changes to a second label value, the focus moves away from the first role option. In response to selecting the second role option, the display device changes the label value of the focus label corresponding to the first role option to the second label value, and changes the label value of the focus label corresponding to the selected second role option to the first label value, thereby shifting the focus from the first role option to the second role option.
[0223] In this embodiment, the focus can be quickly shifted by updating the focus label value.
[0224] In some embodiments, the display device may also, in response to a selection operation for a second role option in an interactive card, move the focus from the first role option to the second role option; and, in response to a dialogue initiation operation triggered by the interactive card, control the display to display a second dialogue prompt message, wherein the second dialogue prompt message is used to prompt a dialogue with the second interactive role represented by the second role option.
[0225] Specifically, when the second role option is selected, the display device can update the interactive card. For example, the background image in the interactive card can be replaced with the background image in the role interaction information corresponding to the second interactive role identifier. The guided dialogue content can also be updated to match the identity of the second interactive role.
[0226] In this embodiment, the second role option allows for free switching between roles, improving interaction efficiency.
[0227] In some embodiments, when the control display shows the first role option corresponding to the first interactive role identifier and the second role option corresponding to the second interactive role identifier in the interactive card, the display device can specifically obtain the role introduction information associated with the first interactive role identifier; control the display to show the first role option corresponding to the first interactive role identifier and the second role option corresponding to the second interactive role identifier in the interactive card, and display the role introduction information in the role introduction card. The role introduction information may include the identity information, character image, age, or appearance information of the first interactive role in the currently played media asset, and may also include information about related works. Related works refer to works that include the first interactive role, for example, works related to the currently played media asset. Work information includes, but is not limited to, the cover image of the work. The role introduction information may also include information used to introduce the currently played media asset.
[0228] The character introduction card can display the character's identity information, image, age, or appearance, as well as information about the work being played. This work information can be used to trigger access to the playback interface of the corresponding work.
[0229] Specifically, once the first interactive role identifier is determined, the display device can send a request to the server to obtain the introduction information carrying the first interactive role identifier. In response to the request, the server queries the role introduction information associated with the first interactive role identifier and returns the queried role introduction information to the display device.
[0230] In some embodiments, when playing currently playing media and receiving an interactive instruction, the display device pauses the playback of the currently playing media and displays an interactive card and a character introduction card. As shown in Figure 17, the interactive card and character introduction card 1701 are displayed.
[0231] In some embodiments, when the second character option in the interactive card is selected, the display device obtains the character introduction information of the second interactive character and updates the information in the character introduction card to the information in the character introduction information of the second interactive character. In this embodiment, the character introduction card is used to present character introduction information, thereby enabling users to further understand the currently playing media asset or the first interactive character based on the character introduction card, which helps in subsequent dialogue with the first interactive character.
[0232] In some embodiments, the display device may also respond to a dialogue initiation operation triggered by an interactive card, controlling the display to show first dialogue prompt information. The first dialogue prompt information prompts the user to engage in dialogue with the first interactive character represented by the first interactive character identifier. It also receives first dialogue content input by the user, controls the display to show the first dialogue content, and outputs second dialogue content as a response from the first interactive character to the first dialogue content. The second dialogue content is generated based on the first interactive character's character setting information and the first dialogue content. The first dialogue prompt information may include opening dialogue content provided by the first interactive character to the user. This opening dialogue content is a greeting to the user and may be consistent with or different from the introductory dialogue content. The opening dialogue content may be preset or generated based on the plot information of the current storyline in the currently playing media and the first interactive character's character setting information. It may also be automatically generated using an AI model.
[0233] The initial dialogue prompt may include preset questions or questions generated based on the current storyline. The prompt may also include a voice input icon, which can be an animated or static image. The voice input icon prompts the user to use voice to interact with the character.
[0234] The first dialogue can be input via remote control or mobile device, and can be either voice or text. The second dialogue can include at least one of a text response and a voice response. Inputting the second dialogue can be either a text response or playing a voice response. Character setting information describes the background, personality, behavioral patterns, or values of the interactive character in the currently playing media asset.
[0235] In some embodiments, upon detecting a dialogue initiation action, the display device hides the interactive card and displays a first dialogue prompt. For example, the display device can display opening dialogue content in text form, and the display device can also play the opening dialogue content using voice.
[0236] In some embodiments, displaying the first dialogue content can be done directly or indirectly. When the first dialogue content is text, the display device can directly display it. When the first dialogue content is a voice command, the display device can convert the first dialogue content into text, or it can send the first dialogue content to another device for conversion into text, and then display the converted text. This achieves indirect display of the first dialogue content.
[0237] In some embodiments, after receiving the first dialogue content, the display device can generate a prompt message based on the first interactive role identifier and the first dialogue content. The prompt message includes the first interactive role identifier and the first dialogue content. Then, the prompt message can be sent to a local or remote intelligent agent. The intelligent agent can query the role setting information associated with the first interactive role identifier, which is the role setting information of the first interactive role. Then, it generates a reply text based on the role setting information of the first interactive role and the first dialogue content. The intelligent agent can also query the voice features associated with the first interactive role identifier, which are the characteristics of the first interactive role's voice in the currently playing media asset, and can convert the second dialogue content into a reply voice based on the voice features of the first interactive role. The second dialogue content may include at least one of reply text and reply voice. The role setting information and voice features of the interactive role are pre-generated and stored.
[0238] In some embodiments, the prompt message may also include a user identifier. The intelligent agent may obtain user profile information associated with the user identifier and generate a reply text based on the role setting information of the first interactive role, the content of the first dialogue, and the user profile information.
[0239] In this embodiment, the display device can respond to a dialogue initiation operation triggered by an interactive card, displaying a first dialogue prompt message. The first dialogue prompt message prompts the user to engage in dialogue with the first interactive role represented by the first interactive role identifier. Thus, through the interactive card and the first dialogue prompt message, the user is guided to interact with the role. The device receives the first dialogue content input by the user, displays the first dialogue content, and outputs the second dialogue content in response to the first interactive role's first dialogue content. Since the second dialogue content is generated based on the role setting information of the first interactive role and the first dialogue content, the second dialogue content is the content of the first interactive role's dialogue with the user in the currently playing media asset. This realizes a way for the user to interact with the role in the media asset in the media asset playback scenario, enriching the interaction methods in the media asset playback scenario.
[0240] In some embodiments, when executing a dialogue initiation operation triggered by an interactive card and controlling the display to show a first dialogue prompt message, the display device may specifically respond to the dialogue initiation operation triggered by the interactive card by generating a first tab based on the character interaction information corresponding to a first interactive character identifier and generating a second tab based on the character interaction information corresponding to a second interactive character identifier; controlling the display to show the first dialogue prompt message, the first tab, and the second tab, and setting the focus on the first tab, while the second tab is used to switch to dialogue with the second interactive character represented by the second interactive character identifier. The first tab may contain at least one of the following: character name, character image pattern, or character feature description information from the character interaction information corresponding to the first interactive character identifier. The second tab may contain at least one of the following: character name, character image pattern, or character feature description information from the character interaction information corresponding to the second interactive character identifier.
[0241] Figure 18 shows a schematic diagram of the interface in dialogue mode. The content within the dashed box 1800 represents the content displayed in dialogue mode, while the remaining content shows the screen of the currently playing media asset. Three tabs are displayed: tab 1801 is the first tab, and tabs 1802 and 1803 are the second tabs. Tab 1801 is selected, while tabs 1802 and 1803 are unselected. The content within the dashed box 1804 represents the first dialogue prompt information, with 1805 being a voice input icon. The prompts "How old are you?" and "Can you camp underwater?" can be preset questions associated with the first interactive character, or they can be generated based on the character's settings or the current storyline's plot information.
[0242] In this embodiment, after entering the dialogue mode, a first tab and a second tab are displayed for the user to select. Thus, in the dialogue mode, the user can freely switch between interactive roles in the dialogue through the tabs, thereby improving the interaction efficiency.
[0243] In some embodiments, the first dialogue prompt information includes the opening dialogue content of the first interactive character. The opening dialogue content may be preset, or it may be generated based on the plot information of the current storyline in the currently playing media and the character setting information of the first interactive character. When the control display shows the first dialogue prompt information, the display device may specifically control the display to show the opening dialogue content and play the opening dialogue content by voice. The opening dialogue content is text and is used to greet the user. As shown in Figure 19, the content in the dashed box 1900 is the content displayed in dialogue mode, and the panel 1901 displays the opening dialogue content and the corresponding voice identifier.
[0244] In some embodiments, the display device may use the introductory dialogue content as the opening dialogue content, or acquire the opening dialogue content in real time.
[0245] In this embodiment, the opening dialogue can be used to greet the user and further guide the user to participate in the conversation.
[0246] In some embodiments, the display device can be a television, which contains a television system, an interactive application, and a homepage application. The television system is the entire television system, supporting basic functions such as receiving user button presses, remote control on / off states, and taking screenshots. The interactive application supports character interaction; it may include a screenshot application or be able to call a screenshot application to take screenshots. Character interaction refers to the user interacting with characters in the viewed media assets, enabling open-ended voice chat. The homepage application provides playback functionality. An image recognition system and an interactive system are deployed on the server. The image recognition system performs image segmentation and recognition, segmenting images into human figures and comparing their similarity to character information in the media asset library, providing the recognition result. The interactive system is a cloud-based system providing an interactive character library, built based on characters from multiple media assets, to support dialogue with users through simulated characters. As shown in Figure 20, a timing diagram corresponding to an interactive method is provided, specifically including:
[0247] S2001, the remote control sends button signals to the TV system in the TV.
[0248] In this embodiment of the application, when the remote control detects that the user has pressed the interactive button, it sends the button signal, i.e., the interactive command, to the television system in the television.
[0249] S2002, the television system receives button signals.
[0250] S2003: When the TV system recognizes that the signal is triggered by pressing the interactive button, it sends an application launch message to the interactive application.
[0251] S2004: After receiving the application launch message, the interactive application launches and captures the current screen to obtain the media asset image.
[0252] Interactive applications can use the screenshot methods provided by the TV system or other applications to capture media asset images.
[0253] S2005, the interactive application calls the TV's home page application to query the media asset information of the currently playing media asset.
[0254] S2006, the home application returns media asset information to the interactive application.
[0255] S2007, the interactive application determines whether the currently playing media asset supports character interaction based on the media asset information.
[0256] S2008, when character interaction is supported, interactive applications call the interactive system to query character interaction information.
[0257] S2009, the interactive system returns character interaction information to the interactive application.
[0258] S2010, the interactive application sends media asset images to the image recognition system.
[0259] Interactive applications can send media asset images to the image recognition system as soon as they are captured, without necessarily waiting until the character interaction information is obtained.
[0260] S2011, the image recognition system segments and identifies media asset images to obtain identification role identifiers.
[0261] S2012, the image recognition system returns the identified role identifier to the interactive application.
[0262] S2013, the interactive application will match the identified role identifier with the interactive role identifier to determine the first interactive role identifier.
[0263] S2014, the interactive application displays interactive cards, and enters dialogue mode when an action is initiated through a dialogue triggered by an interactive card.
[0264] Figure 21 shows a sequence diagram corresponding to another interaction method, specifically including:
[0265] S2101, the remote control sends button signals to the TV system in the TV.
[0266] When the remote control detects a press of the interactive button, it sends the button signal, i.e., the interactive command, to the TV system.
[0267] S2102, the television system receives button signals.
[0268] S2103, when the TV system recognizes that the signal is triggered by the pressing of the interactive button, it sends an application launch message to the interactive application.
[0269] S2104, after receiving the application launch message, the interactive application launches and captures the current screen to obtain the media asset image.
[0270] Interactive applications can use the screenshot methods provided by the TV system or other applications to capture media asset images.
[0271] S2105, the interactive application sends media asset images to the image recognition system.
[0272] S2106, The image recognition system segments and identifies media asset images to obtain identification role identifiers.
[0273] S2107, The image recognition system returns the identified role identifier to the interactive application.
[0274] S2108, the interactive application queries the interactive system for indication information (i.e., second indication information) that identifies the role identifier.
[0275] S2109, The interactive system returns instruction information to the interactive application to identify the role identifier.
[0276] S2110, the interactive application determines whether the instruction information indicates support for role interaction.
[0277] S2111, When character interaction is supported, the interactive application calls the interactive system to query character interaction information.
[0278] S2112, The interactive system returns character interaction information to the interactive application.
[0279] S2113, the interactive application will match the identified role identifier with the interactive role identifier to determine the first interactive role identifier.
[0280] S2114, The interactive application displays an interactive card, and enters dialogue mode when it detects an operation initiated through a dialogue triggered by the interactive card.
[0281] Using the above method, the system automatically identifies human figures based on screenshots and automatically matches interactive characters based on these figures, enabling users to interact with characters in the currently playing film or television work. When applied to daily movie viewing, this provides better interactivity and enriches the interactive methods during playback.
[0282] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0283] Based on the same inventive concept, the device embodiments provided according to some embodiments of this application will be described in detail below with reference to FIG22. It should be understood that the interactive devices in some embodiments of this application can execute various interactive methods of the foregoing embodiments of this application, that is, the specific working processes of the various products below can be referred to the corresponding processes in the foregoing method embodiments.
[0284] Figure 22 is a schematic diagram of an interactive device provided according to some embodiments of this application. It should be understood that the interactive device can execute the interactive methods shown in Figures 1 to 14; the interactive device 2200 includes: a first control unit 2210, used to control the display to show a media asset details page associated with the currently playing media asset; wherein the media asset details page displays digital human interactive controls; a second control unit 2220, used to control the display to show digital human cards of candidate digital humans associated with the currently playing media asset in response to a trigger operation on the digital human interactive controls; wherein each candidate digital human card contains a digital human image corresponding to the candidate digital human and / or the character name in the currently playing media asset represented by the candidate digital human; a receiving unit 2230, used to receive input user interaction voice after the digital human card of the first digital human is selected, wherein the first digital human is one of the candidate digital humans; and a playback unit 2240, used to obtain a first response voice from the first digital human in response to the user interaction voice from the server, and play the first response voice; wherein the timbre in the first response voice is the timbre of the character name corresponding to the first digital human.
[0285] Each unit module of the interactive device can execute the corresponding steps in the above method embodiment, so the details of each unit module will not be elaborated here. Please refer to the description of the corresponding steps above for details.
[0286] Based on the same inventive concept, another device embodiment provided according to some embodiments of this application will be described in detail below with reference to FIG23. It should be understood that the interactive device in some embodiments of this application can execute various interactive methods of the foregoing embodiments of this application, that is, the specific working process of the various products below can be referred to the corresponding process in the foregoing method embodiments.
[0287] Figure 23 is a schematic diagram of an interactive device provided according to some embodiments of this application. It should be understood that this interactive device can execute the interactive methods shown in Figures 15 to 21; the interactive device 2300 includes: a screenshot unit 2310, used to take a screenshot of the screen of the currently playing media asset when playing the media asset and receiving an interactive instruction, to obtain a media asset image; an information acquisition unit 2320, used to acquire an identified role identifier obtained by performing role recognition on the media asset image; and to query the interactive role identifier associated with the media asset identifier of the currently playing media asset and the role interaction information corresponding to the interactive role identifier; and a determination unit 2330, used to match the identified role identifier with the interactive role identifier associated with the media asset identifier of the currently playing media asset to determine a first interactive role identifier. An interactive role identifier; a control unit 2340, configured to control the display to show a first role option corresponding to the first interactive role identifier and a second role option corresponding to the second interactive role identifier in an interactive card, and to set the default focus on the first role option, wherein the second interactive role identifier is an interactive role identifier other than the first interactive role identifier among the interactive role identifiers associated with the media asset identifier, the first role option is generated based on the role interaction information corresponding to the first interactive role identifier, the second role option is generated based on the role interaction information corresponding to the second interactive role identifier, and the first role option is configured to enter a dialogue with the first interactive role after being selected.
[0288] Each unit module of the interactive device can execute the corresponding steps in the above method embodiment. For specific limitations, please refer to the limitations of the interactive method above, which will not be repeated here.
[0289] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0290] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
A display device, comprising: monitor; Memory, configured to store computer programs or instructions; At least one processor, connected to the display and the memory, is configured to execute the computer program or instructions to cause the display device to: The display is controlled to show the media asset details page associated with the currently playing media asset; wherein, the media asset details page displays digital human interactive controls; In response to a trigger operation on the digital human interaction control, the display is controlled to show digital human cards of candidate digital humans associated with the currently playing media asset; wherein, the digital human card of each candidate digital human contains a digital human image corresponding to the candidate digital human and / or the name of the character in the currently playing media asset represented by the candidate digital human; After the first digital human's digital human card is selected, input user interaction voice is received, wherein the first digital human is one of the candidate digital humans; Based on the user's interactive voice, the system obtains the first response voice from the server in response to the user's interactive voice and plays the first response voice; wherein the timbre in the first response voice is the timbre of the character name corresponding to the first digital human. According to claim 1, when the at least one processor executes the step of obtaining the first response voice from the first digital human in response to the user interaction voice from the server and playing the first response voice, it is configured to execute the computer program or instructions to cause the display device to: If the user interaction voice does not contain digital human associated information, the system retrieves the first response voice of the first digital human in response to the user interaction voice from the server and plays the first response voice. In a case where the user interaction voice contains digital person association information, if a second digital person represented by the digital person association information is different from the first digital person, focusing the screen focus on a digital person card corresponding to the second digital person, and according to the user interaction voice, obtaining a second reply voice of the second digital person for the user interaction voice from a server, and playing the second reply voice; wherein, The timbre in the second reply voice is the timbre of the character name corresponding to the second digital persona. The display device according to claim 1 or 2, wherein when the at least one processor executes the step of obtaining the first response voice of the first digital human in response to the user interaction voice from the server, it is configured to execute the computer program or instructions to cause the display device to: Convert the user's interactive voice into user interactive text; Send a first query request to the server, carrying the user interaction text and the digital human identifier of the first digital human; receiving the first reply voice fed back by the server based on the first inquiry request; wherein The first response voice is determined by the server based on the digital human identifier of the first digital human, and invoked by the first digital human according to the user interaction text. The display device according to claim 3, wherein the display device integrates a voice service application and a desktop application for playing the currently playing media; before the least one processor executes the received user interactive voice, it is further configured to execute the computer program or instructions to cause the display device to: controlling the desktop application to send a user wake-up voice and current scene information to the voice service application; wherein The user wake-up voice is used to wake up the voice service application; When the at least one processor performs the conversion of the user-interactive speech into user-interactive text, it is configured to execute the computer program or instructions to cause the display device to: Control the desktop application to send the user interaction voice to the voice service application; After the voice service application recognizes that the current scene information is a role interaction scene, it converts the user interaction voice into user interaction text. After the voice service application recognizes that the current scene information is a non-role interaction scene, it obtains the third reply voice corresponding to the user's interaction voice and plays the third reply voice. According to claim 4, when the at least one processor executes the playback of the first response voice, it is configured to execute the computer program or instructions to cause the display device to: The desktop application is controlled to invoke a multimedia player to play the first reply voice message. According to claim 5, after the at least one processor executes the playback of the first response voice, it is further configured to execute the computer program or instructions to cause the display device to: After the voice service application detects that there has been no voice input for more than a preset time, it sends a return prompt message to the desktop application. The desktop application is controlled to obtain the prompt voice of the third digital human from the server based on the encore prompt information, and then call the multimedia player to play the prompt voice. According to claim 6, when the at least one processor executes the command to control the desktop application to obtain the prompt voice of the third digital human from the server based on the encore prompt information and to call a multimedia player to play the prompt voice, it is configured to execute the computer program or instructions to cause the display device to: The desktop application controls sends a prompt information obtaining request to the server based on the prompt information. The prompt request carries the media asset identifier and the current playback frame image of the currently playing media asset; The desktop application is controlled to obtain the prompt voice of the third digital human fed back by the server; wherein, the third digital human is selected from the candidate digital humans; the prompt voice is determined by the server based on the digital human identifier of the third digital human, and by calling the third digital human according to the media asset identifier and the current playback frame image; When the desktop application controls the third digital human to play the prompt voice when the candidate digital human is the same as the digital human card with the focus on the screen; When the desktop application controls the third digital human and the candidate digital human corresponding to the digital human card where the screen focus is located is different, it focuses the screen focus on the digital human card corresponding to the third digital human and plays the prompt voice. The display device according to any one of claims 1-7, wherein when the at least one processor executes the command to control the display to show a media asset details page associated with the currently playing media asset, it is configured to execute the computer program or instructions to cause the display device to: Send a request to the server to retrieve the details page of the currently playing media asset; Receive detail page data from the server based on the detail page acquisition request; If the details page data contains digital human information, a media asset details page containing the digital human interactive controls is generated based on the details page data, and the display is controlled to show the media asset details page. The display device according to any one of claims 1-7, wherein when the at least one processor executes the command to control the display to show the digital human card associated with the candidate digital human of the currently playing media asset, it is configured to execute the computer program or instructions to cause the display device to: Control the display to show the digital human cards of the candidate digital humans associated with the currently playing media asset, according to the digital human display order; After the at least one processor executes the command to control the display to show the digital human card associated with the candidate digital human of the currently playing media asset, it is further configured to execute the computer program or instructions to cause the display device to: After the first digital human's digital human card is selected, the interactive prompts of the first digital human are obtained and played. The display device according to claim 1, wherein the at least one processor is further configured to execute the computer program or instructions to cause the display device to: When the currently playing media asset is being played and an interactive command is received, a screenshot of the currently playing media asset is taken to obtain the media asset image; Obtain the identified role identifiers obtained by performing role recognition on the media asset images; In addition, query the interactive character identifier associated with the media asset identifier of the currently playing media asset and the character interaction information corresponding to the interactive character identifier; The first interactive role identifier is determined by matching the identified role identifier with the interactive role identifier associated with the media asset identifier of the currently playing media asset. The display is controlled to show the first role option corresponding to the first interactive role identifier and the second role option corresponding to the second interactive role identifier in the interactive card, and the default focus is set on the first role option. The second interactive role identifier is an interactive role identifier other than the first interactive role identifier among the interactive role identifiers associated with the media asset identifier. The first role option is generated based on the role interaction information corresponding to the first interactive role identifier, and the second role option is generated based on the role interaction information corresponding to the second interactive role identifier. The first role option is configured to enter a dialogue with the first interactive role after being selected. According to claim 10, when the at least one processor executes the query of the interactive character identifier associated with the media asset identifier of the currently playing media asset and the character interaction information corresponding to the interactive character identifier, it is configured to execute the computer program or instructions to cause the display device to: Obtain the media asset information of the currently playing media asset, wherein the media asset information includes a media asset identifier and first indication information; If the first indication information indicates that the currently playing media asset supports character interaction, query the interactive character identifier associated with the media asset identifier of the currently playing media asset and the character interaction information corresponding to the interactive character identifier. According to claim 10, when the at least one processor executes the query of the interactive character identifier associated with the media asset identifier of the currently playing media asset and the character interaction information corresponding to the interactive character identifier, it is configured to execute the computer program or instructions to cause the display device to: The query retrieves the second indication information corresponding to the identified role identifier; When the second indication information indicates that the role represented by the identified role identifier supports role interaction, query the interactive role identifier associated with the media asset identifier of the currently playing media asset and the role interaction information corresponding to the interactive role identifier. The display device according to any one of claims 10 to 12, wherein when the at least one processor performs the action of setting the default focus on the first role option, it is configured to execute the computer program or instructions to cause the display device to: Set the label value of the focus label corresponding to the first role option to the first label value, so as to set the default focus on the first role option; wherein, When the label value of the focus label corresponding to the first role option changes to the second label value, the focus moves away from the first role option. According to any one of claims 10 to 12, when the at least one processor executes the command to control the display to display the first role option corresponding to the first interactive role identifier and the second role option corresponding to the second interactive role identifier in the interactive card, it is configured to execute the computer program or instructions to cause the display device to: Obtain the character introduction information associated with the first interactive character identifier; The display is controlled to show the first role option corresponding to the first interactive role identifier and the second role option corresponding to the second interactive role identifier in the interactive card, and to show the role introduction information in the role introduction card. The display device according to any one of claims 10 to 12, wherein the at least one processor is further configured to execute the computer program or instructions to cause the display device to: In response to a dialogue initiation operation triggered by the interactive card, the display is controlled to show a first dialogue prompt message, which is used to prompt a dialogue with the first interactive character represented by the first interactive character identifier. The system receives a first dialogue input from the user, controls the display to show the first dialogue, and outputs a second dialogue response from the first interactive character to the first dialogue. The second dialogue is generated based on the character setting information of the first interactive character and the first dialogue. According to the display device of claim 15, the first dialogue prompt information includes the opening dialogue content of the first interactive character, the opening dialogue content being preset, or generated based on the plot information of the current storyline in the currently playing media asset and the character setting information of the first interactive character; When the at least one processor executes the command to control the display to show the first dialog prompt information, it is configured to execute the computer program or instructions to cause the display device to: Control the display to show the opening dialogue content, and play the opening dialogue content by voice. According to claim 15, when the at least one processor executes the operation of controlling the display to show the first dialogue prompt information in response to a dialogue initiation triggered by the interactive card, it is configured to execute the computer program or instructions to cause the display device to: In response to a dialogue initiation operation triggered by the interactive card, a first tab is generated based on the character interaction information corresponding to the first interactive character identifier, and a second tab is generated based on the character interaction information corresponding to the second interactive character identifier. The display is controlled to show the first dialogue prompt information, the first tab, and the second tab, and the focus is set on the first tab. The second tab is used to switch to dialogue with the second interactive character represented by the second interactive character identifier. The display device according to any one of claims 10 to 12, wherein the at least one processor is further configured to execute the computer program or instructions to cause the display device to: In response to a selection operation for the second role option in the interactive card, the focus is moved from the first role option to the second role option; In response to a dialogue initiation operation triggered by the interactive card, the display is controlled to show a second dialogue prompt message, which prompts the user to engage in dialogue with the second interactive character represented by the second character option. An interactive method applied to a display device, comprising: The display is controlled to show the media asset details page associated with the currently playing media asset; wherein, the media asset details page displays digital human interactive controls; In response to a trigger operation on the digital human interaction control, the display is controlled to show digital human cards of candidate digital humans associated with the currently playing media asset; wherein, the digital human card of each candidate digital human contains a digital human image corresponding to the candidate digital human and / or the name of the character in the currently playing media asset represented by the candidate digital human; After the first digital human's digital human card is selected, input user interaction voice is received, wherein the first digital human is one of the candidate digital humans; Based on the user's interactive voice, the system obtains the first response voice from the server in response to the user's interactive voice and plays the first response voice; wherein the timbre in the first response voice is the timbre of the character name corresponding to the first digital human.