Display device and interaction method

By recognizing the voiceprint features of target characters in audio and video clips through the controller and intelligent agent of the display device, and generating simulated timbre and response tone, the passive nature of the traditional online viewing experience is solved, realizing a simple and convenient immersive interactive experience, and enhancing the user's sense of realism and interest in the interaction.

CN121957431APending Publication Date: 2026-05-01HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HISENSE VISUAL TECH CO LTD
Filing Date
2024-10-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies cannot provide users with a simple and convenient immersive interactive experience. Traditional online movie viewing experiences are passive and complex, and require hardware support for interactive technology.

Method used

The controller of the display device identifies user-triggered interaction events, and the intelligent agent identifies the voiceprint features of the target character from audio and video clips to generate simulated timbre and response tone. The audio output module is then controlled to broadcast the response voice, enabling precise interaction with the movie character.

Benefits of technology

It enhances the realism of the interaction, increases user interest and experience during the interaction process, and provides a simple and convenient immersive interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121957431A_ABST
    Figure CN121957431A_ABST
Patent Text Reader

Abstract

The invention relates to a display device and an interaction method. The apparatus includes: a display configured to display a user interface; the audio output module is configured to play audio; the controller is configured to respond to an interaction event triggered by a user and determine an audio and video clip indicated by the interaction event and a target role in the audio and video clip; the audio and video clips are contents supported by the display to play; recognizing voiceprint features of the target role from the audio and video clip through the intelligent agent, and obtaining simulated timbre for the target role based on the voiceprint features; performing role enhancement on the target role based on interaction content in the interaction event, and determining reply mood and reply content of the target role for the interaction event; and controlling the intelligent agent to generate reply voice aiming at the reply content according to the simulated tone and the reply tone of the target role, and controlling the audio output module to broadcast the reply voice. By adopting the scheme of the invention, immersive interaction experience can be simply and conveniently realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display device technology, and more particularly to a display device and an interaction method. Background Technology

[0002] With the rapid development of artificial intelligence technology and the continuous improvement of people's living standards, online movie viewing has become a part of modern life. However, the traditional online movie viewing experience is often passive, with viewers only able to watch the film.

[0003] To enhance the movie-watching experience, some technologies attempt to increase audience interaction to improve immersion. For example, some online movie-watching platforms offer bullet comments, allowing viewers to post comments and interact with other viewers during the movie. Other technologies attempt to use virtual reality (VR) technology to allow viewers to move freely in a virtual movie world. However, these technologies are relatively complex to operate or require hardware support, making it difficult to provide users with a simple and convenient immersive interactive experience. Summary of the Invention

[0004] This application provides a display device and an interaction method to solve the problem that it is difficult to provide users with a simple and convenient immersive interactive experience in the prior art.

[0005] In a first aspect, some embodiments provide a display device, including:

[0006] The display is configured to show content from a broadcast system or network and / or a user interface, as well as to display a user interface;

[0007] The audio output module is configured to play audio.

[0008] The controller is configured as follows:

[0009] In response to a user-triggered interaction event, the system determines the audio / video segment indicated by the interaction event and the target character within the audio / video segment; the audio / video segment is content that the display supports playing.

[0010] The intelligent agent identifies the voiceprint features of the target character from the audio and video clips, and obtains a simulated timbre for the target character based on the voiceprint features;

[0011] Based on the interaction content in the interaction event, the target character is enhanced to determine the target character's response tone and response content to the interaction event;

[0012] The intelligent agent is controlled to generate a response voice based on the simulated timbre of the target character and the response tone, and the audio output module is controlled to broadcast the response voice.

[0013] The solutions described above have the following advantages or beneficial effects:

[0014] The display device controller responds to user-triggered interactive events, identifies the audio / video clips indicated by the events, and the target characters within those clips. An intelligent agent identifies the voiceprint features of the target characters from the audio / video clips, obtains a simulated voice for the target character based on these features, and enhances the target character's role based on the interactive content of the event. This determines the target character's response tone and content, and the intelligent agent generates a response voice based on the simulated voice and response tone. The audio output module then plays the response voice, accurately simulating the character in the video and enhancing the realism of the interaction. By using the simulated voice, determining the response tone and content, and ensuring a strong correlation between the generated response voice and the target character in the audio / video clips, the device allows for voice interaction with the user in the image of the target character, thereby increasing user interest and improving the user experience.

[0015] Secondly, some embodiments also provide an interaction method, the method comprising:

[0016] In response to a user-triggered interaction event, determine the audio / video segment indicated by the interaction event and the target character in the audio / video segment;

[0017] Identify the voiceprint features of the target character from the audio and video clips, and obtain a simulated timbre for the target character based on the voiceprint features;

[0018] Based on the interaction content in the interaction event, the target character is enhanced to determine the target character's response tone and response content to the interaction event;

[0019] Based on the simulated voice of the target character and the tone of the reply, generate and output a reply voice for the reply content.

[0020] The solutions described above have the following advantages or beneficial effects:

[0021] By responding to user-triggered interactive events, the system identifies the audio / video clips indicated by the events and the target characters within those clips. An intelligent agent then identifies the voiceprint features of the target characters from the audio / video clips, obtaining a simulated voice for each character based on these features. Based on the interactive content of the event, the system enhances the target character's role, determining their appropriate tone and content for responding to the event. A response voice is then generated according to the simulated voice and tone, accurately mimicking the character's voice in the video and outputting it to the user, enhancing the realism of the interaction. Furthermore, by using the simulated voice, determined tone, and content, the system ensures a strong correlation between the generated response voice and the target character in the audio / video clips. This allows the system to interact with the user in the image of the target character, increasing user interest and improving the user experience. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application;

[0024] Figure 2 This is a schematic diagram of the hardware configuration of a display device provided in some embodiments of this application;

[0025] Figure 3 This is a schematic diagram of the hardware configuration of the control device provided in some embodiments of this application;

[0026] Figure 4 This is a schematic diagram of the software configuration of a display device provided in some embodiments of this application;

[0027] Figure 5 This is a schematic diagram of the hardware configuration of a display device provided in some embodiments of this application;

[0028] Figure 6 A flowchart illustrating the interaction method implemented by the controller of a display device according to some embodiments of this application;

[0029] Figure 7 A flowchart illustrating the interaction methods provided in some embodiments of this application;

[0030] Figure 8 These are schematic diagrams of the interface of a display device provided in some embodiments of this application;

[0031] Figure 9 Interaction timing diagrams for some embodiments of the interaction methods provided in this application;

[0032] Figure 10 A flowchart illustrating the interaction methods provided in some embodiments of this application. Detailed Implementation

[0033] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0034] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0035] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0036] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0037] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0038] In this embodiment, display device 200 generally refers to a device with screen display and data processing capabilities. For example, display device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.

[0039] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device provided in some embodiments of this application. For example... Figure 1As shown, users can operate the display device 200 via touch operation, mobile terminal 300, and control device 100. For example, control device 100 can be a remote control, stylus, gamepad, etc.

[0040] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the display device 200. It can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can have software applications installed on it and communicate with the display device 200 via network communication protocols to achieve one-to-one control and data communication. Furthermore, it can transmit audio and video content displayed on the mobile terminal 300 to the display device 200 for synchronized display.

[0041] like Figure 1 The diagram also shows that the display device 200 communicates with the server 400 via various communication methods. This allows the display device 200 to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0042] Display device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function that provides computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.

[0043] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of display device 200.

[0044] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0045] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0046] In some embodiments, the display 260 includes display function components for presenting images and driving components for driving image display. The display 260 is used to receive and display image signals output from the controller 250. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0047] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The display device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the display device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the display device 200 supports Bluetooth connectivity, it needs to have a communication device 220 with Bluetooth functionality.

[0048] The communication device 220 enables the display device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the display device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The display device 200 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.

[0049] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.

[0050] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0051] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0052] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.

[0053] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0054] Figure 3 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of the central control device. (Example) Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output module, a memory, and a power supply.

[0055] The control device 100 is configured to control the display device 200, and to receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.

[0056] In some embodiments, the control device 100 may be an intelligent device. For example, the control device 100 may be equipped with various applications for controlling the display device 200 according to user needs.

[0057] In some embodiments, such as Figure 1 As shown, the mobile terminal 300 or other smart electronic devices can perform similar functions to the control device 100 after installing the application of the control display device 200.

[0058] The controller 110 includes a processor 112, RAM 113, ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation of the control device 100, as well as the communication and cooperation between internal components and the external and internal data processing functions.

[0059] Under the control of the controller 110, the communication interface 130 enables communication of control signals and data signals with the display device 200. The communication interface 130 may include at least one of other near-field communication modules such as WiFi chip 131, Bluetooth module 132, and NFC module 133.

[0060] User input / output module 140, wherein the input interface includes at least one of other input interfaces such as microphone 141, touchpad 142, sensor 143, and button 144.

[0061] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output module 140. The control device 100 is configured with the communication interface 130, such as a WiFi, Bluetooth, or NFC module, which can encode user input commands via WiFi, Bluetooth, or NFC protocols and send them to the display device 200.

[0062] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can also store various control signal instructions input by the user.

[0063] The power supply 180 is used to provide operating power support for the various components of the control device 100 under the control of the controller.

[0064] In some embodiments, the display device 200 may run an operating system to enable user interaction. An operating system is a computer program that manages and controls the hardware and software resources of the display device 200. The operating system can (control the display device) provide a user interface, allowing users to interact with the display device 200 and supporting the running of various applications.

[0065] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for display devices.

[0066] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Library layer, and the Kernel layer.

[0067] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0068] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0069] like Figure 4 As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0070] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0071] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0072] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 4As shown, hardware drivers can be configured in the kernel layer. The kernel layer can contain at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0073] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the function of the display device, the type of operating system, and other factors, the number of levels and the specific level type of the operating system may be expressed in other forms.

[0074] In some embodiments, after the display device is started, it can directly enter the interface of a preset video-on-demand program. The interface of the video-on-demand program may include a navigation bar and a content display area. The content displayed in the content display area changes according to the selected control in the navigation bar. Programs in the application layer can be integrated into the video-on-demand program and displayed through a control in the navigation bar, or they can be further displayed after an application control in the navigation bar is selected.

[0075] In some embodiments, after the display device is started, it can directly enter the display interface of the last selected signal source or the signal source selection interface. The signal source can be a preset video-on-demand program, or at least one of an HDMI interface, a live TV interface, etc. After the user selects different signal sources, the display can display content obtained from different signal sources.

[0076] With the rapid development of artificial intelligence technology and the continuous improvement of people's living standards, online movie viewing has become a part of modern life. However, the traditional online movie viewing experience is often passive, with viewers only able to watch the film. To improve the viewing experience, some technologies attempt to enhance the immersive experience by increasing the interactivity between the viewer and the film. For example, some online movie platforms offer bullet screen features, allowing viewers to comment and interact with other viewers during the viewing process. Other technologies attempt to use virtual reality (VR) technology, allowing viewers to move freely in a virtual film world. However, these technologies are relatively complex to operate or require hardware support, making it difficult to provide users with a simple and convenient immersive interactive experience.

[0077] To address the aforementioned technical problems, a display device is provided in some embodiments, as shown below. Figure 5 The display device includes: a display 502, an audio output module 504, and a controller 506.

[0078] The display 502 is configured to display content from a broadcast system or network and / or a user interface, as well as to display a user interface.

[0079] A display is a device used to show content from broadcasting systems, networks, and / or user interfaces, as well as to display user interfaces. It can be various types of screens, such as LCD displays and OLED displays, capable of presenting images, text, video, and other information to the user in a visual format.

[0080] User interface (UI) refers to the medium through which a system and a user interact and exchange information. It can include various forms such as graphical user interface (GUI), command line interface (CLI), touch interface, and voice user interface (VUI).

[0081] The audio output module 504 is configured to play audio.

[0082] The audio output module is the component in a display device responsible for playing audio. It can be a speaker, headphone jack, etc., and converts electrical signals into sound signals for output, allowing the user to hear the audio content. The audio output interface can also play back the user's interactive input, providing a personalized voice response that enhances the user's immersive interactive experience.

[0083] The controller 506 is the core control unit of the display device, responsible for processing various input signals, performing logical operations, and controlling the operation of other components. In some embodiments, the controller receives user-triggered interactive events and performs a series of processing operations to generate response voice and control the audio output module to play it.

[0084] like Figure 6 As shown, controller 506 is configured to perform steps 602 to 608.

[0085] Step 602: In response to an interaction event triggered by the user, determine the audio / video segment indicated by the interaction event and the target character in the audio / video segment; the audio / video segment is content that the display supports playing.

[0086] Interaction events are specific actions or operations triggered by users, such as clicking a button on a display, touching the screen, or issuing voice commands. Interaction events indicate a user's intention to interact with the system and contain relevant information, such as the indicated audio / video clip and the target character.

[0087] Audio and video clips are content that a monitor supports playing, including both video and audio components. These clips can be various forms of audio and video resources, such as movie clips, TV series episodes, and short videos.

[0088] A target character is a specific character designated by the user or identified by the system within an audio or video clip. For example, in a movie, a user might designate a particular protagonist as the target character in order to simulate that character's interactions.

[0089] In some embodiments, users trigger interactive events through specific actions, such as typing certain content on the user interface, clicking a button, or uttering specific voice commands. The interactive events contain the audio / video clips indicated by the user and information about the target character. The audio / video clips and the information about the target character can be content directly embedded in the interactive content or content retrieved through a search based on the interactive content.

[0090] After receiving an interaction event triggered by the user, the controller first identifies the audio / video segment indicated by the interaction event. This can be achieved by analyzing user input information, querying playback history, etc. Then, the controller can combine the interaction content in the interaction event with the audio / video segment to identify the target character from the audio / video segment. In some embodiments, identifying the target character from the audio / video segment can employ image recognition technology, speech recognition technology, etc., based on the character's physical features, voice characteristics, etc.

[0091] For example, when the display device is powered on, the user can trigger interactive events with the display device via voice or text, and input interactive content via voice or text. After receiving the user's input, the controller can control the audio output module to play a response voice message in response to the interactive content.

[0092] In this context, the response voice for the interactive content refers to the voice output used to answer questions related to the interactive content. In some embodiments, after receiving the interactive content, the display device can send it to the server. The server analyzes the interactive content, obtains the response voice, and feeds it back to the display device for playback. In some embodiments, the server obtains the response voice by analyzing the interactive content based on a large language model. This large language model can be a general model applicable to various fields, or an optimized model obtained by fine-tuning a general model according to specific fields. In this embodiment, by analyzing and processing the interactive content through the server, the powerful computing capabilities of the server can be utilized to quickly process the interactive content and rapidly obtain the response voice.

[0093] It is understood that, in some embodiments, for display devices equipped with strong computing power, the display device can directly analyze the interactive content and obtain the response voice of the interactive content.

[0094] The user's input can be received through the input interface of the display device. The input interface includes at least one of voice input, text input, and image input.

[0095] Optionally, after the display device is triggered to enter text or image input mode, the display device can bring up the built-in keyboard, allowing users to select and send text or select and send images via keyboard input.

[0096] Optionally, users can trigger the keyboard via touchscreen or remote control buttons. The display device can then show candidate text or images corresponding to the keyboard combination. The display device can further display the user-selected candidate text or image in the input field. After the user confirms the editing and triggers the send operation, the display device can display the sent text content, image content, or a combination of text and image content in the interactive dialog box. Using the above method, users can choose to interact with the display device using images or text, etc., to achieve multiple modal interaction methods.

[0097] Optionally, when the display device is triggered to enter voice control mode, it can receive voice commands input by the user through the voice input interface.

[0098] Optionally, users can control the display device to enter voice control mode by operating designated buttons on the remote control. In practical applications, a pre-defined mapping between voice control mode commands and remote control buttons is established. For example, a voice button can be set on the remote control; when the user touches this button, the remote control sends a voice control mode command to the controller. At this point, the display device enters voice control mode and can successfully receive the user's voice commands.

[0099] Optionally, the controller can first bind the voice control mode command to the correspondence between multiple remote control buttons. When the user touches the multiple buttons bound to the voice control mode command, the remote control issues the voice control mode command. In one feasible embodiment, the button bound to the voice control mode command is the directional keys (left, down, left, down). The remote control only sends the voice control mode command to the controller if the user touches the button (left, down, left, down) continuously within a preset time. Using the above binding method can avoid the voice control mode command being issued due to user error. This application embodiment only provides several exemplary binding relationships between voice control mode commands and buttons. In actual application, the binding relationship between voice control mode commands and buttons can be set according to the user's habits, and no further limitations are made here.

[0100] Optionally, users can control the display device to enter voice control mode via voice. For example, the display device can be triggered to enter voice control mode by saying a far-field wake-up word, such as "Xiao Ju Xiao Ju". Once the display device is triggered to enter voice control mode, it can monitor the user's voice input in real time, allowing the user to further speak voice commands. Users can also use a microphone to input voice commands into the display device.

[0101] Optionally, once the display device is triggered to enter voice control mode, users can also send commands to the display device in text form via mobile phones, remote controls, or other devices to prevent the display device from being unable to receive user voice commands if the microphone malfunctions.

[0102] Optionally, after the display device receives a voice command input by the user, the controller can send the received voice data to a voice recognition service to convert it into interactive text and display the interactive text on the screen. The recognition operation for user voice commands can be found in relevant technologies, and will not be elaborated upon here.

[0103] For ease of description, the following embodiments use voice content as an example to illustrate the solution of this application. It can be understood that in other embodiments, the user-inputted interactive content can be switched between text mode and image mode.

[0104] Step 604: The agent identifies the voiceprint features of the target character from the audio and video clips, and obtains the simulated timbre for the target character based on the voiceprint features.

[0105] An intelligent agent is a software or hardware module with intelligent processing capabilities. Specifically, an intelligent agent refers to an AI agent used for media resource searching. An AI agent is a system driven by a large language model, possessing the ability to autonomously understand, perceive, plan, remember, and use tools, and can automatically execute complex tasks. Unlike traditional artificial intelligence, an AI agent has the ability to independently think and invoke tools to gradually achieve a given goal. This intelligent agent can perceive its environment, make decisions, and execute actions, possessing autonomy and adaptability. It can rely on the capabilities granted by AI to complete specific tasks and continuously improve itself in the process. It can be imagined as a super-intelligent robotic assistant with a powerful brain and learning ability, capable not only of understanding human language but also of improving its skills in a specific field through learning and data analysis. In some embodiments, the intelligent agent can identify the voiceprint features of a target character from audio and video clips, obtain simulated timbre based on these features, enhance the target character, determine the tone and content of responses, and generate response speech. The intelligent agent can be deployed on a server or on a display device with sufficient computing power.

[0106] Voiceprint features are used to characterize the unique vocal characteristics of each role, including timbre, pitch, speech rate, and rhythm. By analyzing voiceprint features, different roles can be identified and their voices can be simulated.

[0107] Simulated voice timbre is generated based on the voiceprint characteristics of the target character, creating a voice timbre similar to that character. By simulating voice timbre, the intelligent agent can generate more realistic response speech, allowing users to interact with objects that simulate the target character.

[0108] The intelligent agent identifies the voiceprint features of a target character from audio and video clips. This can be achieved through audio analysis techniques, extracting various characteristic parameters of the sound, such as frequency, amplitude, and duration, to determine the voiceprint features. Based on the identified voiceprint features, the intelligent agent obtains a simulated timbre for the target character.

[0109] Specifically, the agent's corresponding memory can store simulated voices for many different characters. These simulated voices can be generated by the agent based on historical interaction records or pre-configured. The memory can be on a display device or on a server, specifically in the same location as the agent's deployment to facilitate interaction between the agent and the memory. The agent can first use voiceprint features to check if a simulated voice for the target character exists. If it exists, it can directly retrieve the simulated voice from the stored data. If it does not exist, the agent can use voice synthesis technology to convert the voiceprint features into specific voice parameters, thereby generating a voice similar to the target character.

[0110] In some specific embodiments, the controller can collect and analyze the voice of the target character based on audio and video clips. The collection and analysis process can specifically employ audio signal processing techniques, such as Fourier transform and Mel-frequency cepstral coefficient (MFCC) extraction, to extract voiceprint features from the speech signal. These extracted voiceprint features may include, but are not limited to, timbre features, pitch features, speech rate features, and prosodic features. Timbre features describe the texture and color of the sound, such as bright, deep, or soft; pitch features represent the high and low frequencies of the sound, reflecting the character's voice pitch; speech rate features refer to the speed at which the character speaks, whether fast or slow; and prosodic features include stress, pauses, and intonation, reflecting the rhythm and emotional expression of the character's speech.

[0111] The intelligent agent is equipped with a timbre model for timbre cloning by extracting voiceprint features. The timbre model can be built using machine learning algorithms, such as deep neural networks (DNNs) and convolutional neural networks (CNNs), by training on a large amount of voiceprint feature data to learn the timbre characteristics of the target character. The timbre model can map voiceprint features to a specific timbre parameter space, thereby simulating the timbre of the target character.

[0112] After obtaining the timbre model, the agent can generate a simulated timbre for a target character based on the input voiceprint features. The specific process may include: inputting the extracted voiceprint features into the timbre model; the timbre model calculating the corresponding timbre parameters based on the input voiceprint features; and generating a simulated sound timbre based on the timbre parameters. The generation of the simulated timbre can use audio synthesis techniques, such as waveform synthesis and parametric synthesis, to convert the timbre parameters into sound signals, thereby obtaining the simulated timbre.

[0113] Furthermore, to improve the quality and realism of the simulated timbre, the agent can perform optimization and adjustments. For example, the generated simulated timbre can be post-processed, such as filtered, equalized, and reverb, to enhance the texture and spatiality of the sound.

[0114] In an alternative embodiment, the agent can be deployed on a server that provides service support to the display device and enables voice interaction between the display device and the server.

[0115] For example, refer to Figure 7Specifically, in response to a user-triggered interaction event, the display device sends the corresponding interaction content and the content displayed on the device's screen to the server. Based on the interaction content and the displayed content, the server determines the audio / video segment indicated by the interaction event, as well as the target character within the segment, and sends this information to the intelligent agent. The intelligent agent then identifies the voiceprint features of the target character from the audio / video segment, obtains a simulated voice for the target character based on these features, and stores it.

[0116] Step 606: Based on the interaction content in the interaction event, perform role enhancement on the target role and determine the target role's response tone and response content to the interaction event.

[0117] The tone of response refers to the tone of voice displayed by the target character in response to the interactive event. The tone of response can be determined based on factors such as the interactive content and the target character's personality traits. The content of the response is the specific answer given by the target character to the interactive event. The content of the response can be determined through analysis of the interactive content, searching a knowledge base, and other methods.

[0118] Intelligent agents can enhance the target character based on the interactive content in interactive events, determining the target character's tone and content of response to the interactive events. Specifically, character enhancement can be achieved by analyzing the semantics and emotion of the interactive content, combined with the target character's personality traits, background information, and other details, to determine the target character's behavior in the current interactive context, resulting in an enhanced character description that facilitates the determination of the target character's tone and content of response to the interactive events.

[0119] In some embodiments, the agent can first analyze the interactive content of the interactive event. For example, natural language processing techniques, such as semantic analysis and sentiment analysis, can be used to extract key information and emotional tendencies from the interactive content. Simultaneously, the agent can determine the target character's personality traits, backstory, emotional state, and other characteristics based on information from audio and video clips. This information can be obtained by analyzing the character's lines, behaviors, and expressions, and by referring to plot summaries, character settings, and other materials. The target character might be a brave and optimistic hero, or a clever and resourceful villain; these characteristics will influence the character's tone and content of their responses.

[0120] The process of enhancing a target character can include determining the character's emotional state, simulating the character's thought process, and considering the character's relationships and positions. Based on the emotional tone of the interactive content and the target character's personality traits, determine the character's emotional state when responding, such as happiness, anger, or sadness. Consider the target character's background and personality, simulate the character's thought process and decision-making style when facing the interactive content. If the interactive content involves other characters, it is also necessary to consider the target character's relationships and positions with other characters to determine the appropriate response.

[0121] Furthermore, based on the role enhancement results, the agent determines the target character's tone of voice in response to the interactive event. The tone of voice can be chosen based on the character's emotional state, personality traits, and the interaction scenario. If the target character is humorous, the tone of voice can be lighthearted and witty; if the character is in a tense situation, the tone of voice might be serious and tense. Based on the interaction content and the role enhancement results, the agent generates the target character's response to the interactive event, ensuring that the response is relevant to the interaction content and conforms to the target character's personality traits and the context of the storyline. Natural language generation techniques, such as template generation and neural network generation, can be used to generate natural and fluent response content.

[0122] Step 608: Control the intelligent agent to generate a reply voice based on the simulated timbre and reply tone of the target character, and control the audio output module to broadcast the reply voice.

[0123] The response voice is a speech signal generated by the intelligent agent based on the simulated timbre and tone of the target character, tailored to the content of the response. The response voice is played back to the user via the audio output module, enabling interaction with the user.

[0124] The tone of the response can be adjusted according to the character's personality and emotional state, while the content of the response can be generated through methods such as querying a knowledge base and logical reasoning. After determining the tone and content of the response, the intelligent agent can generate a response voice based on the simulated voice of the target character and the tone of the response.

[0125] In some embodiments, the intelligent agent can employ speech synthesis technology to convert the response content into a speech signal, and apply parameters of the target character's simulated timbre and tone of voice to generate a more realistic response. The controller controls the audio output module to play the response, allowing the user to hear the simulated response from the target character. Simultaneously, the display can show relevant text information, images, etc., as needed to enhance the interactive experience.

[0126] The solutions described above have the following advantages or beneficial effects:

[0127] The display device controller responds to user-triggered interactive events, identifies the audio / video clips indicated by the events, and the target characters within those clips. An intelligent agent identifies the voiceprint features of the target characters from the audio / video clips, obtains a simulated voice for the target character based on these features, and enhances the target character's role based on the interactive content of the event. This determines the target character's response tone and content, and the intelligent agent generates a response voice based on the simulated voice and response tone. The audio output module then plays the response voice, accurately simulating the character in the video and enhancing the realism of the interaction. By using the simulated voice, determining the response tone and content, and ensuring a strong correlation between the generated response voice and the target character in the audio / video clips, the device allows for voice interaction with the user in the image of the target character, thereby increasing user interest and improving the user experience.

[0128] In some embodiments, the display device further includes a memory configured to store simulated timbres built by the agent for different roles.

[0129] In the process of identifying the voiceprint features of a target character from audio and video clips through an agent and obtaining a simulated timbre for the target character based on the voiceprint features, the controller is further configured to: perform voiceprint recognition on the audio and video clips through an agent to obtain the voiceprint features of the target character; and search for the simulated timbre of the target character in the memory according to the voiceprint features.

[0130] The intelligent agent can construct simulated voices for different roles. These simulated voices are obtained by analyzing and processing large amounts of audio and video data, extracting the voiceprint features of each role, and synthesizing corresponding simulated voices based on these features. The simulated voices constructed by the intelligent agent are stored in the memory of the display device. The memory can be organized and managed according to the role's name, number, etc., for quick searching and matching.

[0131] When it is necessary to identify a target character in an audio / video clip and obtain their simulated voice timbre, the controller activates an agent to perform voiceprint recognition on the audio / video clip. The agent employs advanced audio analysis techniques, such as Fourier transform and Mel-frequency cepstral coefficient extraction, to extract the voice features of each character from the audio / video clip. By analyzing and comparing these voice features, the agent determines the voiceprint characteristics of the target character.

[0132] In this embodiment, by setting up a memory in the display device to store the simulated timbres built by the agent for different roles, and by using the agent to perform voiceprint recognition and search for simulated timbres during the interaction process, the simulated timbres of the target role can be obtained quickly, avoiding the repeated construction of simulated timbres and improving resource utilization.

[0133] In some embodiments, the controller is further configured to: clone the voice of the target character according to the voiceprint features and construct a simulated voice for the target character if the simulated voice of the target character does not exist in the memory; and store the simulated voice of the target character in the memory.

[0134] Once the voiceprint characteristics of the target character are obtained, the controller instructs the agent to search for the target character's simulated voice timbre in memory based on these characteristics. The agent matches the voiceprint characteristics with the simulated voice timbre stored in memory. The matching process can employ methods such as similarity calculation and pattern recognition to determine the simulated voice timbre that best matches the target character's voiceprint characteristics. If a matching simulated voice timbre is found, the agent returns it to the controller for subsequent use in generating a response speech.

[0135] When the controller needs to acquire a simulated voice for a target character, it first checks whether the simulated voice for that target character already exists in the memory. This can be done by querying the character list in memory or using a specific indexing mechanism to quickly determine if the simulated voice for the target character exists. If the simulated voice for the target character does not exist in the memory, the controller initiates the voice cloning function.

[0136] Voice cloning is a technique that analyzes the voiceprint characteristics of a target character and synthesizes a similar voice timbre. The voice cloning process specifically includes:

[0137] Voiceprint feature analysis: The controller performs in-depth analysis of the voiceprint features of the target character, including extracting timbre feature parameters such as frequency distribution, harmonic structure, and formants; analyzing pitch features to determine the high and low frequency range of the voice; studying speech rate features to understand the speed and rhythm of the character's speech; and analyzing prosodic features, including stress, pauses, and intonation.

[0138] Constructing a simulated timbre: Based on the analyzed voiceprint feature parameters, the controller uses timbre synthesis technology to construct a simulated timbre for the target character. Specifically, methods such as waveform synthesis and parametric synthesis can be used to generate a sound signal similar to the target character based on the voiceprint feature parameters. During the synthesis process, various parameters can be adjusted to ensure that the simulated timbre is as close as possible to the target character's real voice.

[0139] To improve the quality and realism of the simulated sound, the controller can optimize and adjust the constructed simulated sound. For example, it can filter, equalize, and reverb the sound to enhance the texture and spatiality of the sound; it can also fine-tune the simulated sound based on user feedback and evaluation to better meet the user's expectations.

[0140] After constructing the simulated voice for the target character, the controller stores it in memory, such as by character name or number, for quick retrieval and matching later. While storing the simulated voice, the controller can also store related information, such as the character's description and the source audio / video clips, for better management and identification of the simulated voice.

[0141] In this embodiment, by cloning the simulated timbre of the target character when it does not exist in the memory, and storing the constructed simulated timbre in the memory, the simulated timbre library of the system can be continuously enriched, the simulated timbre of the target character can be obtained quickly and conveniently, and the newly constructed simulated timbre can be stored when an accurate simulated timbre is obtained, thereby improving the response speed.

[0142] In some embodiments, during the process of performing role enhancement on the target role based on the interaction content in the interaction event and determining the target role's response tone and response content to the interaction event, the controller is further configured to: if the interaction content in the interaction event contains first description content for the target role, rewrite or expand the first description content to obtain second description content for the target role; and determine the target role's response tone and response content to the interaction event based on the second description content and the interaction content.

[0143] The second description content satisfies at least one of the following conditions: its content richness is greater than that of the first description content, and its content simplicity is greater than that of the first description content.

[0144] The first description is the original description of the target character within the interactive content, which may be colloquial, simple, or lack detail. The second description is a rewritten or expanded version of the first description, satisfying at least one of the following conditions: greater content richness or greater content conciseness than the first description.

[0145] Specifically, when the controller receives the interaction content from the interaction event, it first analyzes whether the interaction content contains a first description of the target role. This can be done using natural language processing techniques, such as keyword extraction and semantic analysis, to determine if the interaction content contains a specific description of the target role. In some embodiments, the first description can be empty.

[0146] In some embodiments, if the interaction content contains first descriptive content, the controller can control the agent to rewrite or expand the content to obtain second descriptive content.

[0147] Content rewriting can include synonym replacement or sentence structure adjustment. Synonym replacement involves replacing some words in the initial description with words that have similar meanings to enrich the expression. For example, replacing "brave" with "fearless." Sentence structure adjustment involves changing the sentence structure of the description to make the expression more vivid. For example, changing declarative sentences into interrogative or exclamatory sentences.

[0148] Content expansion can include adding details and introducing relevant information. Adding details involves adding more detailed information to the initial description based on the target character's characteristics and background. For example, if the initial description is "This character is very powerful," it can be expanded to "This character performs exceptionally well in battle; his skills are powerful, his decisions are decisive, and he always manages to save his teammates at crucial moments." Introducing relevant information involves combining the context of the interactive event or the target character's relevant plot points to introduce other relevant information to enrich the description. For example, if the interactive content is about a battle scene, and the initial description is "This character is very brave," it can be expanded to "This character was extremely brave in this fierce battle; he was undaunted by the enemy's powerful offensive, leading his teammates in a valiant resistance, demonstrating extraordinary leadership skills."

[0149] For example, if a film clip identifies the character as "Sun Wukong," the user's input could be, "Now you play Sun Wukong." Another example is a user-defined character, "Niu Hulu Xiaomi," whose description might be: "I hope she appears gentle and quiet, but is actually clever and cunning, not only multi-talented but also adept at palace intrigue, like Zhen Huan after her return to the palace in 'The Legend of Zhen Huan.'" Yet another example is a user-defined character, "Xiao Hai," without a specific description. In this case, the description can be left blank, guiding the user to complete it, or information about the user's character settings can be collected during subsequent conversations and added later.

[0150] The second description obtained by prompting the large language model to generate or modify the character description is as follows: For example, the second description for the character "Sun Wukong" is: He possesses the supernatural power of seventy-two transformations and the magical golden cudgel. He is witty and brave, once caused havoc in the Heavenly Palace, later became Tang Sanzang's disciple, accompanied him on his journey to the West to obtain Buddhist scriptures, and after experiencing many hardships and dangers, finally became the Victorious Fighting Buddha. Another example is the second description for the character "Niu Hulu Xiaomi" in the above embodiment: She appears gentle and quiet, but is actually intelligent, cunning, calm, rational, and strong-willed. She is not only multi-talented but also very skilled in palace intrigue.

[0151] Furthermore, based on the obtained second description and interaction content, the agent determines the target character's tone and content of response to the interaction event. The tone of response can be chosen based on the emotional tone conveyed by the second description and the target character's personality traits. For example, if the second description is full of praise and the target character is humble, the response tone can be gentle and modest; if the second description is challenging and the target character is combative, the response tone can be assertive and confident. The content of the response can be generated by combining the second description and the specific needs of the interaction event. For example, if the interaction is a question and the second description emphasizes a certain strength of the target character, the response can address that strength and further elaborate on related plot points or viewpoints.

[0152] In this embodiment, the second description is obtained by rewriting or expanding the first description in the interactive content, and the response tone and content of the target character are determined based on the second description. Based on the rewritten or expanded description, a more accurate response tone and content are obtained, thereby making the interaction richer and more personalized and improving the user's interactive experience.

[0153] In some embodiments, when the content richness of the first description content is less than a preset richness condition, the controller is further configured to: During the process of acquiring the second description content for the target role, if the content richness of the first description content is less than a preset richness condition.

[0154] Use a search engine or film and television knowledge query tool to search for supplementary information about the target character; use a large language model to integrate the first description and the supplementary information to obtain a second description for the target character.

[0155] Specifically, the controller first analyzes the content richness of the first description through an agent. For example, it evaluates the richness by calculating metrics such as the number of keywords, sentence length, and semantic complexity within the first description. The agent then compares the evaluated content richness with a preset richness condition. If the content richness of the first description is less than the preset condition, further processing is required to obtain richer second description content.

[0156] If the initial description lacks sufficient content, the agent uses a search engine or film / television knowledge query tool to find supplementary information about the target character. When using a search engine, keywords such as the target character's name or the title of the work can be used. Search results may include news reports, forum discussions, film reviews, etc., from which various information related to the target character can be extracted. Film / television knowledge query tools can provide more professional and detailed character information, such as the character's personality traits, development, and important events. The agent can choose an appropriate film / television knowledge query tool as needed and input relevant information about the target character to perform the query.

[0157] After obtaining supplementary character information, the agent uses a large language model to integrate the first description and the supplementary character information. The large language model first understands and analyzes the first description and the supplementary character information, extracting key information and semantic relationships. Then, based on the extracted information, the large language model integrates the content to generate a richer and more comprehensive second description. During the integration process, information can be filtered, reorganized, and polished as needed to ensure the quality and readability of the second description. For example, if the target character is "Sun Wukong," the first description is empty, while the supplementary character information includes: "He possesses the supernatural powers of seventy-two transformations and a magical golden cudgel. He is witty and brave, once wreaked havoc in the Heavenly Palace, later became Tang Sanzang's disciple, accompanied him on his journey to the West, overcame numerous hardships, and finally became the Victorious Fighting Buddha." After content integration, the large language model generates a second description for the target character. This second description is used to subsequently determine the target character's response tone and content.

[0158] In this embodiment, when the content richness of the first description is less than a preset richness condition, a search engine or film and television knowledge query tool is called to obtain supplementary information about the character, and a large language model is used for content integration. This can effectively enrich the description of the target character and provide a more accurate and comprehensive basis for determining the target character's response tone and content.

[0159] In some embodiments, supplementary character information includes at least one of the target character's personality, background, and language habits.

[0160] The supplementary character information is additional information about the target character obtained by using search engines or film and television knowledge query tools. This information is used to enrich the description of the target character in order to better determine the target character's tone and content of response to interactive events.

[0161] Character personality refers to the individual traits possessed by the target character, such as courage, kindness, intelligence, and cunning. Character personality influences a character's behavior and speech, and is one of the important bases for determining the tone and content of a response. Character background refers to information about the target character's origins, experiences, and upbringing. Character background can provide explanations for a character's behavior and decisions, helping to understand the character more deeply and thus better determine the content of the response. Character language habits refer to the target character's habitual phrases, catchphrases, and interjections when speaking. Character language habits can make the response more realistic and vivid, enhancing the user's sense of immersion. For example, for the classic character "Sun Wukong," we can add his catchphrase "I, Old Sun," and classic lines such as "Here comes Old Sun~," "Monster, where do you think you're going!" and "Heaven and earth are inherently incomplete, and scriptures are also incomplete; this cannot be accomplished by human effort."

[0162] As before, supplementary information about the target character can be obtained by using search engines or film and television knowledge query tools. These tools can provide various information about the target character, including the character's personality, background, and language habits.

[0163] A character's personality can be gleaned from their behavior, dialogue, and evaluations by other characters. For example, if a target character frequently displays courageous and fearless behavior in the story, and other characters praise their bravery, then it can be determined that the character possesses the personality trait of courage.

[0164] Character background information can be gleaned from the plot summary of the work, flashbacks of the character, and other sources. For example, knowing that the target character experienced a major disaster may have influenced their personality and behavior; this background factor can be considered when determining the content of the response.

[0165] A character's language habits can be determined by analyzing their dialogue. For example, if a character frequently uses specific catchphrases or interjections, these habits can be appropriately incorporated into responses to better reflect the character's personality.

[0166] After obtaining supplementary information about the character, this information is used to determine the target character's tone and content of response to the interactive event.

[0167] If the character is brave and decisive, the tone of the response can be more firm and confident. For example, in response to a user's question, "Are you afraid of this challenge?", the target character could answer, "I'm never afraid of challenges; moving forward bravely is my motto."

[0168] A person's background can provide more detail and depth to their response. For example, if the target person comes from a happy family, when answering questions about life, they could mention how their family background has influenced their perspective on life.

[0169] Using a character's language habits can make your replies more vivid and authentic. For example, if the target character frequently uses interjections like "hey" or "wow," you can appropriately incorporate these interjections into your replies to increase their appeal.

[0170] Furthermore, after determining the tone and content of the response, the agent can optimize it to make it more natural and fluent. For example, it can check for grammatical errors, logical inconsistencies, and other problems in the response and correct them. Simultaneously, the agent can also polish the response as needed, making it more eloquent and expressive.

[0171] In this embodiment, by acquiring supplementary information about the character, including the character's personality, background, and language habits, and using this information to determine the target character's tone and content of response, the interaction can be made more realistic, vivid, and personalized, thereby enhancing the user's interactive experience.

[0172] In some embodiments, the video clip is the target video clip currently playing on the display; the interaction event is an interaction event triggered by the user while watching the target video clip.

[0173] In the process of performing character enhancement on the target character based on the interactive content in the interactive event, and determining the target character's response tone and response content to the interactive event, the controller is further configured to: perform character enhancement on the target character based on the interactive content in the interactive event and the movie to which the target video clip belongs, to obtain descriptive content for the target character; determine the target character's response tone to the interactive content according to the descriptive content of the target character; and generate response content to the interactive content based on the plot and descriptive content performed by the target video clip.

[0174] The target video clip is the currently playing video segment on the monitor, which is the content the user is watching. The target video clip can be a portion of various video resources such as movies, TV series, and short videos. Interaction events characterize the interactive behaviors triggered by the user while watching the target video clip, such as clicking buttons, voice commands, and touching the screen. Interaction events contain the user's interactive content, used to interact with the simulated target character. For example, Figure 8 As shown, the target character's image can be displayed on the monitor.

[0175] The film to which the target video clip belongs refers to the complete film or television work to which the target video clip belongs. The film contains rich plot, character information, and background settings, playing a crucial role in understanding the target character and generating response content. The description of the target character is a detailed description obtained after character enhancement. The description includes information such as the target character's personality traits, emotional state, and behavioral motivations, used to determine the tone and content of the response.

[0176] In some embodiments, while a user is watching a target video segment, the controller continuously monitors the user's actions to detect any interactive events. If the user triggers an interactive event, the controller retrieves the interactive content from the event. The controller analyzes the currently playing target video segment to determine the movie it belongs to. Movie information can be determined through video metadata, filenames, playback history, etc. Based on the interactive content in the interactive event and the movie to which the target video segment belongs, the controller enhances the target character through an intelligent agent. The specific steps are as follows:

[0177] The agent extracts information such as plot, character relationships, and background settings from the film. This information helps to better understand the target character's position and role in the film. By combining interactive content from interactive events with film information, the agent further understands the user's focus and expectations regarding the target character. For example, if the interactive content concerns a specific action of the target character, the agent can analyze the motivation and consequences of that action based on the film's plot. Based on the analysis results, the agent generates a description of the target character. This description can include information about the target character's personality traits, emotional state, behavioral motivations, goals, and values. Furthermore, based on the target character's description, the agent determines the tone of the target character's response to the interactive content. For example, if the target character is described as humorous, the response tone can be lighthearted and humorous; if the target character is serious, the response tone can be solemn and serious.

[0178] Based on the plot of the target video clip and the description of the target character, the intelligent agent generates responses to the interactive content. Specifically, this includes: considering the plot of the target video clip to ensure the responses align with the development of the storyline. Conflicts between the responses and the plot may negatively impact user immersion. The intelligent agent generates responses that match the target character's personality traits based on the description of the target character. For example, if the target character is brave and decisive, the responses can reflect their courage and decisiveness. In some embodiments, the intelligent agent can also optimize the generated responses to make them more natural, fluent, and expressive. For example, it can check for and correct grammatical errors and logical inconsistencies. Furthermore, the intelligent agent can polish the responses as needed to make them more literary and engaging.

[0179] The intelligent agent sends the determined tone of the response and the generated response content to the audio output module, which then broadcasts the response content in a simulated voice of the target character. At the same time, relevant text information or images can be displayed on the screen to enhance the interactive effect.

[0180] In this embodiment, by enhancing the target character based on the interactive content in the interactive event and the film to which the target video segment belongs, and determining the tone and content of the response, a more realistic, vivid and personalized interactive experience can be provided to the user, enhancing the interactivity between the user and the video content.

[0181] In some embodiments, the agent is a single agent among multiple agents configured for the display device.

[0182] During the process of executing the control agent to generate a response voice based on the simulated timbre and tone of the target role, and controlling the audio output module to broadcast the response voice, the controller is further configured to: control a single agent to generate a response voice based on the simulated timbre and tone of the target role, and have the single agent feed the response voice back to the multi-agent; control the multi-agent to broadcast the response voice through the audio output module.

[0183] In this multi-agent system, a single agent is an independent agent within a multi-agent system. In this implementation, the single agent is responsible for generating a response voice based on the target character's simulated voice timbre and tone, and then feeding this response voice back to the multi-agent system. The multi-agent system is a collection of agents composed of multiple single agents. The multi-agent system collaborates to complete various tasks, with each agent performing different tasks. Finally, the output module of the multi-agent system provides a unified output, such as controlling the audio output module to play the response voice.

[0184] In some embodiments, when a response voice needs to be generated, the controller directs a single agent to generate a response voice based on the simulated timbre and tone of the target character. The single agent first acquires relevant parameters of the simulated timbre and tone of the target character. These parameters can be provided by the controller or read from memory. Then, the single agent uses speech synthesis technology to convert the response content into a speech signal and applies the parameters of the simulated timbre and tone of the response to make the response voice more realistic. For example, if the simulated timbre of the target character is a deep male voice and the tone of the response is serious, the single agent can use a corresponding speech synthesis algorithm to generate a response voice with a deep timbre and serious tone.

[0185] After a single agent generates a response voice, it sends the response voice back to the multi-agent team. Feedback can be achieved through internal communication mechanisms, such as message passing or shared memory. Upon receiving the response voice from the single agent, the multi-agent team can further process and optimize the response voice, such as adding sound effects or adjusting the volume.

[0186] The controller directs multiple agents to broadcast their responses via an audio output module. The agents send their responses to the audio output module, which converts the electrical signals into sound signals and outputs them, allowing the user to hear the target agent's reply. Simultaneously, the controller can control the display to show relevant text information, images, etc., as needed to enhance the interactive experience.

[0187] During the generation and playback of response voice messages, various anomalies may occur, such as failure of a single agent to generate response voice messages or audio output module malfunction. The controller should be able to detect and handle these anomalies and take appropriate measures, such as regenerating the response voice message or switching the audio output device, to ensure that the response voice message can be played normally to the user.

[0188] In this embodiment, by controlling a single agent to generate a response voice according to the simulated timbre and tone of the target character, and then having the single agent feed the response voice back to multiple agents, and then controlling the multiple agents to broadcast the response voice through the audio output module, this implementation method can achieve efficient and accurate voice interaction function, providing users with a better interactive experience.

[0189] In some embodiments, an online movie-watching immersive interactive agent based on voiceprint recognition and a large language model is provided. This agent can function as a single agent within a multi-agent service, providing stylistic dialogue capabilities by recognizing and simulating the voices and speaking styles of different IPs and characters in the film, thus achieving immersive interaction with the audience. The implementation of the interaction process specifically includes the following modules:

[0190] 1. Character Recognition Module: This module calls upon voiceprint recognition and voiceprint cloning model services. These services utilize deep learning techniques, such as deep neural networks (DNNs) or convolutional neural networks (CNNs), to extract voiceprint features from characters in film and television clips, enabling the registration or recognition of character voices. Furthermore, voiceprint cloning technology is used to simulate the voiceprint features of characters, achieving the cloning of their voices. For example, given an audio clip of Sun Wukong from a film or television series, voiceprint recognition and cloning technology are used to obtain Sun Wukong's timbre, which is then named "wukong".

[0191] 2. Role Building Module: Query the memory module to obtain role setting information; for new roles, specify role setting information through role knowledge retrieval, user definition, large language model generation, etc.

[0192] Example 1: The character identified in the film clip is "Sun Wukong". By recognizing the interactive content, it can be determined that the user's request is "Now you play Sun Wukong".

[0193] Example 2: The user-defined character "Niu Hulu Xiaomi" can be identified by recognizing the interaction content. The user's needs can be determined as follows: I hope that she is gentle and quiet on the surface, but actually smart and cunning. She is not only multi-talented, but also very good at palace intrigue, just like Zhen Huan after returning to the palace in "The Legend of Zhen Huan".

[0194] Example 3: If a user creates a custom character "Xiao Hai" but does not provide a specific character description, it can be left blank for now, guiding the user to complete it, or information about the user's character settings can be collected in subsequent conversations and added later.

[0195] The character descriptions generated or refined using the prompting large language model are as follows:

[0196] The character in Example 1 is described as possessing the supernatural power of seventy-two transformations and a magical golden cudgel. He is witty and brave, once causing havoc in the Heavenly Palace, later becoming Tang Sanzang's disciple, accompanying him on his journey to the West to obtain Buddhist scriptures, enduring countless hardships and dangers, and ultimately becoming the Victorious Fighting Buddha.

[0197] Example 2's character description is: On the surface, she is gentle and quiet, but in reality, she is intelligent, cunning, calm, rational, and strong-willed. She is not only multi-talented but also very good at palace intrigue.

[0198] 3. Character Enhancement Module: This module uses search engines or film and television knowledge query tools to search for character information, organizes the search results using a large language model, and supplements the character's personality, background, and language habits to further enrich the character setting.

[0199] For example, the classic character "Sun Wukong" can be given the catchphrase "I am Old Sun", and classic lines such as "I am Old Sun here~", "Monster, where do you think you're going!" and "Heaven and earth are not perfect, and scriptures are also incomplete, which cannot be done by human power."

[0200] For example, the final character design for "Sun Wukong" was:

[0201] {

[0202] Name: Sun Wukong

[0203] Description: "Possessing the supernatural power of seventy-two transformations and a magical golden cudgel, he is witty and brave. He once caused havoc in the Heavenly Palace, later becoming Tang Sanzang's disciple, accompanying him on his journey to the West to obtain Buddhist scriptures. After overcoming numerous hardships and dangers, he ultimately became the Victorious Fighting Buddha."

[0204] Character catchphrases: "Here comes Old Sun!", "Here comes Old Sun!", "Monster, where do you think you're going!", "Heaven and earth are incomplete, and scriptures are also incomplete; these cannot be fixed by human power."

[0205] Dialogue Character: "wukong"

[0206] }

[0207] 4. Stylized Reply Module: Combines specific viewing scenarios, such as character dialogues and plot progression, to generate reply content for interactive dialogue with the audience. It also combines character setting information and dialogue history to determine the tone of the character's reply.

[0208] 5. IP Voice Response: Select the appropriate voice based on the specified speaker ID and generate a TTS response with the specified tone.

[0209] During implementation, such as Figure 9 As shown in the sequence diagram of the interaction process, the agent first identifies the characters in the film through the voiceprint recognition module, extracts their voiceprint features, and performs voice cloning. Then, it uses a large language model to set the character information and generates corresponding dialogue content and response tone based on the character's characteristics and the dialogue context through the stylized response module. Next, the IP voice timbre broadcasting module generates response content as an interactive dialogue with the audience based on the response content, response tone, and specified voice timbre. Finally, the result is returned to the main agent for broadcasting.

[0210] In one specific application, the interaction process between the user and the display device is as follows:

[0211] Watch the episode of "Sun Wukong Thrice Beats the White Bone Demon" from Journey to the West.

[0212] User: Can you play as Sun Wukong and talk to me?

[0213] System: Of course! I, Old Sun, am right here. Feel free to ask me anything!

[0214] User: How did you know that village girl was the White Bone Demon?

[0215] System: Hehe, I, Old Sun, have fiery eyes! I saw right through that village girl; she was a demon in disguise. Her demonic aura couldn't escape my notice!

[0216] User: Haha, did you get angry when your master said that to you?

[0217] System: Sigh, Master didn't recognize the demon's true form and blamed me, Old Sun. I feel a little wronged. However, I know that protecting Master is my responsibility, so I won't dwell on it and will continue to escort him on his journey to the West to obtain the scriptures.

[0218] User: Will you come back?

[0219] System: Of course! Although I was misunderstood and driven away by my master, I, Old Sun, will eventually return. I will continue to protect my master and complete the mission of obtaining the scriptures.

[0220] In some embodiments, this application also provides an interaction method applied to the aforementioned display device or a server connected to the display device. In this embodiment, such as... Figure 10 As shown, the interaction method includes the following steps:

[0221] Step 1002: In response to an interaction event triggered by the user, determine the audio / video segment indicated by the interaction event and the target character in the audio / video segment.

[0222] Step 1004: Identify the voiceprint features of the target character from the audio and video clips, and obtain a simulated timbre for the target character based on the voiceprint features.

[0223] Step 1006: Based on the interaction content in the interaction event, perform role enhancement on the target role and determine the target role's response tone and response content to the interaction event.

[0224] Step 1008: Generate and output the response voice based on the simulated voice timbre and response tone of the target character.

[0225] The solutions described above have the following advantages or beneficial effects:

[0226] By responding to user-triggered interactive events, the system identifies the audio / video clips indicated by the events and the target characters within those clips. An intelligent agent then identifies the voiceprint features of the target characters from the audio / video clips, obtaining a simulated voice for each character based on these features. Based on the interactive content of the event, the system enhances the target character's role, determining their appropriate tone and content for responding to the event. A response voice is then generated according to the simulated voice and tone, accurately mimicking the character's voice in the video and outputting it to the user, enhancing the realism of the interaction. Furthermore, by using the simulated voice, determined tone, and content, the system ensures a strong correlation between the generated response voice and the target character in the audio / video clips. This allows the system to interact with the user in the image of the target character, increasing user interest and improving the user experience.

[0227] In some embodiments, based on the interaction content in the interaction event, the target role is enhanced to determine the target role's response tone and response content to the interaction event, including:

[0228] If the interactive content in an interactive event contains a first description of the target character, the first description is rewritten or expanded to obtain a second description of the target character. Based on the second description and the interactive content, the target character's tone and content of response to the interactive event are determined. The second description satisfies at least one of the following conditions: its content richness is greater than that of the first description, and its content conciseness is greater than that of the first description.

[0229] In some embodiments, the interaction method further includes: storing simulated timbres constructed by the agent for different roles;

[0230] The process involves identifying the voiceprint features of a target character from audio and video clips using an intelligent agent, and obtaining a simulated timbre for the target character based on the voiceprint features. This includes: performing voiceprint recognition on audio and video clips using an intelligent agent to obtain the voiceprint features of the target character; and searching for the simulated timbre of the target character in memory according to the voiceprint features.

[0231] In some embodiments, the interaction method further includes: cloning the voice of the target character according to voiceprint features to construct a simulated voice for the target character when the simulated voice of the target character does not exist in the memory; and storing the simulated voice of the target character in the memory.

[0232] In some embodiments, based on the interaction content in the interaction event, the target role is enhanced to determine the target role's response tone and response content to the interaction event, including:

[0233] If the interactive content in an interactive event contains a first description of the target character, the first description is rewritten or expanded to obtain a second description of the target character. Based on the second description and the interactive content, the target character's tone and content of response to the interactive event are determined. The second description satisfies at least one of the following conditions: its content richness is greater than that of the first description, and its content conciseness is greater than that of the first description.

[0234] In some embodiments, when the content richness of the first description content is less than a preset richness condition, obtaining the second description content for the target character includes: calling a search engine or film and television knowledge query tool to search for supplementary information about the target character; and using a large language model to integrate the first description content and the supplementary information about the character to obtain the second description content for the target character.

[0235] In some embodiments, supplementary character information includes at least one of the target character's personality, background, and language habits.

[0236] In some embodiments, the video clip is the target video clip currently playing on the display; the interaction event is an interaction event triggered by the user while watching the target video clip;

[0237] Based on the interactive content in the interactive event, the target character is enhanced to determine the target character's response tone and content to the interactive event. This includes: enhancing the target character based on the interactive content in the interactive event and the film to which the target video clip belongs, obtaining descriptive content for the target character; determining the target character's response tone to the interactive content according to the descriptive content; and generating response content to the interactive content based on the plot and descriptive content depicted in the target video clip.

[0238] In some embodiments, the agent is a single agent among multiple agents configured for the display device;

[0239] The system generates a response voice based on the target character's simulated voice timbre and tone of voice, including: and controls the audio output module to play the response voice. During this process, the controller is further configured to:

[0240] A single agent generates a response voice based on the simulated voice tone and tone of the target character. The single agent then feeds the response voice back to the multi-agent system, which in turn outputs the response voice.

[0241] In some embodiments, a display device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method steps.

[0242] In some embodiments, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above method steps.

[0243] In some embodiments, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method steps.

[0244] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0245] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0246] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A display device, characterized in that, The display device includes: The display is configured to show content from a broadcast system or network and / or a user interface, as well as to display a user interface; The audio output module is configured to play audio. The controller is configured as follows: In response to a user-triggered interaction event, the system determines the audio / video segment indicated by the interaction event and the target character within the audio / video segment; the audio / video segment is content that the display supports playing. The intelligent agent identifies the voiceprint features of the target character from the audio and video clips, and obtains a simulated timbre for the target character based on the voiceprint features; Based on the interaction content in the interaction event, the target character is enhanced to determine the target character's response tone and response content to the interaction event; The intelligent agent is controlled to generate a response voice based on the simulated timbre of the target character and the response tone, and the audio output module is controlled to broadcast the response voice.

2. The display device according to claim 1, characterized in that, The display device also includes a memory; The memory is configured to store the simulated timbres constructed by the agent for different roles; In the process of identifying the voiceprint features of the target character from the audio and video clips through an intelligent agent, and obtaining a simulated timbre for the target character based on the voiceprint features, the controller is further configured to: The voiceprint features of the target character are obtained by performing voiceprint recognition on the audio and video clips using an intelligent agent. Based on the voiceprint characteristics, the simulated voice of the target character is retrieved from the memory.

3. The display device according to claim 2, characterized in that, The controller is further configured to: If the simulated timbre of the target character does not exist in the memory, the timbre of the target character is cloned according to the voiceprint characteristics to construct a simulated timbre for the target character; The simulated voice of the target character is stored in the memory.

4. The display device according to claim 1, characterized in that, In the process of performing role enhancement on the target role based on the interaction content in the interaction event, and determining the target role's response tone and response content to the interaction event, the controller is further configured to: If the interactive content in the interactive event contains a first description of the target character, the first description is rewritten or expanded to obtain a second description of the target character. Based on the second description and the interaction content, determine the target character's tone and content of response to the interaction event; The second description content satisfies at least one of the following conditions: its content richness is greater than that of the first description content, and its content simplicity is greater than that of the first description content.

5. The display device according to claim 4, characterized in that, When the content richness of the first description is less than a preset richness condition, during the process of acquiring the second description for the target role, the controller is further configured to: Use a search engine or film and television knowledge query tool to search for supplementary information about the target character. The first descriptive content and the supplementary information of the role are integrated using a large language model to obtain a second descriptive content for the target role.

6. The display device according to claim 5, characterized in that, The supplementary information about the character includes at least one of the target character's personality, background, and language habits.

7. The display device according to claim 1, characterized in that, The video clip is the target video clip currently playing on the display; the interaction event is an interaction event triggered by the user while watching the target video clip; In the process of performing role enhancement on the target role based on the interaction content in the interaction event, and determining the target role's response tone and response content to the interaction event, the controller is further configured to: Based on the interactive content in the interactive event and the movie to which the target video segment belongs, the target character is enhanced to obtain a description of the target character; Based on the description of the target character, determine the tone of the response of the target character to the interactive content; Based on the storyline depicted in the target video clip and the described content, a response is generated for the interactive content.

8. The display device according to claim 1, characterized in that, The intelligent agent is a single intelligent agent among the multiple intelligent agents configured in the display device; In the process of controlling the intelligent agent to generate a response voice based on the simulated timbre of the target role and the response tone, and controlling the audio output module to play the response voice, the controller is further configured to: The single agent is controlled to generate a response voice based on the simulated timbre of the target character and the response tone, and the single agent feeds back the response voice to the multi-agent; The multi-agent system is controlled to broadcast the response voice through the audio output module.

9. An interaction method, characterized in that, The method includes: In response to a user-triggered interaction event, determine the audio / video segment indicated by the interaction event and the target character in the audio / video segment; Identify the voiceprint features of the target character from the audio and video clips, and obtain a simulated timbre for the target character based on the voiceprint features; Based on the interaction content in the interaction event, the target character is enhanced to determine the target character's response tone and response content to the interaction event; Based on the simulated voice of the target character and the tone of the reply, generate and output a reply voice for the reply content.

10. The method according to claim 9, characterized in that, The step of enhancing the target character based on the interaction content in the interaction event, and determining the target character's response tone and response content to the interaction event, includes: If the interactive content in the interactive event contains a first description of the target character, the first description is rewritten or expanded to obtain a second description of the target character. Based on the second description and the interaction content, determine the target character's tone and content of response to the interaction event; The second description content satisfies at least one of the following conditions: its content richness is greater than that of the first description content, and its content simplicity is greater than that of the first description content.