Television apparatus and conversation processing method

By integrating voice acquisition, voiceprint recognition, and generative models into television devices, and combining them with a digital human-driven system, the problem of poor foreign language dialogue quality on television devices has been solved, enabling high-quality foreign language dialogue practice and improving users' foreign language learning experience.

CN119905094BActive Publication Date: 2025-11-04HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411919258.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-11-04
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing television equipment cannot provide high-quality dialogue content when engaging in foreign language conversations with users, which affects the user's foreign language learning outcomes.

Method used

By integrating voice acquisition components, voiceprint recognition, network communication devices, and generative models into television equipment, the system can recognize, convert, and generate personalized dialogue content for users' voices. It also utilizes a digital human-driven system for motion synchronization, providing high-quality foreign language dialogue practice.

Benefits of technology

It enables high-quality foreign language dialogue practice between TV devices and users, provides personalized dialogue content, and improves users' foreign language speaking skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119905094B_ABST
    Figure CN119905094B_ABST
Patent Text Reader

Abstract

The application relates to a television device and a conversation processing method. The television device comprises a display component, a remote control signal receiving component, an audio output component, a network communication device, a voice collecting component, and a controller configured to perform voiceprint recognition on current user voice to determine a voiceprint identifier corresponding to the current user and display voice text, extract user feature information from a knowledge graph based on the voiceprint identifier, and obtain dialogue text based on a generative model in a first server based on the user feature information, obtain dialogue voice data, display the dialogue text through a first interface, output the dialogue voice data through the audio output component, and drive a preset digital human image to make corresponding actions through a digital human driving system. The television device and the conversation processing method can provide a non-Chinese voice dialogue function with better quality.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of television equipment, and particularly relates to a television equipment and a conversation processing method. BACKGROUND

[0002] At present, there is a development trend of intelligence for television equipment, and a user can perform dialogue training with television equipment to improve the user's foreign language ability, such as English oral language. However, at present, the dialogue quality provided by the television equipment is poor. SUMMARY

[0003] The present application provides a television equipment and a conversation processing method to improve the quality of conversation between the television equipment and the user.

[0004] In a first aspect, some embodiments provide a television equipment, comprising:

[0005] a display component;

[0006] a remote control signal receiving component;

[0007] an audio output component;

[0008] a network communication device;

[0009] a voice collecting component configured to collect a current user voice when the remote control signal receiving component receives a key signal corresponding to a target key in a remote controller; the target key is a preset key for starting a voice recognition function of the television;

[0010] a controller configured to:

[0011] in a case where the display component displays a first interface and the voice collecting component collects the current user voice, perform voiceprint recognition on the current user voice, determine a voiceprint identifier corresponding to the current user based on a voiceprint recognition result, extract user feature information corresponding to the current user from a preset knowledge graph based on the voiceprint identifier, and convert the current user voice into corresponding voice text, and display the voice text through the first interface; the first interface is a preset interface for voice dialogue with a digital person;

[0012] initiate a first request to a first server through the network communication device; the first request carries a dialogue text request instruction, voice text and user feature information; receive dialogue text returned by the first server; wherein the dialogue text is generated based on the first request and a preset generative model in the first server;

[0013] initiate a second request to a second server through the network communication device; the second request carries a text-to-speech instruction and the dialogue text;

[0014] receive the dialogue voice data returned by the second server; the dialogue voice data is generated based on the second request and a preset text-to-speech service in the second server;

[0015] send the dialogue text to the display component, and send the dialogue voice data to the audio output component and the preset digital human driving system respectively, so as to display the dialogue text through the first interface, output the dialogue voice data through the audio output component, and drive the preset digital human image to make actions corresponding to the dialogue voice data through the digital human driving system.

[0016] In one of the embodiments, the controller is further configured to:

[0017] In the case where the remote control signal receiving component receives the confirmation instruction sent by the remote controller, if the current display focus in the display component is located on the preset dialogue mode option, control the display component to display the first interface; or,

[0018] In the case where the remote control signal receiving component receives the confirmation instruction sent by the remote controller, if the current display focus in the display component is located on the preset guessing picture mode option, control the display component to display the second interface; the second interface is a preset interface for voice guessing picture with the digital human;

[0019] The first interface and the second interface are interfaces included in the target application; the target application is an application including a digital human voice dialogue function.

[0020] In one of the embodiments, the controller is further configured to:

[0021] Based on the historical user voice collected by the voice collection component, analyze the voiceprint information corresponding to each historical user, and determine the voiceprint identifier of each historical user based on the voiceprint information;

[0022] Convert the historical user voice into corresponding historical voice text, and extract entity information corresponding to a pre-defined entity word and the association relationship between different entity information from the historical voice text; the pre-defined entity word includes one or more of name, gender, age, interest, and character relationship;

[0023] According to the voiceprint identifier, the entity information, and the association relationship of at least one historical user, construct a knowledge graph; the nodes in the knowledge graph represent the voiceprint identifier and the entity information, and the edges between the nodes represent the association relationship.

[0024] In one of the embodiments, extracting the entity information corresponding to the pre-defined entity word and the association relationship between different entity information from the historical voice text includes:

[0025] Initiate a third request to the first server through the network communication device; the third request carries the historical voice text and an entity information extraction instruction;

[0026] receive the entity information and the association relationship between different entity information returned by the first server; the entity information and the association relationship are generated based on the third request, a preset entity word set in the first server and a generative model.

[0027] In one of the embodiments, the controller is further configured to:

[0028] After initiating the first request to the first server through the network communication device, receive the entity information contained in the speech text of the current user's speech and the association relationship between different entity information returned by the first server based on the first request, and update the knowledge graph.

[0029] In one of the embodiments, the user feature information corresponding to the current user is extracted from the preset knowledge graph based on the voiceprint identification, including:

[0030] Determine the node corresponding to the voiceprint identification from the preset knowledge graph as a target node;

[0031] Determine at least one associated node of the target node, and there is an edge between the associated node and the target node;

[0032] Obtain the entity information represented by the associated node and the association relationship between the associated node and the target node as the user feature information corresponding to the current user.

[0033] In one of the embodiments, the controller is further configured to:

[0034] In the case that the display component switches from other interfaces to display the second interface, extract a preset number of pictures from the preset picture database and cache; the preset number is an integer greater than 1;

[0035] Determine one picture from the cached pictures as a first target picture, and initiate a fourth request to the first server through the network communication device; the fourth request carries an opening text request instruction and the first target picture;

[0036] Receive the opening text returned by the first server; the opening text is generated based on the fourth request and a preset generative model in the first server;

[0037] Initiate a fifth request to the second server through the network communication device; the fifth request carries a text-to-speech instruction and the opening text;

[0038] Receive the opening speech data returned by the second server; the opening speech data is generated based on the fifth request and a preset text-to-speech service in the second server;

[0039] send the first target picture and the opening text to the display component, send the opening voice data to the audio output component and the digital human driving system respectively, so as to display the first target picture and the opening text through the second interface, output the opening voice data through the audio output component, and drive the preset digital human image to make actions corresponding to the dialogue voice data through the digital human driving system;

[0040] The opening text and the opening voice data correspond to the same non-Chinese language.

[0041] In one of the embodiments, the controller is further configured to:

[0042] In the case that the display component displays the second interface and the voice collection component collects the current user voice, convert the current user voice into corresponding picture guessing answer text, and display the picture guessing answer text through the second interface;

[0043] initiate a sixth request through the network communication device to the first server; the sixth request carries a result determination request instruction, the picture guessing answer text and the first target picture;

[0044] receive the determination result text returned by the first server; the determination result text is generated based on the sixth request and a preset generative model in the first server;

[0045] initiate a seventh request through the network communication device to the second server; the seventh request carries a text-to-speech instruction and the determination result text;

[0046] receive the picture guessing result voice data returned by the second server; the picture guessing result voice data is generated based on the seventh request and a preset text-to-speech service in the second server;

[0047] determine another picture from the cached pictures as a second target picture;

[0048] obtain a preset prompt text and corresponding voice data thereof; the prompt text and the corresponding voice data thereof are used to prompt the user to continue guessing the picture;

[0049] send the picture guessing result text, the second target picture and the prompt text to the display component, send the combination data of the picture guessing result voice data and the voice data corresponding to the prompt text to the audio output component and the digital human driving system respectively, so as to display the picture guessing result text, the second target picture and the prompt text through the second interface, output the combination data through the audio output component, and drive the preset digital human image to make actions corresponding to the opening voice data through the digital human driving system;

[0050] The picture guessing result text and the picture guessing result voice data correspond to the non-Chinese language.

[0051] In one of the embodiments, the first server and the second server are servers deployed in the cloud, and the first server and the second server are the same server or different servers; the preset generative model in the first server is a model based on a Transformer architecture and having a text-to-text function.

[0052] The television device has the following effects:

[0053] In the case where the television device enters a function page related to digital foreign language training, when the user operates a target key in the remote controller, it indicates that the voice recognition function of the television device needs to be started to further have a foreign language oral conversation with the digital person. In this regard, the television device collects the current user voice through the voice collection component, and converts the current user voice into voice text through the controller and displays it on the first interface, so that the user knows the content of the user's speech captured by the television device. At the same time, the controller performs voiceprint recognition to determine the corresponding voiceprint identifier, and extracts the corresponding user feature information from the knowledge graph according to the voiceprint identifier. The controller of the television device sends a first request to the first server to obtain a conversation text for the user's speech content. Since the conversation text is generated based on the first request and the preset generative model in the first server, and the first request carries the voice text and the user feature information, a conversation text with better quality can be obtained. The controller of the television device sends a second request to the second server to obtain conversation voice data. The obtained conversation voice data is consistent with the content of the conversation text generated by the generative model. Finally, the display component of the television device displays the conversation text, the audio output component outputs the conversation voice data, and the digital person driving system drives the digital person to make corresponding actions. Thus, based on the cooperation of the controller of the television device, the generative model in the first server, and the text-to-speech service in the second server, the user can have a foreign language conversation practice with the digital person. Since the conversation text of the digital person for the user's speech content is generated based on the generative model and the user feature information, the user can be provided with personalized conversation content, and the quality of the conversation text is better. At the same time, through the text-to-speech service and the digital person driving system, the content of the conversation text, the action of the digital person, and the voice of the digital person are consistent, so that a better foreign language oral conversation experience can be provided.

[0054] In a second aspect, some embodiments further provide a conversation processing method applied to the controller provided in the first aspect, and the method comprises:

[0055] In a case where the display component of the television device displays a first interface and the voice collection component of the television device collects a current user voice, performing voiceprint recognition on the current user voice, determining a voiceprint identifier corresponding to the current user based on a voiceprint recognition result, extracting user feature information corresponding to the current user from a preset knowledge graph based on the voiceprint identifier, and converting the current user voice into corresponding voice text, and displaying the voice text through the first interface; the first interface is a preset interface for voice conversation with the digital human;

[0056] initiating a first request to the first server through the network communication device; the first request carries a dialogue text request instruction, voice text and user feature information; receiving dialogue text returned by the first server; wherein the dialogue text is generated based on the first request and a preset generative model in the first server;

[0057] initiating a second request to the second server through the network communication device; the second request carries a text-to-speech instruction and the dialogue text;

[0058] receiving dialogue voice data returned by the second server; the dialogue voice data is generated based on the second request and a preset text-to-speech service in the second server;

[0059] sending the dialogue text to the display component and sending the dialogue voice data to the audio output component of the television device and the preset digital human driving system respectively, so as to display the dialogue text through the first interface, output the dialogue voice data through the audio output component, and drive the preset digital human image to make corresponding actions through the digital human driving system.

[0060] The above-mentioned electric conversation processing method has the following effects:

[0061] In a case where the television device enters a function page related to digital oral language training, when a user operates a target key in a remote controller, it indicates that the voice recognition function of the television device needs to be started to further have an oral language conversation with the digital person. In this case, a controller in the television device converts the current user voice into voice text and displays it on a first interface, so that the user knows the content of the user's speech captured by the television device. At the same time, the controller performs voiceprint recognition to determine the corresponding voiceprint identifier, and extracts the corresponding user feature information from the knowledge graph according to the voiceprint identifier. The controller sends a first request to a first server through a network communication device to obtain a conversation text for the user's speech content. Since the conversation text is generated based on the first request and a preset generative model in the first server, and the first request carries the voice text and the user feature information, a better quality conversation text can be obtained. The controller sends a second request to a second server through the network communication device to obtain conversation voice data. The obtained conversation voice data is consistent with the content of the conversation text generated by the generative model. Finally, the controller sends the above conversation text to a display component of the television device to display the conversation text, sends the conversation voice data to an audio output component to output the conversation voice data, and sends the conversation voice data to a digital person driving system to drive the digital person to make corresponding actions. Thus, based on the cooperation of the controller of the television device, the generative model in the first server, and the text-to-speech service in the second server, the user can have an oral language conversation practice with the digital person. Since the conversation text of the digital person for the user's speech content is generated based on the generative model and the user feature information, the user can be provided with personalized conversation content, and the quality of the conversation text is better. At the same time, through the text-to-speech service and the digital person driving system, the content of the conversation text, the action of the digital person, and the voice of the digital person are consistent, so that a better oral language conversation experience can be provided. BRIEF DESCRIPTION OF DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor.

[0063] Figure 1 The schematic diagram of the operation scene between the display device and the control device provided by some embodiments of the present application is shown.

[0064] Figure 2 The hardware configuration schematic diagram of the display device provided by some embodiments of the present application is shown.

[0065] Figure 3Hardware configuration schematic diagram of the control device provided for some embodiments of the present application;

[0066] Figure 4 Software configuration schematic diagram of the display device provided for some embodiments of the present application;

[0067] Figure 5 Flowchart schematic diagram of the controller configured to be executed in the television device provided for some embodiments of the present application;

[0068] Figure 6 Another flowchart schematic diagram of the controller configured to be executed in the television device provided for some embodiments of the present application;

[0069] Figure 7 Timing diagram of the non-Chinese voice dialogue practice provided for some embodiments of the present application;

[0070] Figure 8 Statistical data diagram of the non-Chinese voice dialogue practice provided for some embodiments of the present application;

[0071] Figure 9 Timing diagram of the voice guessing picture provided for some embodiments of the present application;

[0072] Figure 10 Interface display diagram of the voice guessing picture provided for some embodiments of the present application;

[0073] Figure 11 First schematic diagram of the knowledge graph provided for some embodiments of the present application;

[0074] Figure 12 Second schematic diagram of the knowledge graph provided for some embodiments of the present application;

[0075] Figure 13 Third schematic diagram of the knowledge graph provided for some embodiments of the present application;

[0076] Figure 14 Fourth schematic diagram of the knowledge graph provided for some embodiments of the present application. DETAILED DESCRIPTION

[0077] The embodiments will be described in detail below with reference to the drawings. When the following description refers to arrangements in the drawings, identical numbers on different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following embodiments are not meant to represent all implementations consistent with the present application. Rather, they are merely examples of implementations consistent with some aspects of the present application as detailed in the appended claims.

[0078] It should be noted that the brief description of the terms in this application is only for the convenience of understanding the implementation described next, and is not intended to limit the implementation of the application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.

[0079] The terms "first", "second", "third" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchanged under appropriate circumstances. Unless otherwise specified, the plurality referred to herein can be understood as two or more. Unless otherwise specified, the id and ID in this document can be the same.

[0080] The terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device including a series of components does not necessarily limit to all components clearly listed, but can include other components not clearly listed or inherent to these products or devices.

[0081] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware or / and software code capable of performing functions related to the element.

[0082] In the embodiments of the present application, the display device 200 generally refers to a device with picture display and data processing capability. For example, the display device 200 includes but is not limited to smart TV, mobile terminal, computer, monitor, advertising screen, wearable device, virtual reality device, augmented reality device, etc.

[0083] Figure 1 The display device provided for some embodiments of the present application and the control device between the operation scene diagram. As Figure 1 As shown in the middle, the user can operate the display device 200 through touch operation, mobile terminal 300 and control device 100. For example, the control device 100 can be a remote controller, a touch pen, a handle, etc.

[0084] The mobile terminal 300 can be used as a kind of control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a kind of communication device for establishing communication connection with the display device 200 and performing data interaction. In some embodiments, the mobile terminal 300 can install software application with the display device 200, realize connection communication through network communication protocol, realize one-to-one control operation and data communication purpose. The mobile terminal 300 can also display audio and video content on the display device 200, realize synchronous display function.

[0085] As Figure 1It is also shown that the display device 200 also communicates data with the server 400 through various communication modes. The display device 200 can be allowed to communicate through a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0086] The display device 200 can provide a broadcast receiving television function, and can also provide an intelligent network television function of computer support function, including but not limited to, network television, smart television, Internet protocol television (IPTV), etc.

[0087] Figure 2 The display device 200 provided in some embodiments of the present application Figure 1 The hardware configuration block diagram of the display device 200 is shown in the following.

[0088] In some embodiments, the display device 200 can include at least one of a tuning demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0089] In some embodiments, the detector 230 is used to collect signals of external environment or external interaction. For example, the detector 230 includes a light receiver for collecting ambient light intensity; or the detector 230 includes an image collector such as a camera, which can be used to collect external environment scenes, user attributes or user interaction gestures; or the detector 230 includes a sound collector such as a microphone, etc., for receiving external sound.

[0090] In some embodiments, the display 260 includes a display function component for presenting a picture, and a driving component for driving image display. The display 260 is used to receive image signals output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, and components of a menu control interface, as well as user control UI interface, etc.

[0091] In some embodiments, the communication device 220 is a component for communicating with external devices or the server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 according to different supported communication modes. For example, when the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 containing WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 containing Bluetooth function.

[0092] The communication device 220 enables the display device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the display device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The display device 200 can directly establish a connection with external devices or indirectly through gateways, routers, or other connection devices.

[0093] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.

[0094] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0095] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface receives user input commands through the graphical user interface (GUI).

[0096] In some embodiments, the audio output device 270 can be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the display device 200 to output sound from the display device 200.

[0097] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0098] Figure 3 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of the central control device. (Example) Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0099] The control device 100 is configured to control the display device 200, and can receive the input operation instruction of the user, and convert the operation instruction into an instruction that the display device 200 can recognize and respond to, and play a role of an intermediary in the interaction between the user and the display device 200.

[0100] In some embodiments, the control device 100 can be a smart device. For example, the control device 100 can install various applications for controlling the display device 200 according to the user's needs.

[0101] In some embodiments, as shown in FIG. 3, the mobile terminal 300 or other smart electronic device can play a similar function of the control device 100 after installing the application for controlling the display device 200. Figure 1

[0102] The controller 110 includes a processor 112 and a RAM 113 and a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, and the communication and cooperation between the internal components, and the data processing function of the external and internal.

[0103] The communication interface 130 realizes the communication of the control signal and the data signal between the display device 200 under the control of the controller 110. The communication interface 130 can include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC module 133 and other near field communication modules.

[0104] The user input / output interface 140, wherein the input interface includes at least one of a microphone 141, a touchpad 142, a sensor 143, a key 144 and other input interfaces.

[0105] In some embodiments, the control device 100 includes at least one of the communication interface 130 and the input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC module, which can encode the user input instruction through a WiFi protocol, or a Bluetooth protocol, or an NFC protocol, and send it to the display device 200.

[0106] The storage 190 is used to store the driving and control of the controller of the control device 100 various running programs, data and applications. The storage 190 can store various control signal instructions input by the user.

[0107] The power supply 180 is used to provide operating power support for the elements of the control device 100 under the control of the controller.

[0108] ​To perform user interactions, in some embodiments, the display device 200 can run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system can provide a user interface to allow a user to interact with the display device 200 and support running various application programs.

[0109] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.

[0110] The operating system can be divided into different modules or levels according to the implemented functions, for example, as shown in FIG. 2, in some embodiments, the system is divided into four layers, from top to bottom, the Applications layer (referred to as the “application layer”), the Application Framework layer (referred to as the “framework layer”), the system library layer, and the kernel layer. Figure 4

[0111] In some embodiments, the application layer is used to provide services and interfaces for application programs, so that the display device 200 can run application programs and interact with users based on the application programs. At least one application program can run in the application layer, which can be a window (Window) program, a system setting program, or a clock program provided by the operating system, or an application program developed by a third-party developer. In specific implementation, the application programs in the application layer are not limited to the above examples.

[0112] The framework layer provides application programming interfaces (APIs) and programming frameworks for application programs. The application framework layer includes some pre-defined functions. The application framework layer is equivalent to a processing center that decides which application program in the application layer to act. The application program can access resources in the system and obtain services of the system through the API interface in the execution.

[0113] As shown in FIG. 2, in some embodiments, the system library layer includes a number of libraries, such as a media library, a graphics library, a database library, a network library, and a location library. Figure 4 ​As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, textboxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0114] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0115] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0116] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 4 As shown, hardware drivers can be configured in the kernel layer. The drivers included in the kernel layer can be at least one of the following: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0117] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the display device 200 in this application embodiment. Depending on the function of the display device, the type of operating system, and other factors, the number of levels and the specific level type of the operating system may be expressed in other forms.

[0118] In one exemplary embodiment, a television device is provided, which allows users to interact with the device via remote control, voice, or other means. The television device can provide dialogue training to improve the user's foreign language skills. The television device includes: a display component, a remote control signal receiving component, an audio output component, a network communication device, a voice acquisition component, a controller, etc. The television device can be the display device described above. For an understanding of the television device and its components, please refer to the relevant descriptions of display devices above. For example, for an understanding of the display component, please refer to the display component in the previous section on display devices; for an understanding of the network communication device, please refer to the communication device in the previous section on display devices; for an understanding of the remote control, please refer to the description of control devices above. The same content will not be repeated here. The following mainly describes the specific technical solutions related to how the television device provides users with better dialogue quality:

[0119] Regarding the voice acquisition component in the television device: it is used to acquire the current user's voice when the remote control signal receiving component receives the button signal corresponding to the target button on the remote control; the target button is a preset button used to activate the television's voice recognition function.

[0120] The target button can be a physical button or a virtual button. It can be a dedicated button or a generic button, such as a "Confirm" or "OK" button, as long as the user can activate the TV's voice recognition function by pressing the target button.

[0121] In some possible embodiments, the voice acquisition component can continuously acquire the current user's voice by pressing the target button once, until the user presses the target button again, at which point the voice acquisition component will stop acquiring the current user's voice.

[0122] In some possible embodiments, the voice acquisition component can acquire the current user's voice when the user presses and holds the target button and speaks, and the voice acquisition component stops acquiring the current user's voice when the user stops pressing and releasing the target button.

[0123] like Figure 5 As shown, the controller in the television device is configured to execute steps S101 to S106, specifically as follows:

[0124] Step S101: In the case that the display component displays a first interface and the voice collection component collects a current user voice, performing voiceprint recognition on the current user voice, determining a voiceprint identifier corresponding to the current user based on the voiceprint recognition result, extracting user feature information corresponding to the current user from a preset knowledge graph based on the voiceprint identifier, and converting the current user voice into corresponding voice text, and displaying the voice text through the first interface; the first interface is a preset interface for voice dialogue with the digital human.

[0125] The voiceprint identifier can be obtained based on voiceprint recognition on the collected user language. Different users have different voiceprints, and thus the voiceprint identifiers corresponding to different users are different.

[0126] In some possible embodiments, to provide foreign language spoken dialogue service to the user, the television device can be pre-provisioned with a foreign language spoken dialogue function, in which the user can have a voice dialogue with the digital human in the foreign language. The function can be started or entered in various ways, for example, a shortcut key can be provided on the remote controller, and when the user presses the shortcut key, the function is entered or started. For another example, an interface option for starting the function can be provided on the startup interface or the initial default interface of the television device, for example, when the user selects the interface option through the remote controller and confirms the selection, the function is entered or started. The function can have a corresponding interface, and the television device displays the corresponding interface to the user.

[0127] The first interface can be an interface corresponding to the function of providing foreign language spoken dialogue training or voice dialogue to the user by the television device, and the content of the dialogue between the user and the digital human can be displayed in the interface. In some possible embodiments, an image of the digital human is also displayed in the first interface. The image of the digital human can be determined based on a selection operation of the user, or can be determined based on related configuration information, or can be determined by the large model from a plurality of pre-set images suitable for the current scene.

[0128] In some possible embodiments, all the text or characters displayed in the first interface are the same non-Chinese language, for example, English.

[0129] In some possible embodiments, converting the current user voice into corresponding language text can be performed locally on the television device, or can be performed by using a server (in the cloud).

[0130] Step S102: Initiating a first request to a first server through a network communication device; the first request carries a dialogue text request instruction, voice text, and user feature information.

[0131] The knowledge graph can be pre-constructed information related to user voiceprint identification and user characteristics, such as the user's name, hobbies, family relationships, and the like.

[0132] The dialogue text request instruction is an instruction for instructing the first server to perform a related action, for example, the dialogue text request instruction of the present application, for instructing the first server to generate corresponding dialogue text according to the voice text and the user characteristic information and return.

[0133] The understanding of the first server can refer to the description of the "server 400" above.

[0134] In some possible embodiments, the knowledge graph can be a user characteristic related graph constructed around the voiceprint identification. In the case of determining the voiceprint identification, the corresponding user characteristic information can be extracted according to the determined voiceprint identification.

[0135] In some possible embodiments, the knowledge graph can be updated according to the dialogue record of the user with the television device and the digital person.

[0136] For example, the user characteristic information corresponding to the current user can be extracted from the knowledge graph according to the voiceprint identification, and then the user characteristic information, the voice text, and the corresponding dialogue text request instruction can be packaged and arranged into the first request, and the first request can be initiated to the first server.

[0137] Step S103: receiving the dialogue text returned by the first server; wherein the dialogue text is generated based on the first request and a preset generative model in the first server.

[0138] The dialogue text can be a text for responding to the current user's voice, and also a data basis for the digital person to have a spoken dialogue with the user, thereby forming a dialogue exchange with the user in terms of content.

[0139] The generative model can be a preset artificial intelligence model in the first server, which has the function of generating text from text, and can generate dialogue text according to the first request. For example, the voice text in the first request includes "how are you?", and the user characteristic information includes the user's name "Tom", then the generative model can generate "Hi, Tom, I'm fine, and you?".

[0140] In some possible embodiments, the user characteristic information can be integrated into the prompt information of the generative model, for example, generating a prompt word based on the current user's voice text, user characteristic information, and dialogue text request instruction, to prompt the generative model to generate more accurate dialogue text according to the voice text.

[0141] After the foregoing steps, data for the user can be obtained, that is, the dialogue text, but it is in written form, which is the data basis for the digital person to have a spoken dialogue with the user. In order to complete the spoken dialogue with the user, audio data of the digital person is also needed, that is, a dialogue form that provides both text and audio. For this, the following steps can be taken.

[0142] Step S104: initiating a second request to the second server through the network communication device; the second request carries a text-to-speech instruction and the dialogue text.

[0143] The text-to-speech instruction can be used to indicate the specific request to the second server, that is, to request conversion of the dialogue text into speech and return.

[0144] Exemplarily, the received dialogue text and text-to-speech instruction can be packaged into the second request, and the second request can be initiated to the second server.

[0145] Step S105: receiving dialogue speech data returned by the second server; the dialogue speech data is generated based on the second request and a preset text-to-speech service in the second server.

[0146] The text-to-speech service can be a service preset in the second server to convert text into speech.

[0147] The understanding of the second server can also be referred to the description of the "server 400" above.

[0148] In some possible embodiments, the second server can be the same server as the first server, which can provide both text-to-speech service and dialogue text generation service. Exemplarily, the server can be preset with a large model, which has both text-to-speech function and dialogue text generation function.

[0149] Step S106: sending the dialogue text to the display component, and sending the dialogue speech data to the audio output component and the preset digital person driving system, respectively, to display the dialogue text through the first interface, output the dialogue speech data through the audio output component, and drive the preset digital person image to make actions corresponding to the dialogue speech data through the digital person driving system.

[0150] The digital person can be understood as a virtual character image displayed in the television device. Correspondingly, the digital person driving system can be a system for driving the digital person, for example, driving the digital person to make actions. The "action" can be understood in a broad sense, which can be a large-amplitude action such as waving hands or bowing, or a micro-action such as blinking eyes, tilting head, smiling, crying, yawning, etc. It can be a body action or a facial action, etc.

[0151] In some possible embodiments, the user voice, the voice text, the dialogue text and the dialogue voice data can correspond to the same non-Chinese language, for example, English. In some possible embodiments, on the basis of the same non-Chinese language, corresponding translation content can also be provided, for example, while the dialogue text in English is displayed, the corresponding Chinese translation text is also displayed.

[0152] In some possible embodiments, after the dialogue text and the dialogue voice data are obtained through the foregoing steps, the dialogue text can be displayed by the display component for the user to read, the corresponding audio can be played by the audio output component, that is, the dialogue voice data is output, and the digital person is driven to make corresponding changes.

[0153] In the embodiment, when the television device enters the digital person foreign language oral training function page, if the user operates the target key of the remote controller, it indicates that the voice recognition function of the television device needs to be started and then the digital person needs to be engaged in foreign language oral dialogue. In this case, the television device collects the current user voice through the voice collection component, and converts the current user voice into voice text through the controller and displays the voice text on the first interface, so that the user knows the content of the user's speech captured by the television device. At the same time, the controller performs voiceprint recognition to determine the corresponding voiceprint identifier, and extracts the corresponding user feature information from the knowledge graph according to the voiceprint identifier. The controller of the television device sends a first request to the first server to obtain the dialogue text for the user's speech content. Since the dialogue text is generated based on the first request and the preset generative model in the first server, and the first request carries the voice text and the user feature information, the dialogue text with better quality can be obtained. The controller of the television device sends a second request to the second server to obtain dialogue voice data. The obtained dialogue voice data is consistent with the content of the dialogue text generated by the generative model. Finally, the dialogue text is displayed through the display component of the television device, the dialogue voice data is output through the audio output component, and the digital person is driven to make corresponding actions through the digital person driving system. Thus, based on the cooperation of the controller of the television device, the generative model in the first server and the text-to-speech service in the second server, the user can practice foreign language oral dialogue with the digital person. Since the dialogue text of the digital person for the user's speech content is generated based on the generative model and the user feature information, the user can be provided with personalized dialogue content, and the quality of the dialogue text is better. At the same time, through the text-to-speech service and the digital person driving system, the content of the dialogue text, the action of the digital person and the voice of the digital person are consistent, so that a better foreign language oral dialogue experience can be provided.

[0154] In one of the foregoing embodiments, the controller is further configured to: in a case where the remote control signal receiving component receives a confirmation instruction issued by the remote controller, if the current display focus in the display component is located on the preset dialogue mode option, control the display component to display a first interface; or in a case where the remote control signal receiving component receives a confirmation instruction issued by the remote controller, if the current display focus in the display component is located on the preset guessing picture mode option, control the display component to display a second interface; the second interface is a preset interface for voice guessing picture with the digital person.

[0155] The first interface and the second interface are interfaces included in the target application; the target application is an application including a digital person voice dialogue function.

[0156] The confirmation instruction issued by the remote controller can be an instruction issued by the user through the remote controller, and the corresponding remote control operation is not limited, for example, the user can click the "OK / confirmation" button in the remote controller, of course, it can also be other buttons, the button can be physical, virtual, etc., which are not limited here.

[0157] For example, when the user presses the confirmation button in the remote controller, the remote controller issues a confirmation instruction.

[0158] As described above, the digital person dialogue function can be pre-installed in the television device, and the interface of the digital person dialogue function includes a mode option, which specifically includes a dialogue mode option and a guessing picture mode option. The dialogue mode option is used to enter an interface for voice chatting with the digital person, and the guessing picture mode option is used to enter an interface for guessing picture interaction with the digital person, and the interaction process is used to practice oral English.

[0159] The guessing picture mode will be described below. For example, the television device can display a picture, text description and / or play related audio to guide the user to answer the content of the picture through voice, and give feedback on the correctness of the user's answer. The process uses the same foreign language, thereby helping the user to improve oral English ability and vocabulary in the process.

[0160] In some possible embodiments, the dialogue mode option and the guessing picture mode option can be displayed simultaneously in the same interface of the television device, or only the dialogue mode option can be displayed in the voice dialogue mode, and only the guessing picture mode option can be displayed in the guessing picture mode. The user selects the currently displayed mode option to pop up another option for the user to select.

[0161] The current display focus can be a position coordinate or a position coordinate range of the display component. For example, the current display focus can be a selection box of the operating system in the television device, and the color of the selection box can be white, yellow, etc. to highlight the selection box; for another example, the current display focus can be a highlighted, enlarged or other highlighted part in the display picture.

[0162] When the current display focus is on the preset dialogue mode option, if a determination instruction of the remote controller is received, the controller in the television device controls the display component to display the first interface to the user.

[0163] In some possible embodiments, the voice dialogue interface corresponding to the dialogue mode option can be provided by a target application, and the voice guessing interface corresponding to the guessing mode option can be provided by the same target application. Accordingly, the first interface can be an interface included by the target application, and the second interface can be an interface included by the same target application; and the target application can be an application including a digital human voice dialogue function.

[0164] In this embodiment, by locating the current display focus in the display component on the preset dialogue mode option or the guessing mode option, and receiving the confirmation instruction of the remote controller by the remote signal receiving component, the display component is controlled to display the corresponding first interface or the second interface, so that the user can select to enter the corresponding interactive mode through the remote controller, and practice spoken language with the digital human through various interactive modes, which is beneficial to enrich the user experience.

[0165] In one of the embodiments, as shown in FIG. 2, the controller in the foregoing embodiment is further configured to perform steps S201 to S204: Figure 6

[0166] Step S201: Based on the historical user voice collected by the voice collecting component, the voiceprint information corresponding to each historical user is analyzed, and the historical user voice is converted into the corresponding historical voice text.

[0167] In this embodiment, the historical user can be a user who uses the television device in the past time; accordingly, the historical user voice can be the user voice collected by the voice collecting component in the past time. For example, the historical user voice can be the user voice collected when the television device enters the first interface or the second interface.

[0168] Step S202: Based on the voiceprint information corresponding to each historical user, the voiceprint identifier of each historical user is determined.

[0169] For example, the voiceprint identifier corresponding to the same user is the same and unique, and the voiceprint identifiers corresponding to different users are different. If there are multiple historical users talking with the digital human, the voiceprint identifier corresponding to each user can be analyzed.

[0170] Step S203: Extracting the entity information corresponding to the pre-defined entity word and the association relationship between different entity information from the historical voice text; the pre-defined entity word includes one or more of the name, the gender, the age, the interest, and the relationship between characters.

[0171] ​Among them, entity information can be names, personal information, hobbies, etc.; correspondingly, association can be a relationship that represents the relationship between different entity information.

[0172] In some possible implementations, artificial intelligence-related technologies, such as large language models, can be used to analyze historical speech texts to obtain entity information and relationships.

[0173] Step S204: Construct a knowledge graph based on the voiceprint identifier, entity information, and association relationships of at least one historical user; the nodes in the knowledge graph represent voiceprint identifiers and entity information, and the edges between nodes represent association relationships.

[0174] In some possible embodiments, the voiceprint identifier itself is also a node. Two nodes representing voiceprint identifiers are not directly connected, but can be indirectly connected through the same or one or more nodes representing entity information.

[0175] like Figure 14 As shown, a possible knowledge graph is provided, in which nodes representing voiceprint identifiers include "Voiceprint id1" and "Voiceprint id2"; nodes representing entity information include "Alice", "Tom", "badminiton", etc.; for the representation of association relationships, for example, the name of "Voiceprint id1" is "Tom", the friend of "Voiceprint id1" is "Alice", that is, Tom's friend is Alice, and the love of "Voiceprint id1" is "badminiton", that is, Tom loves to play badminton, etc.

[0176] In this embodiment, the television device processes user voice recordings collected in the past to determine the user's voiceprint identifier, extracts entity information and relationships from historical voice text, and then constructs a knowledge graph based on the voiceprint identifier, entity information, and relationships. This knowledge graph can form a comprehensive information network based on the voiceprint identifier, including the user's name, gender, interests, relationships, and other aspects. Based on this knowledge graph, more comprehensive, three-dimensional, and accurate user characteristic information can be provided for subsequent steps.

[0177] In one embodiment, the "extracting entity information corresponding to predefined entity words from historical speech text and the association between different entity information" in the aforementioned embodiments may include: initiating a third request to a first server through a network communication device; the third request carrying instructions for extracting historical speech text and entity information; receiving entity information and the association between different entity information returned by the first server; the entity information and the association are generated based on the third request, the pre-set set of entity words in the first server, and a generative model.

[0178] As described above, the entity information and the association relationship can be extracted from the historical voice text. In order to improve the accuracy and efficiency of the extraction, a large model can be used for the extraction. In the present embodiment, as described above, the generative model can be configured in the first server, and thus the extraction of the entity information and the association relationship can be realized by interacting with the first server.

[0179] Exemplarily, the historical voice text and the entity information extraction instruction can be packaged into a third request and sent to the first server. After receiving the third request, the first server parses the third request, calls the built-in generative model to generate the corresponding entity information and association relationship, and returns the entity information and the association relationship to the television device.

[0180] In some possible embodiments, the entity word set can be preset in the first server or other servers. When needed, the first server can directly obtain or obtain from other servers. Of course, the entity word set can also be preset in the television device. When needed, the television device can send the entity set to the first server, for example, the entity word set is also put into the third request. In order to improve the use efficiency of each hardware in the television device and reduce the hardware requirements, the entity word set can be preset in the first server.

[0181] In some possible embodiments, the entity word set can be updated, such as adding new entity words.

[0182] In the present embodiment, by sending a request to the first server, the entity information returned by the first server and the association relationship between different entity information can be obtained. Since the entity information and the association relationship are generated by the entity word set and the generative model preset in the first server, the entity information and the association relationship can more accurately reflect the real situation of each historical user of the television device.

[0183] As described in the foregoing embodiments, the knowledge graph can be constructed based on the historical user voice. In order to ensure its timeliness, the knowledge graph can also be updated based on the voice information of the current user. In one of the embodiments, the controller in the foregoing embodiments is further configured to: after initiating the first request to the first server through the network communication device, receive the entity information contained in the voice text of the current user voice and the association relationship between different entity information returned by the first server based on the first request, and update the knowledge graph.

[0184] In the embodiment, for the current user, in the process of obtaining the dialogue text of the digital person through the first server, the entity information contained in the speech text of the current user's speech and the association relationship between different entity information are returned through the first server, so as to update the knowledge graph. For example, if a new family member is mentioned in the current user's speech, the name of the family member can be updated as a new node in the knowledge graph, and the association relationship between the node and the voiceprint identifier corresponding to the current user is added as family. This enables the knowledge graph to be updated in real time and dynamically, thereby ensuring the accuracy and real-time performance of the knowledge graph.

[0185] As described in the foregoing embodiments, there is a node representing the voiceprint identifier in the knowledge graph, and the voiceprint identifier corresponding to the current user can be extracted through the steps in the foregoing embodiments, so that the node representing the voiceprint identifier can be found in the knowledge graph, and the user feature information corresponding to the current user can be extracted according to the node. In one embodiment, the “extracting the user feature information corresponding to the current user from the preset knowledge graph based on the voiceprint identifier” in the foregoing embodiments can include: determining the node corresponding to the voiceprint identifier from the preset knowledge graph as a target node; determining at least one associated node of the target node, and there is an edge between the associated node and the target node; obtaining the entity information represented by the associated node and the association relationship between the associated node and the target node as the user feature information corresponding to the current user.

[0186] In this embodiment, the understanding of the target node can be combined with the examples of the foregoing embodiments Figure 14 For example, if the node corresponding to the voiceprint identifier is “voiceprint id 1”, then “voiceprint id 1” can be the target node, and correspondingly, “Tom” and the like can be the associated nodes.

[0187] In some possible embodiments, there is an edge between the associated node and the target node, which can correspond to the association relationship between the associated node and the target node. For example, the edge between “voiceprint id 1” and “Tom” corresponds to the association relationship “name”, that is, the name of “voiceprint id 1” is “Tom”.

[0188] In this embodiment, the node corresponding to the voiceprint identifier in the knowledge graph is taken as the target node, and the associated nodes of the target node are determined, so as to obtain the entity information represented by the associated nodes and the association relationship between the associated nodes and the target node, and then obtain the user feature information corresponding to the current user. Since the user feature information is determined based on the voiceprint identifier, confusion of the user feature information between different users is avoided; meanwhile, the user feature information contains the entity information and the association relationship of the target node and the associated nodes, so that the user feature information can reflect more comprehensive information related to the user, thereby helping to provide better non-Chinese dialogue service.

[0189] In the foregoing embodiment, the television device can provide the dialogue service for the user through the first interface. In addition to the dialogue service, the television device can also provide the voice guessing picture service, which can be non-Chinese, thereby helping to provide the training service of foreign language ability for the user.

[0190] In one of the embodiments, the controller in the foregoing embodiment is further configured to perform the following steps:

[0191] Step one: in the case where the display component switches from other interfaces to display the second interface, a preset number of pictures are extracted from the preset picture database and cached; the preset number is an integer greater than 1; a picture is determined from the cached pictures as a first target picture, and a fourth request is initiated to the first server through the network communication device; the fourth request carries a opening text request instruction and the first target picture.

[0192] The preset picture database can be pre-set in the television device, and the controller can extract and cache a preset number of pictures from the preset picture database directly or indirectly.

[0193] The case where the other interfaces are switched to display the second interface indicates that the digital human initiates the first round of dialogue after entering the second interface, based on which, the opening text is the text of the first round, and the opening text request instruction can be an instruction representing the opening text requested from the first server. In the first interface and the second interface of the embodiment of the application, the dialogue of the first round is initiated by the digital human.

[0194] In some possible embodiments, new picture data can be added locally and / or the preset picture database can be updated by interacting with the cloud server.

[0195] In some possible embodiments, a picture can be determined as the first target picture from the cached pictures in a random selection manner.

[0196] In some possible embodiments, to avoid the same picture being displayed to the user through the second interface for multiple times in a short time, when the first target picture is determined, the selected picture can be put back into the preset picture database or no longer participate in subsequent steps; of course, an upper limit of the number of times the same picture can be selected within a preset time range can be set, for example, 1 time in 10 minutes, when the upper limit of the number of times is reached, the picture cannot be selected and enters a waiting period, for example, 10 minutes, after the waiting period ends, the picture can be selected again, and the cycle is repeated.

[0197] Step two: receiving the opening text returned by the first server; the opening text is generated based on the fourth request and a preset generative model in the first server.

[0198] The opening text can be information used to explain or describe to the user that the second interface displayed by the television device is an interface for voice guessing picture with the digital person, so as to guide the user to participate in the voice guessing picture.

[0199] The opening text for voice guessing picture can be obtained through the above steps. In order to improve the interaction quality with the user, the corresponding audio can be added at the same time, and the following steps can be used to achieve this:

[0200] Step three: initiating a fifth request to the second server through the network communication device; the fifth request carries a text-to-speech instruction and the opening text.

[0201] Step four: receiving the opening voice data returned by the second server; the opening voice data is generated based on the fifth request and a preset text-to-speech service in the second server. The opening voice data is the voice data initiated by the digital person in the first round.

[0202] Step five: sending the first target picture and the opening text to the display component, sending the opening voice data to the audio output component and the digital person driving system respectively, so as to display the first target picture and the opening text through the second interface, output the opening voice data through the audio output component, and drive the preset digital person image to make corresponding actions through the digital person driving system; wherein the opening text and the opening voice data correspond to the same non-Chinese language.

[0203] In this embodiment, when the television device enters the voice guessing picture mode, that is, the display component switches from other interfaces to display the second interface, a plurality of pictures are extracted from the preset picture database and cached at one time to determine the first target picture, and an opening text is requested from the first server. The opening text is generated based on a preset generative model, so that a better quality opening text can be obtained. At the same time, the television device requests text-to-speech from the second server, so that the opening speech data corresponding to the opening text returned by the second server can be received. Finally, the opening text is displayed through the display component, the opening speech data is output through the audio output component, and the digital person makes corresponding actions through the digital person driving system, thereby realizing the cooperation of the opening text, digital person actions, and digital person speech to provide foreign language conversation services for the user. Since the contents of the opening text, digital person actions, and digital person speech are consistent, a better quality non-Chinese conversation can be provided.

[0204] In one of the embodiments, the controller of the foregoing embodiment is further configured to perform the following steps:

[0205] Step one: in the case that the display component displays the second interface and the voice collection component collects the current user voice, the current user voice is converted into corresponding guessing picture answer text, and the guessing picture answer text is displayed through the second interface.

[0206] As can be seen from the foregoing embodiments, when the display component displays the second interface, it can be indicated that the television device currently provides related services of voice guessing picture to the user. The television device shows the first target picture, the opening text, and plays the opening speech data, etc. to the user. In response to this, the user can speak out the content in the first target picture. In response to this, the television device can collect the current user voice through the voice collection component, and convert the current user voice into corresponding guessing picture answer text, and display the guessing picture answer text through the second interface.

[0207] When the guessing picture answer text is obtained, the digital person service of the television device can make a corresponding response, and the response content can be obtained by the television device from the server in the cloud. The following steps can be used to achieve this:

[0208] Step two: a sixth request is initiated through the network communication device to the first server; the sixth request carries a result determination request instruction, the guessing picture answer text, and the first target picture.

[0209] Step three: receiving the determination result text returned by the first server; the determination result text is generated based on the sixth request and a preset generative model in the first server.

[0210] The determination result text can be text representing whether the guess or description about the first target picture is accurate. In addition, the description text of subsequent interaction related to the determination result can also be included. For example: guessed correctly, let's go to the next picture guessing.

[0211] Step four: initiating a seventh request to the second server through the network communication device; the seventh request carries the text-to-speech instruction and the determination result text.

[0212] Step five: receiving the picture guessing result voice data returned by the second server; the picture guessing result voice data is generated based on the seventh request and the preset text-to-speech service in the second server.

[0213] Step six: determining another picture from the cached pictures as a second target picture. The second target picture is different from the first target picture.

[0214] For the determination of another picture from the cached pictures, reference can be made to the related technical solutions of the aforementioned determination of a picture from the cached pictures, such as setting an upper limit of the number of selections.

[0215] In some possible embodiments, the determination of another picture from the cached pictures can be synchronous or asynchronous with the aforementioned steps of sending requests to the first server or the second server.

[0216] Step seven: obtaining the preset prompt text and its corresponding voice data; the prompt text and its corresponding voice data are used to prompt the user to continue picture guessing.

[0217] In some possible embodiments, the prompt text and its corresponding voice data can be pre-stored in the television device.

[0218] Step eight: sending the picture guessing result text, the second target picture, and the prompt text to the display component, and sending the combination data of the picture guessing result voice data and the voice data corresponding to the prompt text to the audio output component and the digital human driving system, so as to display the picture guessing result text, the second target picture, and the prompt text through the second interface, output the combination data through the audio output component, and drive the preset digital human image to make actions corresponding to the current voice data through the digital human driving system; wherein the picture guessing result text and the picture guessing result voice data correspond to a non-Chinese language.

[0219] In this embodiment, the second interface is displayed on the display component, and the current user voice is collected. According to the current user voice, a corresponding guessing picture answer text is generated. According to the guessing picture answer text, a judgment result text and guessing picture result voice data corresponding to the judgment result text are requested from the server, and another picture is determined as a second target picture from the cached pictures. A preset prompt text and voice data corresponding to the prompt text are obtained. Finally, the judgment result text, the second target picture and the prompt text are displayed on the display component, the guessing picture result voice data and the voice data corresponding to the prompt text are output through the audio output component, and the digital human makes a corresponding action according to the current voice data through the digital human image. This makes the digital human service of the television device able to make corresponding feedback according to the user's answer when providing voice guessing picture service to the user, such as whether the guessing picture is accurate and guessing the next picture. At the same time, since the judgment result text is generated by the generative model in the first server, the accuracy of the judgment result text is better.

[0220] In one of the embodiments, the first server and the second server are servers deployed in the cloud, and the first server and the second server are the same server or different servers. The preset generative model in the first server is a model based on a Transformer architecture and having a text-to-text function.

[0221] The Transformer architecture can be an architecture of an artificial intelligence model, which can include an input part, an encoder, a decoder and an output part.

[0222] The text-to-text function can be the ability to generate another piece of text from a piece of text. In terms of content, the two pieces of text can correspond and connect. For example, a piece of text is "Hi, I am Tom", and another piece of text generated by the text-to-text function can be "Nice to meet you, Tom, I'm Alice".

[0223] In this embodiment, the first server and the second server are deployed in the cloud, so the television device can interact with the first server and the second server through a network communication device to obtain corresponding data to serve the user to provide foreign language conversation. At the same time, the preset generative model in the first server is a model based on a Transformer architecture and having a text-to-text function, so the television device can obtain data such as better quality conversation text generated based on the model from the first server, thereby providing better quality foreign language conversation for the user.

[0224] In some possible embodiments, the current user voice can be converted into corresponding text locally on the television device, without the need to interact with the server to realize voice-to-text conversion, thereby improving processing efficiency; the text returned by the first server, such as opening text, judgment result text, and dialogue text, can be converted into voice by using the second server, so that more accurate voice is obtained, and the demand of the digital human driving system for corresponding data can also be met.

[0225] In some possible embodiments, when requesting the text from the first server, the corresponding prompt word can be constructed based on the current requested text type, to guide the generative model to generate data with better quality, such as dialogue text.

[0226] In addition to the television device provided in the foregoing embodiments, a conversation processing method is further provided in the following through embodiments, which can be applied to the controller of the television device, and specifically as follows.

[0227] In an exemplary embodiment, the method applied to the controller of the television device in the foregoing embodiments includes the following.

[0228] When the display component of the television device displays the first interface, and the voice collection component of the television device collects the current user voice, the current user voice is subjected to voiceprint recognition, the voiceprint identifier corresponding to the current user is determined based on the voiceprint recognition result, the user feature information corresponding to the current user is extracted from the preset knowledge graph based on the voiceprint identifier, and the current user voice is converted into corresponding voice text, and the voice text is displayed through the first interface; the first interface is a preset interface for voice dialogue with the digital human;

[0229] The first request is initiated to the first server through the network communication device; the first request carries the dialogue text request instruction, the voice text, and the user feature information; the dialogue text returned by the first server is received; the dialogue text is generated based on the first request and the preset generative model in the first server;

[0230] The second request is initiated to the second server through the network communication device; the second request carries the text-to-voice instruction and the dialogue text;

[0231] The dialogue voice data returned by the second server is received; the dialogue voice data is generated based on the second request and the preset text-to-voice service in the second server;

[0232] The dialogue text is sent to the display component, and the dialogue voice data is sent to the audio output component of the television device and the preset digital human driving system respectively, so as to display the dialogue text through the first interface, output the dialogue voice data through the audio output component, and drive the preset digital human image to make corresponding actions through the digital human driving system.

[0233] For understanding of the technical solutions of the present embodiment, reference can be made to the relevant parts in the foregoing embodiments, which will not be repeated here.

[0234] In the present embodiment, when the television device enters a function page related to digital foreign language training, and the user operates a target key on the remote controller, it indicates that the user needs to start the voice recognition function of the television device to have a foreign language conversation with the digital person. In response to this, the controller in the television device converts the current user voice into voice text and displays it on the first interface, so that the user knows the content of the user's speech captured by the television device. At the same time, the controller performs voiceprint recognition to determine the corresponding voiceprint identifier, and extracts the corresponding user feature information from the knowledge graph according to the voiceprint identifier. The controller sends a first request to the first server through the network communication device to obtain the conversation text for the user's speech content. Since the conversation text is generated based on the first request and the preset generative model in the first server, and the first request carries the voice text and user feature information, a better quality conversation text can be obtained. The controller sends a second request to the second server through the network communication device to obtain the conversation voice data. The obtained conversation voice data is consistent with the content of the conversation text generated by the generative model. Finally, the controller sends the above conversation text to the display component of the television device to display the conversation text, sends the conversation voice data to the audio output component to output the conversation voice data, and sends the conversation voice data to the digital person driving system to drive the digital person to make corresponding actions. Thus, based on the cooperation of the controller of the television device, the generative model in the first server, and the text-to-speech service in the second server, the user can practice foreign language conversation with the digital person. Since the conversation text of the digital person for the user's speech content is generated based on the generative model and the user feature information, the user can be provided with personalized conversation content, and the quality of the conversation text is better. At the same time, through the text-to-speech service and the digital person driving system, the content of the conversation text, the action of the digital person, and the voice of the digital person are consistent, so that a better foreign language conversation experience can be provided.

[0235] In one exemplary embodiment, a television device is provided, which is configured with an English oral coach to provide English oral training services to users. The English oral coach can be a virtual character presented in the television device, i.e., a digital person or a 3D digital person. The driving principle of the 3D digital person mainly includes speech recognition, motion capture, and expression recognition technologies. These technologies enable the digital person to respond to user inputs in terms of actions and expressions, enabling interaction with the user. Speech recognition technology can convert user speech into text, motion capture technology can obtain user motion data, and expression recognition technology can recognize user expressions, thereby driving the digital person to make corresponding actions and expression responses.

[0236] Current English chat products are based on traditional chat models, and the conversation content is not smooth enough. The pronunciation and grammar errors in the conversation process cannot be timely feedback. The technical solutions of the embodiment are based on speech recognition, large model text generation, speech evaluation technology, 3D digital human technology, voiceprint recognition technology, and dialogue related knowledge graph construction and use, etc. The conversation is more smooth and natural, the body language of the 3D digital human in the conversation process makes the conversation more human, and the conversation record can let the user clearly understand his own weaknesses and learning progress. The technical solutions of the embodiment mainly involve three parts: 1. Conversation practice; 2. Guessing game; 3. Knowledge graph. The specific content is as follows:

[0237] 1. Conversation practice

[0238] As shown in Figure 7 , a possible timing for implementing conversation practice / chat interaction is provided, wherein the digital human control related content corresponds to the controller related content in the foregoing embodiment, the large model service and the TTS service (Text-to-Speech Service) can correspond to the first server and the second server related content in the foregoing embodiment, and the digital human service can correspond to the digital human driving system related content in the foregoing embodiment.

[0239] The main functions of the conversation practice include English conversation, grammar correction, advanced expression (prompting the user to use more advanced expression), conversation duration statistics, word quantity statistics, and pronunciation fluency score statistics. The specific process is as follows:

[0240] 1. The user initiates a voice request, and the television terminal calls the ASR (Automatic Speech Recognition) service to perform speech recognition and pronunciation evaluation, and obtains corresponding text and pronunciation score;

[0241] 2. The television calls the digital human control service, and the digital human control calls the large model service to obtain the dialogue result; at the same time, the large model is used to obtain the grammar correction result and the advanced expression result, which are sent to the television terminal for display;

[0242] 3. The digital human control calls the TTS service to obtain the audio according to the dialogue result of the large model, and uses the audio to call the digital human service to obtain the driving data, which is sent to the television terminal to drive the digital human; The interaction process of the digital human control and the large model, and the digital human service is streaming, and will not be blocked, and the dialogue expression of the digital human is more smooth;

[0243] 4. During the interaction process, the dialogue duration, dialogue content, user pronunciation score and other data are saved and statistically analyzed. As shown in Figure 8As shown, a possible form of statistical data is provided.

[0244] II. Guessing game

[0245] The guessing game corresponds to the voice guessing game mode in the foregoing embodiment. The guessing game is based on the dialogue practice and adds the function of guessing English word phrases from pictures. The interaction logic here also includes the interaction process with the large model service, the TTS service, and the digital human driving service. For related details, please refer to the foregoing embodiments and the accompanying drawings. Figure 9 As shown, a possible timing of the guessing game is provided. Figure 9 As shown, a possible interface display of the guessing game is provided. As shown in the figure, the “XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX” displayed in the figure can be an opening text in a foreign language, such as “Hello, let’s start the game of guessing the English content from pictures”. As for the first interface in the foregoing embodiment, the overall framework can also refer to Figure 10 As shown, a possible interface display of the guessing game is provided. As shown in the figure, the “XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX” displayed in the figure can be an opening text in a foreign language, such as “Hello, let’s start the game of guessing the English content from pictures”. As for the first interface in the foregoing embodiment, the overall framework can also refer to Figure 10 , but the specific display content can be different, for example, the “theme” in the figure can display the dialogue mode, there can be no picture in the dialogue box, the opening text is different, etc. The specific flow of the guessing game is as follows:

[0246] 1. Enter the guessing game, the television initiates a request to the digital human central control, triggers the card generation logic, and the digital human central control randomly obtains 100 pieces of card information from the preset database and puts it into redis (Remote Dictionary Server, i.e., Remote Dictionary Service), that is, extracts a preset number of pictures from the preset picture database and caches them; randomly obtains 1 piece of picture information from the redis, that is, determines the first target picture, assembles the prompt with the picture information, asynchronously requests the large model service to obtain the opening speech, and returns the television terminal for display, that is, displays the first target picture and the opening text through the second interface, outputs the opening voice data through the audio output component, and drives the preset digital human image to make corresponding actions through the digital human driving system.

[0247] 2. The user answers the card content. The user's answer and the original card information are assembled into a prompt. An asynchronous request is made to the large model service to obtain the judgment result, i.e., the judgment result text. At the same time, a card information assembly result is asynchronously retrieved from Redis, i.e., the second target image is determined. If the data is not in Redis, it is retrieved from the database and added to Redis, and then sent to the TV terminal for display. That is, the second target image and the guessing result text are displayed through the second interface, and the guessing result audio data is output through the audio output component. The digital human driving system drives the preset digital human image to perform actions corresponding to the guessing result audio data.

[0248] 3. Repeating the above process constitutes the entire interactive logic of the picture guessing game.

[0249] III. Knowledge Graph

[0250] This embodiment utilizes voiceprint technology to distinguish different users and establish corresponding knowledge graphs, forming long-term memory capabilities, thereby helping to improve dialogue quality. The specific solution is as follows:

[0251] 1. Use voiceprint recognition to distinguish different users. When a user is having a conversation, use voiceprint recognition to identify the user's voiceprint ID and report it to the digital human control center along with the identified text.

[0252] 2. Utilize large-scale model services for entity extraction and content summarization, storing the extracted entities in a graph database. Extracted entities include name, gender, age, interests, and interpersonal relationships. The television terminal reports the voiceprint ID and text content to the digital human central control unit. The digital human central control unit asynchronously extracts entity relationships from the text based on existing logic, establishing a knowledge graph for the task profile.

[0253] The following example illustrates the construction of knowledge graphs:

[0254] User 1 inputs the following using voice or other methods: "my name is Tom," meaning my name is Tom. For example... Figure 11 As shown, a corresponding knowledge graph is provided, where voiceprint ID 1 corresponds to user 1.

[0255] User 1 inputs the following using voice or other methods: "I love to play badminton," which means "I love playing badminton." Figure 12 As shown, a corresponding knowledge graph is provided.

[0256] User 1 used voice or other means to say: "My friend's name is Alice." Figure 13 As shown, a corresponding knowledge graph is provided.

[0257] At this time, user 2 speaks, user 2: hi, I am Alice. As shown in Figure 14 A corresponding knowledge graph is provided.

[0258] In some possible embodiments, different colors can be used to represent different labels, such as yellow for voiceprint id and blue for name.

[0259] 3. The extracted entities and relationships are saved into an image database. When a user initiates a conversation, the corresponding user information can be retrieved according to the voiceprint id and the entities extracted from the user conversation, and integrated into the prompt of the large model, thereby realizing long-term memory based on the voiceprint id and improving the quality of the conversation.

[0260] In a multi-person conversation scenario, different users are distinguished based on voiceprint id, and personalized conversations are conducted based on the collected user information.

[0261] In some possible embodiments, the large model service can be a self-developed specific large model using a server-side deployment method.

[0262] In some possible embodiments, the digital human 3D model is a terminal deployment method, and the terminal pre-downloads related components of the 3D model according to the image selected by the user.

[0263] In some possible embodiments, the definition process and instruction issuing process of the digital human are as follows:

[0264] 1. Define the digital human image

[0265] Different clothes, hairstyles, accessories, shoes, and props are combined to form different digital human images. The same digital human image can be designed for different users and different emotions by changing the color of the clothes and the action.

[0266] 2. Define the action animation of the digital human

[0267] The parameters of the action animation of the digital human are defined in advance, including arm swing angle, knee bending angle, facial expression parameters, etc. The actual situation of the digital human model is not limited.

[0268] 3. Establish the relationship between the topic, emotion, and digital human image

[0269] Different digital human images are used in different topics, and a digital human sports suit and a football prop are used in a sports topic scene; in the same chat topic, different digital human images are used according to different emotions of a user, for example, in a casual chat mode, if the emotion of the user is happy, a happy digital human image is used, and if the emotion of the user is sad, a co-emotion image is used, so that the user feels the emotional resonance of the digital human; wherein the user emotion column common represents that the emotion is not limited.

[0270] 4. According to the chat topic selected by the user and the emotion of the user recognized by the large model, a corresponding digital human image and action are matched, the digital human image parameters, action parameters and broadcast speech are input to a digital human driving system, the specific parameters of the digital human are obtained according to the digital human image parameters, the lip shape parameters are obtained through the digital human face driving algorithm, and the image parameters, action parameters, face, expression parameters and voice are issued to the terminal, and the terminal drives the digital human model to make corresponding expressions and actions.

[0271] In some possible embodiments, the digital human corresponding image can be multiple, for example, a happy image in a sports topic, a co-emotion image in casual chat and the like; each digital human image can have a respective digital human image id; different digital human images can be different in appearance, for example, clothes, accessories, actions and the like can be different.

[0272] In some possible embodiments, a matching digital human image can be determined from multiple digital human images according to a topic, an intention of a user, an emotion and the like to show to the user.

[0273] The technical solution in the embodiment is based on speech recognition, large model text generation, speech evaluation technology, 3D digital human technology, so that the dialogue is more fluent and natural, the body language of the 3D digital human in the dialogue process makes the dialogue more human, and the dialogue record can enable the user to clearly understand his own weaknesses and learning progress. By using speech intelligent evaluation, pronunciation scores are provided for the user, so that the user can correct pronunciation; English dialogue is performed in combination with the large model, so that the dialogue content is more fluent and natural; the interaction process is more humanized by using the 3D digital human driving technology; a learning record statistics function is provided, and the user can more intuitively understand the learning process. Meanwhile, different users are distinguished by using voiceprint technology, a knowledge graph is constructed, and long-term memory based on a voiceprint id is realized by combining the large model technology and prompt technology, so as to improve the dialogue quality; in a multi-person dialogue scene, different users are distinguished based on the voiceprint id, and personalized dialogue is performed according to collected user information.

[0274] In one exemplary embodiment, a conversation processing apparatus is provided, applied to a controller of a television device, comprising:

[0275] The identification module is configured to, when the display component of the television device displays a first interface and the voice collection component of the television device collects a current user voice, perform voiceprint identification on the current user voice, determine a voiceprint identifier corresponding to the current user based on a voiceprint identification result, extract user feature information corresponding to the current user from a preset knowledge graph based on the voiceprint identifier, and convert the current user voice into corresponding voice text and display the voice text through the first interface. The first interface is a preset interface for voice dialogue with the digital person.

[0276] The dialogue content generation module is configured to initiate a first request to a first server through a network communication device, the first request carrying a dialogue text request instruction, voice text, and user feature information, receive dialogue text returned by the first server, wherein the dialogue text is generated based on the first request and a preset generative model in the first server, initiate a second request to a second server through the network communication device, the second request carrying a text-to-speech instruction and the dialogue text, and receive dialogue voice data returned by the second server, wherein the dialogue voice data is generated based on the second request and a preset text-to-speech service in the second server.

[0277] The dialogue content output module is configured to send the dialogue text to the display component, send the dialogue voice data to an audio output component of the television device and a preset digital person driving system, display the dialogue text through the first interface, output the dialogue voice data through the audio output component, and drive the preset digital person image to make corresponding actions through the digital person driving system.

[0278] The modules in the above session processing apparatus can also have other functions, or the above session processing apparatus can further include other modules, so that the above session processing apparatus can have functions corresponding to the session processing method in the foregoing embodiments.

[0279] The modules in the above session processing apparatus can be implemented by software, hardware, or a combination thereof. The modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the modules.

[0280] In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above method embodiments when executing the computer program.

[0281] In one embodiment, a computer readable storage medium is provided, storing a computer program, and the computer program is executed by a processor to implement the steps in the above method embodiments.

[0282] In an embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of any of the above method embodiments.

[0283] It should be understood that the various steps, processes, etc. in the above-described embodiments are not necessarily performed in the order indicated. Unless specifically stated, the execution of these steps, processes, etc. is not strictly sequential and they can be performed in other orders. Also, at least some of the steps in the above-described embodiments can comprise multiple steps or stages, which are not necessarily performed at the same time, but can be performed at different times, and the order of execution of these steps or stages is not necessarily sequential, but can be round-robin or alternating with other steps or steps or stages in other steps.

Claims

1. A television apparatus, characterized by comprising: The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device.

2. The television apparatus of claim 1, wherein, The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device.

3. The television apparatus of claim 1, wherein, The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application relates to a voice interaction method and device. The application convert the historical user voice into corresponding historical voice text, and extract entity information corresponding to a pre-defined entity word and an association relationship between different entity information from the historical voice text; the pre-defined entity word includes one or more of a name, a gender, an age, an interest, and a character relationship; construct the knowledge graph according to at least one voiceprint identifier of the historical user, the entity information, and the association relationship; a node in the knowledge graph represents the voiceprint identifier and the entity information, and an edge between nodes represents the association relationship.

4. The television apparatus of claim 3, wherein, The extracting of the entity information corresponding to the pre-defined entity word and the association relationship between different entity information from the historical voice text includes: initiating a third request to a first server through the network communication device; the third request carries the historical voice text and an entity information extraction instruction; receiving entity information and an association relationship between different entity information returned by the first server; the entity information and the association relationship are generated based on the third request, a pre-set entity word set in the first server, and a generative model.

5. The television apparatus of claim 4, wherein, The controller is further configured to: after initiating the first request to the first server through the network communication device, receive entity information contained in voice text of the current user voice and an association relationship between different entity information returned by the first server based on the first request, and update the knowledge graph.

6. The television apparatus of claim 3, wherein, The extracting of user feature information corresponding to the current user from a pre-set knowledge graph based on the voiceprint identifier includes: determining a node corresponding to the voiceprint identifier from the pre-set knowledge graph as a target node; determining at least one associated node of the target node, the associated node and the target node having an edge therebetween; obtaining entity information represented by the associated node and an association relationship between the associated node and the target node as user feature information corresponding to the current user.

7. The television apparatus of claim 2, wherein, The controller is further configured to: in a case where the display component switches from another interface to display the second interface, extract a pre-set number of pictures from a pre-set picture database and cache; the pre-set number is an integer greater than 1; determine one picture from the cached pictures as a first target picture; initiate a fourth request to a first server through the network communication device; the fourth request carries an opening text request instruction and the first target picture; receive an opening text returned by the first server; the opening text is generated based on the fourth request and a pre-set generative model in the first server; the opening text is a first-round text; initiate a fifth request to a second server through the network communication device; the fifth request carries a text-to-speech instruction and the opening text; receive opening voice data returned by the second server; the opening voice data is generated based on the fifth request and a pre-set text-to-speech service in the second server; the opening voice data is first-round voice data; send the first target picture and the opening text to the display component, send the opening voice data to the audio output component and the digital human driving system respectively, so as to display the first target picture and the opening text through the second interface, output the opening voice data through the audio output component, and drive the preset digital human image to make actions corresponding to the opening voice data through the digital human driving system; The opening text and the opening voice data correspond to the same non-Chinese language.

8. The television apparatus of claim 7, wherein, The controller is further configured to: In the case that the display component displays the second interface and the voice collection component collects the current user voice, convert the current user voice into corresponding picture guessing answer text, and display the picture guessing answer text through the second interface; initiate a sixth request to a first server through the network communication device; the sixth request carries a result determination request instruction, the picture guessing answer text, and the first target picture; receive a determination result text returned by the first server; The determination result text is generated based on the sixth request and a preset generative model in the first server. initiate a seventh request to a second server through the network communication device; the seventh request carries a text-to-speech instruction and the determination result text; receive picture guessing result voice data returned by the second server; The picture guessing result voice data is generated based on the seventh request and a preset text-to-speech service in the second server. Determine another picture from the cached pictures as a second target picture; Obtain a preset prompt text and corresponding voice data thereof; The prompt text and corresponding voice data thereof are used to prompt the user to continue guessing pictures; send the picture guessing result text, the second target picture, and the prompt text to the display component, and send the combination of the picture guessing result voice data and the voice data corresponding to the prompt text to the audio output component and the digital human driving system respectively, so as to display the picture guessing result text, the second target picture, and the prompt text through the second interface, output the combination through the audio output component, and drive the preset digital human image to make actions corresponding to the opening voice data through the digital human driving system; The picture guessing result text and the picture guessing result voice data correspond to the non-Chinese language.

9. The television device according to any one of claims 1-8, wherein The first server and the second server are servers deployed in the cloud, and the first server and the second server are the same server or different servers. The generative model preset in the first server is a model based on a Transformer architecture and having a text-to-text function.

10. A session processing method, characterized by, A controller applied to a television device, The method comprises: In a case where a display component of a television device displays a first interface and a voice collection component of the television device collects a current user voice, performing voiceprint recognition on the current user voice, determining a voiceprint identifier corresponding to the current user based on a voiceprint recognition result, extracting user feature information corresponding to the current user from a preset knowledge graph based on the voiceprint identifier, and converting the current user voice into corresponding voice text, and displaying the voice text through the first interface; the first interface is a preset interface for voice conversation with a digital person; initiating a first request to a first server through a network communication device; the first request carries a conversation text request instruction, the voice text, and the user feature information; receiving a conversation text returned by the first server; the conversation text is generated based on the first request and a preset generative model in the first server; initiating a second request to a second server through the network communication device; the second request carries a text-to-speech instruction and the conversation text; receiving conversation voice data returned by the second server; the conversation voice data is generated based on the second request and a preset text-to-speech service in the second server; sending the conversation text to the display component and sending the conversation voice data to an audio output component of the television device and a preset digital person driving system respectively, so as to display the conversation text through the first interface, output the conversation voice data through the audio output component, and drive a preset digital person image to make corresponding actions through the digital person driving system.

Citation Information

Patent Citations

  • Virtual digital human interaction method and device, electronic equipment and storage medium

    CN117493501A

  • Personalized three-dimensional digital human holographic interaction forming system and method

    CN117523088A