Server, display device and digital human processing method

CN120937376APending Publication Date: 2025-11-11HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202480019607.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-27
Filing Date
2024-05-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing digital human technology has limited application scenarios and image display, mainly confined to single scenarios such as virtual news anchors and educational video lecturers, and users cannot customize their digital human image.

Method used

A server and display device are provided. By receiving voice data input by the user, identifying entity data to obtain digital human data and media asset data, the data is sent to the display device for playback. The system supports user-defined digital human avatars and interactions, and utilizes the server to generate and push digital human data, enabling multi-scenario applications.

Benefits of technology

It enables the application of digital humans in multiple scenarios, supports users to customize digital human appearance and interaction, and improves the flexibility of digital human technology and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937376A_ABST
    Figure CN120937376A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a server, display equipment and a digital human processing method. The method comprises the following steps: receiving voice data input by a user and sent by the display equipment; recognizing the voice data to obtain a recognition result; if the identification result comprises entity data, media asset data corresponding to the identification result and digital human data corresponding to the entity data are obtained, and the entity data comprise human names and / or media asset names; and sending the digital person data and the media asset data to the display device, playing audio and video data or displaying interface data, and playing images and voices of the digital person according to the digital person data. According to the embodiment of the invention, the entity data included in the voice data uploaded by the display device is recognized, the digital person data corresponding to the entity data is issued to the display device, corresponding scene display is performed in combination with semantic understanding, and the interesting experience of voice interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

A server, display device and digital human processing method

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] The present disclosure claims priority to Chinese applications filed on June 25, 2023, with application numbers 202310758892.0; filed on September 27, 2023, with application numbers 202311256230.X; filed on September 27, 2023, with application numbers 202311256277.6; filed on September 27, 2023, with application numbers 202311259355.8; filed on September 27, 2023, with application numbers 202311267720.X; and filed on September 27, 2023, with application numbers 202311258706.3, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of digital human interaction technology, and in particular to a server, a display device, and a digital human processing method. Background Art

[0004] With the continuous development of artificial intelligence (AI), digital humans have become a highly sought-after technology. Digital humans are virtual characters generated by computer programs and algorithms. They can simulate human language, behavior, emotions, and other characteristics, and possess a high degree of intelligence and interactivity. Currently, digital human technology is primarily used in gaming, education, healthcare, finance, and other fields.

[0005] Digital human applications are relatively limited to single scenarios, such as virtual news anchors and educational video lecturers. Digital human image presentation is also relatively simple, simply replacing the traditional voice assistant image, with users choosing from a selection of digital human images.

[0006] Summary of the Invention

[0007] In a first aspect, some embodiments of the present disclosure provide a server that can be configured to: receive voice data input by a user sent by a display device; recognize the voice data to obtain a recognition result; if the recognition result includes entity data, obtain media data corresponding to the recognition result, and digital human data corresponding to the entity data; wherein the entity data includes a character name and / or a media name, the digital human data includes the image data and broadcast voice of the digital human, and the media data includes audio and video data or interface data; send the digital human data and the media data to the display device, so that the display device plays the audio and video data or displays the interface data, and plays the image and voice of the digital human according to the digital human data.

[0008] In a second aspect, some embodiments of the present disclosure provide a display device, which may include: a display configured to display an image and / or a user input interface; a user input interface configured to receive instructions from a user; a Bluetooth module configured to perform operations related to the Bluetooth protocol; a communication device configured to communicate with an external device according to a predetermined protocol; a memory configured to store computer instructions and data associated with the display device; at least one processor connected to the display, user input interface, Bluetooth module, communication device and memory, and configured to execute computer instructions so that the display device performs the following: receiving voice data input by the user; sending the voice data to a server through the communication device; receiving digital human data issued by the server based on the voice data; and playing the image and voice of the digital human according to the digital human data.

[0009] On the third aspect, some embodiments of the present disclosure provide a digital human processing method, which may include: receiving voice data input by a user sent by a display device; recognizing the voice data to obtain a recognition result; if the recognition result includes entity data, obtaining media data corresponding to the recognition result, and digital human data corresponding to the entity data, the entity data including a character name and / or a media name, the digital human data including the image data and broadcast voice of the digital human, and the media data including audio and video data or interface data; sending the digital human data and the media data to the display device, so that the display device plays the audio and video data or displays the interface data, and plays the image and voice of the digital human according to the digital human data. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] FIG1 illustrates an operation scenario between a display device and a control apparatus according to some embodiments;

[0011] FIG2 is a block diagram of a hardware configuration of a control device according to some embodiments;

[0012] FIG3 is a block diagram of a hardware configuration of a display device according to some embodiments;

[0013] FIG4 is a diagram illustrating a software configuration in a display device according to some embodiments;

[0014] FIG5 is a diagram illustrating another software configuration in a display device according to some embodiments;

[0015] FIG6 is a flowchart of digital human interaction according to some embodiments;

[0016] FIG7 is a schematic diagram of a digital human entry interface provided according to some embodiments;

[0017] FIG8 is a schematic diagram of a digital human selection interface provided according to some embodiments;

[0018] FIG9 is a flow chart of displaying a digital human interface according to some embodiments;

[0019] FIG10 is a flow chart of adding a digital human interface according to some embodiments;

[0020] FIG11 is a schematic diagram of a video recording preparation interface provided according to some embodiments;

[0021] FIG12 is a schematic diagram of a tone setting interface provided according to some embodiments;

[0022] FIG13 is a schematic diagram of an audio recording preparation interface provided according to some embodiments;

[0023] FIG14 is a schematic diagram of a digital human naming interface according to some embodiments;

[0024] FIG15 is a schematic diagram of another digital human selection interface provided according to some embodiments;

[0025] FIG16 is a flowchart of a digital human customization method according to some embodiments;

[0026] FIG17 is another flowchart of digital human interaction according to some embodiments;

[0027] FIG18 is a schematic diagram of a live data streaming process provided according to some embodiments;

[0028] FIG19 is a schematic diagram of a user interface according to some embodiments;

[0029] FIG20 is another digital human interaction timing diagram provided according to some embodiments;

[0030] FIG21 is another flowchart of digital human interaction according to some embodiments;

[0031] FIG22 is a flowchart of generating a digital human image model according to some embodiments;

[0032] FIG23 is a schematic diagram of another digital human data playback interface provided according to some embodiments;

[0033] FIG24 is another flowchart of digital human interaction according to some embodiments;

[0034] FIG25 is a schematic diagram of another digital human data playback interface provided according to some embodiments;

[0035] FIG26 is a schematic diagram of another digital human data playback interface provided according to some embodiments;

[0036] FIG27 is a schematic diagram of another digital human data playback interface provided according to some embodiments;

[0037] FIG28 is a schematic diagram of another digital human data playback interface provided according to some embodiments;

[0038] FIG29 is a flow chart of a server performing voice interaction according to some embodiments;

[0039] FIG30 is a schematic diagram of an emotional speech model provided according to some embodiments;

[0040] FIG31 is a flow chart of obtaining emotion type and emotion intensity according to some embodiments;

[0041] FIG32 is a schematic diagram of another emotional speech model provided according to some embodiments;

[0042] FIG33 is another flowchart of digital human interaction according to some embodiments;

[0043] FIG34 is a schematic diagram of a personal center interface provided according to some embodiments;

[0044] FIG35 is a schematic diagram of a family relationship according to some embodiments;

[0045] FIG36 is a flowchart of voiceprint recognition according to some embodiments;

[0046] FIG37 is a schematic diagram of another digital human data playback interface provided according to some embodiments;

[0047] FIG38 is a schematic diagram of a digital human driving process according to some embodiments;

[0048] FIG39 is another schematic diagram of a digital human driving process according to some embodiments;

[0049] FIG40 is another schematic diagram of a digital human driving process according to some embodiments;

[0050] FIG41 is another schematic diagram of a digital human driving process according to some embodiments;

[0051] FIG42 is another schematic diagram of a digital human driving process according to some embodiments;

[0052] FIG43 is another schematic diagram of a digital human driving process according to some embodiments;

[0053] FIG44 is another schematic diagram of a digital human driving process according to some embodiments;

[0054] FIG45 is another schematic diagram of a digital human driving process according to some embodiments;

[0055] Figure 46 is a structural diagram of a chip system provided according to some embodiments. DETAILED DESCRIPTION

[0056] The display device provided in the embodiments of the present disclosure can have various implementation forms, for example, it can be a television, a smart TV, a laser projection device, a monitor, an electronic bulletin board, an electronic table, etc. Figures 1 and 2 illustrate a specific embodiment of the display device of the present disclosure.

[0057] Fig. 1 is a schematic diagram of an operation scenario between a display device and a control apparatus according to an embodiment. As shown in Fig. 1 , a user can operate a display device 200 via a terminal 300 or a control apparatus 100 .

[0058] In some embodiments, the control device 100 may be a remote controller. Communication between the remote controller and the display device may include infrared protocol communication, Bluetooth protocol communication, or other short-range communication methods, and the display device 200 may be controlled wirelessly or wired. The user may control the display device 200 by inputting user commands through buttons on the remote controller, voice input, control panel input, and the like.

[0059] In some embodiments, a terminal 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) may also be used to control the display device 200. For example, the display device 200 may be controlled using an application running on the terminal 300.

[0060] In some embodiments, the display device may not use the aforementioned terminal 300 or control apparatus 100 to receive instructions, but may receive user control through touch or gestures.

[0061] In some embodiments, the display device 200 can also be controlled in a manner other than the control device 100 and the terminal 300. For example, the user's voice command control can be directly received through a module for obtaining voice commands configured inside the display device 200, or the user's voice command control can be received through a voice control device set outside the display device 200.

[0062] In some embodiments, the display device 200 also communicates data with the server 400. The display device 200 may be connected to a local area network (LAN), a wireless local area network (WLAN), or other networks. The server 400 may provide various content and interactions to the display device 200. The server 400 may be a single cluster or multiple clusters, and may include one or more types of servers.

[0063] Figure 2 is a block diagram of the configuration of a control device 100 according to an exemplary embodiment. As shown in Figure 2, the control device 100 includes a processor 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 receives user input commands and converts them into commands that the display device 200 can recognize and respond to, acting as an intermediary for interaction between the user and the display device 200.

[0064] 3 , the display device 200 may include at least one of a tuner 210 , a communication device 220 , a detector 230 , an external device interface 240 , a processor 250 , a display 260 , an audio output interface 270 , a memory, a power supply, and a user input interface.

[0065] In some embodiments, the processor may include one or more processors, for example, a video processor, an audio processor, a graphics processor, RAM, ROM, and first to nth interfaces for input / output.

[0066] The display 260 includes a display screen component for presenting images, a driving component for driving image display, a component for receiving image signals output from a processor, and a component for displaying video content, image content, and a menu control interface and a user control UI interface.

[0067] The display 260 may be a liquid crystal display, an OLED display, or a projection display, and may also be a projection device and a projection screen.

[0068] The display 260 may further include a touch screen, which is used to receive control instructions input by actions such as sliding or clicking a user's finger on the touch screen.

[0069] The communication device 220 is a component used to communicate with external devices or servers using various communication protocols. For example, the communication device may include at least one of a Wi-Fi module, a Bluetooth module, a wired Ethernet module, or other network communication protocol chip or a near-field communication protocol chip, as well as an infrared receiver. The display device 200 can use the communication device 220 to send and receive control signals and data signals with the external control device 100 or the server 400.

[0070] The user input interface can be used to receive control signals from the control device 100 (such as an infrared remote controller, etc.).

[0071] Detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 may include a light receiver, a sensor for collecting ambient light intensity; or, detector 230 may include an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or, detector 230 may include a sound collector, such as a microphone, for receiving external sounds.

[0072] The external device interface 240 may include, but is not limited to, any one or more of the following: a high-definition multimedia interface (HDMI), an analog or digital high-definition component input interface (component), a composite video input interface (CVBS), a USB input interface (USB), an RGB port, etc. It may also be a composite input / output interface formed by multiple of the above interfaces.

[0073] The tuner-demodulator 210 receives broadcast television signals via a wired or wireless reception method, and demodulates audio and video signals, such as EPG data signals, from a plurality of wireless or wired broadcast television signals.

[0074] In some embodiments, the processor 250 and the tuner / demodulator 210 may be located in different separate devices, that is, the tuner / demodulator 210 may also be located in an external device of the main device where the processor 250 is located, such as an external set-top box.

[0075] Processor 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. Processor 250 controls the overall operation of display device 200. For example, in response to receiving a user command to select a UI object to be displayed on display 260, processor 250 may perform operations related to the object selected by the user command.

[0076] In some embodiments, the processor may include at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (Random Access Memory, RAM), ROM (Read-Only Memory, ROM), a first interface to an nth interface for input / output, a communication bus (Bus), etc.

[0077] The user may input a user command through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input command through the graphical user interface (GUI). Alternatively, the user may input a user command through a specific voice or gesture, and the user input interface may recognize the voice or gesture through a sensor to receive the user input command.

[0078] A "user interface" is a medium interface for interaction and information exchange between an application or operating system and a user. It converts information between its internal form and a form acceptable to the user. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operations that uses a graphical display. It can be an icon, window, control, or other interface element displayed on the display of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0079] In some embodiments, as shown in FIG. 4 , the system of the display device can be divided into three layers, namely, an application layer, a middleware layer, and a hardware layer from top to bottom.

[0080] The application layer mainly includes commonly used applications on the TV and the application framework. Among them, commonly used applications are mainly browser-based applications, such as HTML5 apps, and native apps.

[0081] An application framework can be a complete program model that has all the basic functions required by standard application software, such as file access, data exchange, etc., as well as the user interface for these functions (toolbars, status bars, menus, dialog boxes).

[0082] Native apps can support online or offline, message push or local resource access.

[0083] The middleware layer includes various television protocols, multimedia protocols, and system components. Middleware can leverage the fundamental services (functions) provided by system software to connect various parts of the application system or different applications on the network, enabling resource and function sharing.

[0084] The hardware layer includes the HAL interface, hardware, and drivers. The HAL interface is a unified interface for all TV chipsets, while the specific logic is implemented by each chip. Drivers mainly include: audio driver, display driver, Bluetooth driver, camera driver, Wi-Fi driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.

[0085] Referring to Figure 5, in some embodiments, the system is divided into four layers, namely, from top to bottom, the application layer (referred to as "application layer"), the application framework layer (referred to as "framework layer"), the Android runtime (Android runtime) and system library layer (referred to as "system runtime library layer"), and the kernel layer.

[0086] In some embodiments, at least one application runs in the application layer. These applications can be window programs, system settings programs, clock programs, etc. that come with the operating system, or applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.

[0087] The framework layer provides applications with an application programming interface (API) and programming framework. The application framework layer may include predefined functions. The application framework layer acts as a processing center, determining the actions taken by applications in the application layer. Through the API, applications can access system resources and services during execution.

[0088] As shown in Figure 5, the application framework layer in the embodiment of the present disclosure may include managers, content providers, etc., wherein the manager may include at least one of the following modules: an activity manager (Activity Manager) for interacting with all activities running in the system; a location manager (Location Manager) for providing system services or applications with access to system location services; a package manager (Package Manager) for retrieving various information related to the application packages currently installed on the device; a notification manager (Notification Manager) for controlling the display and clearing of notification messages; and a window manager (Window Manager) for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0089] In some embodiments, the activity manager is used to manage the lifecycle of each application and common navigation back functions, such as controlling the exit, opening, and backing of an application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling display window changes (such as shrinking the display window, shaking the display, distorting the display, etc.).

[0090] In some embodiments, the system runtime layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system will run the C / C++ library contained in the system runtime layer to implement the functions to be implemented by the framework layer.

[0091] In some embodiments, the kernel layer is a layer between hardware and software. As shown in FIG5 , the kernel layer includes at least one of the following drivers: an audio driver, a display driver, a Bluetooth driver, a camera driver, a Wi-Fi driver, a USB driver, an HDMI driver, a sensor driver (such as a fingerprint sensor, a temperature sensor, a pressure sensor, etc.), and a power driver.

[0092] With the continuous development of artificial intelligence (AI), digital humans have become a highly sought-after technology. Digital humans are virtual characters generated by computer programs and algorithms. They can simulate human language, behavior, emotions, and other characteristics, and possess a high degree of intelligence and interactivity. Currently, digital human technology is primarily used in gaming, education, healthcare, finance, and other fields.

[0093] Digital human applications are relatively limited to single scenarios, such as virtual news anchors and educational video lecturers. Digital human image presentation is also relatively simple, simply replacing the traditional voice assistant image, with users choosing from a selection of digital human images.

[0094] The present disclosure provides a method for processing a digital human, as shown in FIG6 , which may include the following steps:

[0095] Step S501: The terminal 300 establishes an association relationship with the display device 200 through the server 400;

[0096] In some embodiments, the server 400 establishes connection relationships with the display device 200 and the terminal 300 respectively, so that the display device 200 and the terminal 300 establish an association relationship.

[0097] The step of establishing a connection between the server 400 and the display device 200 may include:

[0098] The server 400 establishes a long connection with the display device 200;

[0099] The purpose of establishing a long connection between the server 400 and the display device 200 is that the server 400 can push the customized status of the digital human to the display device 200 in real time.

[0100] A persistent connection allows multiple data packets to be sent continuously over a single connection. If no data packets are sent during the connection, both parties must send a link detection packet. A persistent connection only requires a single connection establishment for multiple communications, saving network overhead. A persistent connection only requires a single handshake and authentication to maintain communication, improving communication efficiency. A persistent connection enables two-way data transmission, allowing the server to proactively send customized digital human data to the display device, achieving real-time communication.

[0101] In some embodiments, after receiving a power-on message from the display device 200 , the server 400 establishes a persistent connection with the display device 200 .

[0102] In some embodiments, after receiving a message from the display device 200 initiating the voice digital human service, the server 400 establishes a long connection with the display device 200 .

[0103] In some embodiments, after receiving the instruction to add a digital human from the display device 200 , the server 400 establishes a long connection with the display device 200 .

[0104] The server 400 receives the request data sent by the display device 200 ; wherein the request data may include the device identification of the display device 200 .

[0105] After receiving the request data, the server 400 determines whether there is an identification code corresponding to the device identifier in the database; wherein the identification code is used to represent the device information of the display device 200, and the identification code can be a number of random numbers or letters, a barcode, or a QR code.

[0106] If the identification code corresponding to the device identifier exists in the database, the identification code is sent to the display device 200 so that the display device 200 displays the identification code on the adding digital human interface.

[0107] If the identification code corresponding to the device identifier does not exist in the database, an identification code corresponding to the device identifier is created, the device identifier and the identification code are correspondingly saved in the database, and the identification code is sent to the display device 200 so that the display device 200 displays the identification code on the add digital human interface.

[0108] In order to clarify the interactive process of establishing a connection between the server 400 and the display device 200, the following embodiments are disclosed:

[0109] After receiving the instruction input by the user to open the digital human entrance interface, the display device 200 controls the display 260 to display the digital human entrance interface; wherein the digital human entrance interface may include a voice digital human control;

[0110] In some embodiments, as shown in FIG. 7 , the digital human portal interface may include a voice digital human control 61 , a natural conversation control 62 , a wake-up word-free control control 63 , and a focus 64 .

[0111] It should be noted that controls refer to visual objects displayed in various display areas of the user interface in the display device 200 to represent corresponding content such as icons, thumbnails, video clips, links, etc. These controls can provide users with various traditional program contents received through data broadcasting, as well as various applications and service contents set by content manufacturers.

[0112] The display format of a control is generally diverse. For example, a control may include text content and / or an image for displaying a thumbnail related to the text content, or a video clip related to the text. In another example, a control may be text and / or an icon for an application.

[0113] Focus is used to indicate that any one of the controls has been selected. On the one hand, a control can be selected or controlled by controlling the movement of a focus object displayed on the display device 200 according to the user's input through the control device 100. For example, a user can control the movement of a focus object between controls by pressing the up arrow key on the control device 100 to select and control a control. On the other hand, a control can be selected or controlled by controlling the movement of each control displayed on the display device 200 according to the user's input through the control device 100. For example, a user can control each control to move left and right at the same time by pressing the up arrow key on the control device 100 to select and control a control while keeping the focus object's position unchanged.

[0114] Focus identification often takes various forms. For example, the focus object's location can be identified by magnifying the item, by setting the item's background color, or by changing the border, size, color, transparency, outline, and / or font of the text or image of the focused item.

[0115] After receiving the user input to select the voice digital human control, the display device 200 controls the display 260 to display the digital human selection interface; wherein, the digital human selection interface may include at least one digital human control and an add control. The digital human control is displayed with the digital human image and the name corresponding to the digital human image. The add control is used to add a new digital human image, voice and name.

[0116] In some embodiments, as shown in FIG7 , after receiving a user input command to select the voice-activated digital human control 61, the display device 200 displays a digital human selection interface. As shown in FIG8 , the digital human selection interface may include a default image control 71, a Ding Ding control 72, a bottle control 73, an add control 74, and a focus 75. The user can select the desired digital human to respond to the voice command by moving the focus 75.

[0117] In some embodiments, the process of displaying the digital human interface on the display device 200 is shown in FIG9 , and may include the following steps:

[0118] Step S901: The Digital Human application requests homepage data from the voice zone;

[0119] Step S902: The voice zone obtains homepage configuration information from the operator;

[0120] Step S903: The operator returns the homepage data to the voice zone;

[0121] Step S904: the voice zone returns the display device data protocol to the digital human application;

[0122] Step S905: The Digihuman application requests the Digihuman account data from the voice zone;

[0123] Step S906: The voice zone obtains operation preset data from the operation terminal;

[0124] Step S907: The voice zone obtains the digital human account data stored in the cloud from the algorithm service;

[0125] Step S908: The algorithm service returns the cloud-stored digital human account data to the voice zone;

[0126] Step S909: The voice zone determines whether to supplement the default parameters;

[0127] Step S910: Based on the supplementary results of the default parameters, the voice zone returns the display device data protocol to the Digital Human application. In steps S901-S910, after the Digital Human application of the display device 200 receives a user input command to open the Digital Human portal interface (homepage), the Digital Human application requests the homepage data from the voice zone. The voice zone obtains the homepage configuration information (homepage data) from the operator end and sends the homepage data to the Digital Human application, so that the Digital Human application controls the display 260 to display the Digital Human homepage. The Digital Human application can directly send a Digital Human account request. After receiving the virtual Digital Human account request, the voice zone obtains preset data, such as default Digital Human account information, from the operator end. It also obtains the cloud-stored Digital Human account data from the algorithm service of the server 400. If default supplementary parameters are present, the preset data, cloud-stored Digital Human account data, and supplementary parameters are sent to the Digital Human application together. If no default supplementary parameters are present, the preset data and cloud-stored Digital Human account data are sent to the Digital Human application, so that the Digital Human application controls the display 260 to display the Digital Human selection interface after receiving the instruction to display the Digital Human selection interface. After displaying the Digital Human homepage, the Digital Human application can also send a virtual Digital Human account request after receiving the user's input instruction to display the Digital Human selection interface, and directly display the Digital Human selection interface after receiving the preset data, the Digital Human account data stored in the cloud and the supplementary parameters.

[0128] The voice zone, oriented towards server 400 and based on the operations support platform, implements operationally configurable management of default backend data items and configuration items, and completes the protocol delivery of data required by display device 200. The voice zone is connected to the algorithm service interaction between display device 200 and server 400. By obtaining data parameters reported by display device 200, it completes instruction parsing, completes algorithm backend interaction transfer, parses and delivers backend stored data, and ultimately realizes a full-link data docking process.

[0129] After receiving the instruction input by the user to select to add a control, the display device 200 sends request data carrying the device identification of the display device 200 to the customized central control service of the server 400.

[0130] The customized central control service calls the target application program interface to determine whether an identification code corresponding to the device identifier exists in the database. If an identification code corresponding to the device identifier exists in the database, the identification code is sent to the display device 200. If an identification code corresponding to the device identifier does not exist in the database, an identification code is created and sent to the display device 200. The target application refers to an application with an identification code recognition function.

[0131] The display device 200 receives the identification code sent by the server 400 and displays it on the add digital human interface.

[0132] In some embodiments, after receiving the user input instruction to select the add control 74 in FIG8 , the display device 200 will display the add digital human interface as shown in FIG10 , which may include a QR code 91 .

[0133] The step of establishing a connection between the server 400 and the terminal 300 may include: the server 400 receiving the identification code uploaded by the terminal 300; determining whether there is a display device 200 corresponding to the identification code;

[0134] If there is a display device 200 corresponding to the identification code, an association relationship between the terminal 300 and the display device 200 is established, so that the data uploaded by the terminal 300 is processed by the server 400 and then sent to the display device 200.

[0135] In order to clarify the interactive process of establishing a connection between the server 400 and the terminal 300, the following embodiments are disclosed:

[0136] After receiving the user input instruction to open the target application, the terminal 300 starts the target application and displays the home page interface corresponding to the target application. The home page interface may include a scan control.

[0137] After receiving the instruction input by the user to select the scan control, the terminal 300 displays the code scanning interface.

[0138] After scanning the identification code displayed on the display device 200, such as a QR code, the terminal 300 uploads the identification code to the server 400. The user can aim the camera of the terminal 300 at the identification code displayed on the digital human interface of the display device 200.

[0139] If the identification code is in the form of numbers or letters, the home page interface may include an identification code control. After receiving an instruction from the user to select the identification code control, the identification code input interface is displayed, and the numbers or letters displayed on the display device 200 are input into the identification code input interface to upload the identification code to the server 400.

[0140] Server 400 determines whether a display device corresponding to the identification code exists. If so, server 400 establishes an association between terminal 300 and display device 200, processing the data uploaded by terminal 300 on server 400 before sending it to display device 200. If no display device 200 corresponding to the identification code exists, a message indicating recognition failure is sent to terminal 300, causing terminal 300 to display an error message.

[0141] After determining that there is a display device 200 corresponding to the identification code, the server 400 sends a message indicating successful identification to the terminal 300. The terminal 300 displays a startup page, which starts the digital human customization process.

[0142] In some embodiments, the startup page may include a digital human avatar selection interface. The digital human avatar selection interface includes at least one default avatar control and a custom avatar control. Upon receiving a user input indicating the selection of a custom avatar control, the terminal 300 displays a video recording preparation interface, which may include recording controls. In some embodiments, as shown in FIG11 , the video recording preparation interface may include video recording instructions 101 and a start recording control 102.

[0143] In some embodiments, the startup page may also be a video recording preparation interface.

[0144] In some embodiments, the step of establishing an association between the terminal 300 and the display device 200 through the server 400 may include: the server 400 receives the user account and password uploaded by the terminal 300 and sends a login success message after verifying that the user account and password are correct, so that the terminal 300 can obtain the data corresponding to the user account.

[0145] The server 400 receives the user account and password uploaded by the display device 200, and after verifying that the user account and password are correct, sends a login success message so that the display device 200 can obtain the data corresponding to the user account. Among them, the terminal 300 and the display device 200 have the same login user account. The terminal 300 and the display device 200 establish an association relationship by logging in with the same user account so that the data updated by the terminal 300 can be synchronized to the display device 200. For example: the digital human-related data customized at the terminal 300 can be synchronized to the display device 200. Step S502: The terminal 300 uploads the image data and audio data to the server 400;

[0146] The image data may include videos or pictures taken by the user, videos or pictures selected by the user in the photo album, and videos or pictures downloaded from a website.

[0147] In some embodiments, the terminal 300 uploads the received video or picture taken by the user to the server 400 .

[0148] In some embodiments, after receiving a user input command to select the start recording control 102 in FIG11 , the video is recorded using the terminal 300 media component video. To avoid multiple recordings due to unsatisfactory face detection, the recording interface displays a suggested face location, allowing the terminal 300 to perform a preliminary face location detection. After recording ends, the recorded video can be repeatedly previewed. After receiving a user input command confirming upload, the user-recorded video is sent to the server 400.

[0149] In some embodiments, the terminal 300 may send the captured user photo to a server.

[0150] In some embodiments, the terminal 300 may select a user photo or a user video from an album and upload the user photo or the user video to the server 400 .

[0151] The server 400 receives the image data uploaded by the terminal;

[0152] Detect whether the facial points in the image data are qualified;

[0153] After receiving the image data uploaded by the terminal, the customized central control service calls the algorithm service to verify the facial points.

[0154] If the facial points in the image data are detected to be qualified, a message indicating that the image detection is qualified is sent to the terminal 300;

[0155] If it is detected that the facial points in the image data are unqualified, an image detection unqualified message is sent to the terminal, so that the terminal 300 prompts the user to upload again.

[0156] Among them, face point detection can be to use algorithms to detect whether the key points of the face are all within the specified area.

[0157] After receiving the image detection qualified message, the terminal 300 displays the online special effects page.

[0158] On the online special effects page, users can upload their original video or photo to server 400, using the original video or photo as their digital avatar. Alternatively, they can select a preferred special effects style, drag or click the effect intensity, and upload the post-effected video or photo to server 400, using the post-effected video or photo as their digital avatar. During the special effects creation process, users can touch the lower right corner of the effect image to compare the difference with the original image at any time. Image preloading is used during special effects creation, monitoring the loading progress of image resources, and setting image hierarchical relationships.

[0159] After the image data passes the facial point verification and is successfully uploaded to the server 400, the terminal 300 displays a timbre setting interface. The timbre setting interface may include at least one preset recommended timbre control and a custom timbre control;

[0160] The terminal 300 receives the user input instruction to select the preset recommended timbre control, sends the identifier corresponding to the preset recommended timbre to the server 400, and displays the digital human naming interface.

[0161] In some embodiments, after receiving an instruction from the user to select a custom tone control, the terminal 300 displays an audio recording selection interface, which may include adult controls and child controls.

[0162] In some embodiments, as shown in FIG12 , the timbre setting interface may include a Xiaowan control 111, a Xiaosheng control 112, and a custom timbre control 113. Upon receiving user input selecting the custom timbre control 113, an audio recording preparation interface is displayed, as shown in FIG13 . The audio recording selection interface may include recording notes 121, adult controls 122, and child controls 123. Upon receiving user input selecting adult controls 122 or child controls 123, the corresponding processes are entered. Upon receiving user input selecting the Xiaowan control 111, a digital human naming interface is displayed, as shown in FIG14 .

[0163] After receiving the user input instruction to select the adult control, the terminal 300 displays the ambient sound detection interface.

[0164] The terminal 300 collects the ambient sound of a preset duration and sends the user-recorded ambient sound to the server 400 .

[0165] The server 400 receives the ambient recorded sound uploaded by the terminal 300;

[0166] Check whether the recorded sound of the environment is qualified;

[0167] After receiving the ambient recorded sound uploaded by the terminal 300, the customized central control service calls the algorithm service to detect whether the ambient recorded sound is qualified.

[0168] The steps of detecting whether the ambient recorded sound is qualified may include: obtaining the noise value of the ambient recorded sound; determining whether the noise value exceeds a preset threshold; if the noise value exceeds the preset threshold, determining that the ambient recorded sound is unqualified; if the noise value does not exceed the preset threshold, determining that the ambient recorded sound is qualified.

[0169] If it is detected that the ambient sound recording is qualified, a message of the ambient sound being qualified and the target text required for the recorded audio is sent to the terminal 300;

[0170] If it is detected that the recorded ambient sound is unqualified, a message that the ambient sound is unqualified is sent to the terminal 300 so that the terminal 300 prompts the user to select a quiet space to re-record.

[0171] After receiving the environmental sound qualification message and the target text required for recording audio, the terminal 300 displays the target text; wherein, the target text can be selected as a text that reflects the user's voice characteristics.

[0172] The terminal 300 receives the audio of the user reading the target text and sends the audio to the server 400. The terminal 300 can send the audio data of the preset length to the server 400 after receiving it, so that the server 400 can send the recognition result back to the terminal 300 to achieve the effect of real-time recognition of the text being read.

[0173] The server 400 receives the audio of the user reading the target text;

[0174] Identify the user text corresponding to the audio;

[0175] Calculate the pass rate based on the target text and the user text;

[0176] The step of calculating the pass rate based on the target text and the user text may include: comparing the target text with the user text to obtain the number of correct characters in the user text; and determining the pass rate as the ratio of the number of correct characters to the number of characters in the target text.

[0177] Determine whether the qualified rate is less than the preset value;

[0178] If the pass rate is less than a preset value, a voice upload failure message is sent to the terminal 300, so that the terminal 300 prompts the user to re-record the audio of reading the target text;

[0179] In some embodiments, when real-time recognition is performed on the read text, the target text is compared with the user text to determine erroneous, over-read and missed texts, and the erroneous, over-read and missed texts are marked and sent to the terminal 300 so that the terminal 300 displays the erroneous, over-read and missed texts in different colors or fonts.

[0180] If the pass rate is not less than the preset value, a voice upload success message is sent to the terminal 300 so that the terminal 300 displays the next target text or voice recording completion information.

[0181] After a preset number of target texts are read and passed, the audio collection process ends and the terminal 300 displays the digital human naming interface.

[0182] The server 400 receives a preset number of audio data corresponding to target texts.

[0183] After receiving the user input instruction to select the child control, the terminal 300 also displays the ambient sound detection interface. The ambient sound detection steps are the same as when the adult control is selected.

[0184] If it is detected that the user's recorded environmental sound is qualified, a message of environmental sound qualification and the reading audio required for recording the audio are sent to the terminal 300.

[0185] Terminal 300 can automatically play the audio of reading, can listen to repeatedly.When receiving the instruction of user pressing recording key, start recording the audio of user following reading, and send audio to server 400.

[0186] The server 400 receives the audio of the user reading aloud;

[0187] Identify the user text corresponding to the audio;

[0188] Calculate the pass rate based on the target text corresponding to the leading reading audio and the user text corresponding to the follow-up reading audio;

[0189] Determine whether the pass rate is less than the preset value;

[0190] If the pass rate is less than the preset value, send a voice upload failure message to the terminal 300 so that the terminal 300 prompts the user to re-record the audio corresponding to the leading reading audio; when real-time recognizing the reading text, compare the target text with the user text to determine the error, extra reading, and missing reading texts, mark the error, extra reading, and missing reading texts and send them to the terminal 300 so that the terminal 300 displays the error, extra reading, and missing reading texts in different colors or fonts.

[0191] If the pass rate is not less than the preset value, send a voice upload success message to the terminal 300 so that the terminal 300 plays the next leading reading audio or voice recording completion information.

[0192] After receiving the voice recording completion, the terminal 300 displays the digital human naming interface.

[0193] In some embodiments, after the terminal 300 receives the instruction from the user to select the custom voice control, it can choose to upload a piece of audio data. After receiving the audio data, the server 400 detects the noise value. If the noise value exceeds the preset threshold, it sends an upload failure message to the terminal 300 so that the terminal 300 prompts the user to re-upload. If the noise value does not exceed the preset threshold, it sends an upload success message to the terminal 300 so that the terminal 300 displays the digital human naming interface.

[0194] After receiving the digital human name input by the user, the terminal 300 sends the digital human name to the server​​​​

[0196] After receiving the user's input indicating that the creation control 133 has been completed, the name of the digital person is sent to the server 400. After the server 400 detects that the name of the digital person sent by the user has been reviewed, it sends a message indicating a successful creation to the terminal 300. The terminal 300 may display a prompt indicating a successful creation. If the server 400 detects that the name of the digital person sent by the user has not been reviewed, it sends a message indicating a failed creation and the reason for the failure to the terminal 300. The terminal 300 may display a prompt indicating the reason for the failed creation and a new name.

[0197] Step S503: The server 400 determines the digital human image data based on the image data, and determines the digital human voice features based on the audio data.

[0198] Image preprocessing is performed on user-uploaded second-level videos or user photos to obtain digital human image data. Image preprocessing is the process of sorting each image and handing it over to the recognition module for identification. In image analysis, it is the processing performed on the input image before feature extraction, segmentation, and matching. The main purpose of image preprocessing is to eliminate irrelevant information in the image, restore useful real-world information, enhance the detectability of relevant information, and maximize data simplification, thereby improving the reliability of feature extraction, image segmentation, matching, and recognition. The embodiments of the present disclosure use relevant algorithms to achieve customized high-fidelity, high-definition interactive images.

[0199] In some embodiments, the digital human image data may include a 2D digital human image and facial key point coordinate information, and the facial key point coordinate information provides data support for the digital human voice key point driving.

[0200] In some embodiments, the digital human image data may include digital human parameters, such as 3D BS (Blend Shape) parameters. The digital human parameters provide offsets of key facial points based on a base model, allowing the display device 200 to render the digital human image based on the base model and the digital human parameters.

[0201] The voice cloning model is trained using user-uploaded audio data to generate timbre parameters that match the user's voice. During speech synthesis, the announcement text can be input into the voice cloning model embedded with the timbre parameters to generate an announcement that matches the user's voice.

[0202] To support digital human voice interaction, the disclosed embodiments add phoneme duration prediction to the general speech synthesis architecture, providing downstream digital human facial keypoint driving. To support digital human image customization, a multi-speaker speech synthesis model is used to implement small-sample timbre customization. Voice cloning is achieved by fine-tuning a small number of model parameters using 1-10 user speech samples.

[0203] The digital human image can be a real person image or a cartoon image, or you can choose to create a real person image and a cartoon image at the same time.

[0204] Upon receiving the image data uploaded by the terminal 300 (without facial point detection), the server 400 can notify the user to start training with their real-life or cartoon avatar. This means that real-life or cartoon avatar training and facial point detection are performed simultaneously. If facial point detection fails, real-life or cartoon avatar training is terminated. If facial point detection succeeds, the waiting time for digital human training can be shortened.

[0205] In some embodiments, the server 400 sends the trained real-life image and cartoon image to the terminal 300, so that the terminal 300 displays the digital human image and provides the user with the option to use it.

[0206] After receiving and displaying the trained real-person image, the terminal 300 can provide the user with operations such as beautifying the real-person image and adding special effects. It can also provide options such as making a cartoon image and re-recording a video, so that the user can get the digital human image he wants.

[0207] Step S504: the server 400 sends the digital human image data to the display device 200 associated with the terminal 300, so that the display device 200 displays the digital human image based on the digital human image data.

[0208] In some embodiments, after receiving the 2D digital human image, the digital human image may be directly displayed on the digital human selection interface.

[0209] In some embodiments, after receiving the digital human parameters, a digital human image is drawn based on the basic model and the digital human parameters, and the digital human image is displayed on the digital human selection interface.

[0210] In some embodiments, the server 400 may also send the digital human name corresponding to the digital human image data to the display device 200 associated with the terminal 300, so that the display device 200 displays the digital human name at the corresponding position of the digital human image.

[0211] In some embodiments, after receiving the digital human's name uploaded by the terminal 300, the server 400 sends the initial image and the digital human's name to the display device 200 and displays it on the digital human selection interface. The digital human is marked "training" and may also indicate the training time. In some embodiments, the digital human selection interface is shown in Figure 15. After training is completed, the server 400 sends the final image obtained through training to the display device 200 for updating the display.

[0212] In some embodiments, a target voice (e.g., a greeting) generated based on the digital human's voice characteristics can also be sent to the display device 200, so that when the user moves the focus to the control corresponding to the digital human, the voice corresponding to the digital human's timbre can be played. For example, in Figure 8, when the focus 75 moves to the Ding Ding control 72, the voice "Hello, I am Ding Ding" with the Ding Ding timbre is played.

[0213] In some embodiments, a target voice is generated based on the digital human's voice features, and a key point sequence is determined based on the target voice. Image data is synthesized based on the key point sequence and the digital human's image data, and the image data and the target voice are sent to the display device 200, where the display device 200 stores the data locally. The digital human control is displayed using the first frame (first parameter) or a specified frame (specified parameter) in the image data, or an image is drawn based on the first parameter or specified parameter in the image data. When the user moves the focus to the digital human control, the image and target voice are played.

[0214] In some embodiments, when the display device 200 displays the digital human selection interface, a user input instruction for managing the digital human is received;

[0215] In response to the user inputting an instruction to manage the digital human, the control display 260 displays a digital human management interface, which may include at least one deletion control, modification control, and disabling control corresponding to the digital human.

[0216] If a user input instruction to select a delete control is received, the relevant data corresponding to the digital person is deleted.

[0217] If a user input instruction to disable a control is received, the relevant data corresponding to the digital person is retained and marked as disabled.

[0218] If a user input instruction to select a modification control is received, the control display 260 is controlled to display a modification identification code. After the modification identification code is scanned through the terminal 300, the user video or photo can be re-uploaded on the terminal 300 to change the image of the digital human, and / or, the user audio can be re-uploaded on the terminal 300 to change the voice characteristics of the digital human, and / or, the name / wake-up word of the digital human can be changed on the terminal 300.

[0219] It should be noted that during the above-mentioned digital human customization process, the user can exit the customization process at any time. The target application of terminal 300 will record and cache the user's data in real time to the server. If the user returns midway, the target application will retrieve the previously recorded data from the server, providing convenience for the user to continue the operation and avoiding re-recording. If the user is not satisfied with the continued recording, they can choose to re-record at any time.

[0220] The embodiment of the present disclosure does not limit the order of video recording, audio recording, and naming of the digital person.

[0221] In some embodiments, a schematic diagram of digital human interaction is shown in Figure 16. Display device 200 displays a QR code. Terminal 300 scans the QR code and receives the user's recorded video and audio. Terminal 300 sends the recorded video and audio to server 400. Server 400 uses voice cloning and image preprocessing technologies to obtain customized data for the digital human, including the digital human's image and voice features. Server 400 then sends the digital human's image to terminal 300 and display device 200, respectively. Display device 200 displays the digital human image on the user interface.

[0222] In some embodiments, the display device 200 and the terminal 300 do not need to establish an association. The interface for adding a digital human in Figure 10 may also include a local upload control 92. Upon receiving a user input indicating the selection of the local upload control 92, the display device 200's camera is activated, and the camera captures the user's image data, or displays local videos and pictures. The user selects locally stored image data, which is uploaded to the server 400. The server 400 then performs facial point detection and generates digital human image data. The display device 200 then displays the digital human image based on the digital human image data sent by the server 400. Similarly, the display device 200's sound collector can also collect ambient sound, which is then sent to the server 400 for ambient sound detection. The display device 200's sound collector or the control device 100's voice collection function can also transmit audio of the user reading the target text to the server 400, which then generates the digital human's voice features.

[0223] In some embodiments, the present disclosure further improves some functions of the server 400. The server 400 performs the following steps, as shown in FIG17 .

[0224] Step S1601: receiving voice data input by the user sent by the display device 200;

[0225] After starting the digital human interaction program, the display device 200 receives voice data input by the user;

[0226] In some embodiments, the step of starting the digital human interaction program may include: when the display device 200 displays the user interface, receiving an instruction input by the user to select a control corresponding to the digital human application; wherein the user interface includes controls corresponding to the application installed by the display device 200; in response to the instruction input by the user to select a control corresponding to the digital human application, displaying the digital human entrance interface as shown in Figure 7.

[0227] In response to the user inputting an instruction to select the natural dialogue control 62, the digital human interaction program is started, waiting for the user to input voice data through the control device 100 or controlling the sound collector to start collecting the user's voice data. The natural dialogue may include a chat mode, that is, the user can chat with the digital human.

[0228] In some embodiments, the step of starting the digital human interaction program may include: receiving ambient voice data collected by a sound collector; when it is detected that the ambient voice data is greater than or equal to a preset volume or the duration of the ambient voice data sound signal is greater than or equal to a preset threshold, determining whether the ambient voice data includes a wake-up word corresponding to the digital human;

[0229] If the ambient voice data includes the wake-up word corresponding to the digital human, the digital human interaction program is started, the sound collector is controlled to start collecting the user's voice data, and a voice receiving box is displayed in a floating layer on the current user interface;

[0230] If the ambient voice data does not include the wake-up word corresponding to the digital human, the related operations of displaying the voice receiving box will not be performed.

[0231] In some embodiments, the digital human interaction program and the voice assistant can be installed simultaneously in the display device 200. Upon receiving a user instruction to set the digital human interaction program as the default interaction program, the digital human interaction program is set as the default interaction program. Received voice data can be sent to the digital human interaction program, which then sends the voice data to the server 400. The digital human interaction program can also receive voice data and send the voice data to the server 400.

[0232] In some embodiments, after the digital human interaction program is started, voice data input by the user pressing the voice key of the control device 100 is received.

[0233] The voice data collection starts after the user starts pressing the voice key of the control device 100 , and ends after the user stops pressing the voice key of the control device 100 .

[0234] In some embodiments, after the digital human interaction program is launched, when a voice receiving box is displayed in a floating layer on the current user interface, the sound collector is controlled to start collecting voice data input by the user. If no voice data is received for a long time, the digital human interaction program can be closed and the voice receiving box can be removed.

[0235] In some embodiments, the display device 200 receives voice data input by the user and sends the voice data and the digital human identifier selected by the user to the server 400. The digital human identifier is used to represent the image, voice characteristics, name, etc. of the digital human.

[0236] In some embodiments, after receiving the voice data input by the user, the display device 200 sends the voice data and the device identifier of the display device 200 to the server 400. The server 400 obtains the digital human identifier corresponding to the device identifier from the database. It should be noted that when the display device 200 detects that the user has changed the digital human of the display device 200, it will send the changed digital human identifier to the server 400, so that the server 400 will change the digital human identifier corresponding to the device identifier in the database to the modified digital human identifier. In the embodiment of the present disclosure, the user does not need to upload the digital human identifier every time, and can obtain it directly from the database.

[0237] In some embodiments, the user can select the digital human they want to use through the digital human image displayed on the digital human selection interface as shown in FIG8 .

[0238] In some embodiments, each created digital person has a unique digital person name, which can be set as a wake-up word, and the digital person selected by the user can be determined based on the wake-up word included in the ambient voice data.

[0239] In some embodiments, the voice data received by the display device 200 as user input is essentially streaming audio data. After receiving the voice data, the display device 200 sends the voice data to the sound processing module, which performs acoustic processing on the voice data. Acoustic processing may include sound source localization, denoising, and sound quality enhancement. Sound source localization is used to enhance or retain the signal of the target speaker when multiple speakers are speaking, suppress the signals of other speakers, track the speaker, and perform subsequent voice directional pickup. Denoising is used to remove ambient noise from the voice data. Sound quality enhancement is used to increase the voice intensity of the speaker when the voice intensity is low. The purpose of acoustic processing is to obtain a relatively clean and clear voice of the target speaker in the voice data. The acoustically processed voice data is sent to the server 400.

[0240] In some embodiments, after receiving voice data input by the user, the display device 200 directly sends it to the server 400. The server 400 performs acoustic processing on the voice data and sends the processed voice data to the semantic service. After performing speech recognition, semantic understanding, and other processing on the received voice data, the server 400 sends the processed voice data to the display device 200.

[0241] Step S1602: Generate a broadcast text based on the voice data;

[0242] After receiving the voice data, the semantic service of server 400 uses voice recognition technology to identify the text content corresponding to the voice data. It performs semantic understanding, business distribution, vertical domain analysis, and text generation on the text content to obtain the broadcast text.

[0243] Step S1603: Generate digital human data based on the broadcast text, the digital human voice features and the digital human image data;

[0244] In some embodiments, the semantic service of server 400 can send the broadcast text or semantic results to the display device 200, and the display device 200 completes the voice interaction transfer and connects to the push streaming central control service of server 400, that is, the display device initiates a request to the push streaming central control service of server 400, and the request carries the broadcast text or semantic results, and the push streaming central control service completes voice synthesis, key point prediction, image synthesis and live interaction, etc.

[0245] In some embodiments, the semantic service of server 400 can send the broadcast text directly to the streaming control service, which can perform speech synthesis, key point prediction, image synthesis, and live broadcast interaction.

[0246] In some embodiments, the digital human data may include digital human image data and broadcast voice. The streaming control service generates digital human data based on the broadcast text, digital human voice features, and digital human image data. This may include synthesizing the broadcast voice based on the voice features corresponding to the digital human identifier and the broadcast text; inputting the broadcast text into a trained human voice cloning model corresponding to the digital human identifier to generate the broadcast voice with the digital human's timbre. The broadcast voice is a sequence of audio frames.

[0247] The key point sequence is determined based on the broadcast speech. The broadcast speech undergoes data preprocessing, such as denoising, to obtain speech features. These features are input into the encoder to generate high-level semantic features, which are then input into the decoder. The decoder combines the actual joint point sequence to generate a predicted joint point sequence, which then generates the digital human's body movements.

[0248] synthesizing digital human image data according to the key point sequence and the digital human image data;

[0249] In some embodiments, a digital human image frame sequence is synthesized based on the key point sequence and the digital human image corresponding to the digital human identifier. Image synthesis is performed using an image synthesis service based on the predicted key point sequence and the digital human image data (digital human image) to obtain digital human data, namely, all image frame sequences and audio frame sequences.

[0250] In some embodiments, a digital human parameter sequence is generated based on a key point sequence and digital human image data (digital human parameters); wherein the digital human parameter sequence includes a sequence of parameters such as the digital human image, lip shape, expression, and movement. Based on the predicted key point sequence and digital human image data (digital human parameters), digital human data is obtained, namely, a sequence of all digital human parameter sequences and an audio frame sequence.

[0251] Step S1604: Send the digital human data to the display device 200, so that the display device 200 plays the digital human image and voice according to the digital human data.

[0252] In some embodiments, the streaming control service relies on the live broadcast channel to encode the image frame sequence and broadcast voice and push them to the live broadcast room to complete the digital human streaming.

[0253] In some embodiments, the live streaming data push process is shown in Figure 18. Terminal 300 sends a request to establish a live streaming channel, creates a live streaming channel room, and sends the live streaming channel room to the push streaming control service. The push streaming control service then sends the live streaming data, generated through steps such as speech synthesis, key point prediction, and image synthesis, to display device 200 via the live streaming channel in a live streaming manner, for display device 200 to play.

[0254] The streaming control service is an important part of the driving display and terminal presentation of the digital human. It is responsible for the driving and display of the virtual image, reflecting the customization and driving effect of the entire digital human.

[0255] The push-stream control service receives three types of display device requests: 1) restart. The push-stream control service interrupts the current video playback, re-applies for the room instance, verifies the validity and sensitivity of the customized image, records the instance status, creates a live broadcast room and publishes the broadcast, completing the live broadcast preparation actions; 2) query. The push-stream control service asynchronously processes the request content and performs actions such as speech synthesis, key point prediction, image synthesis, and live broadcast room push until the image frame group and audio frame group are pushed, the live broadcast is completed, the room is destroyed, and the instance is recycled; 3) stop. The push-stream control service interrupts the current video playback, destroys the room, and recycles the instance.

[0256] In order to ensure the real-time nature of digital human driving, live broadcast technology is used to synthesize digital human data in real time based on the received request content and push the stream to the live broadcast room, realizing instant playback on the playback end.

[0257] In addition, the streaming control service utilizes an instance pool mechanism. Each user with the same authentication information applies for a unique instance. The instance pool automatically recycles used instances to make them available for other devices. Instances that experience exceptions or have not been recycled for an extended period of time are automatically detected and destroyed by the instance pool, with new instances created to maintain a healthy pool of instances.

[0258] The display device 200 injects the received encoded image frame sequence and broadcast voice into the decoder for decoding, and synchronously plays the decoded image frames and broadcast voice, that is, the image and voice of the digital human.

[0259] In some embodiments, the server 400 sends the digital human parameter sequence and the broadcast voice to the display device 200. The display device 200 draws and renders the digital human image based on the digital human parameters and the basic model, and synchronously displays the drawn digital human image while playing the broadcast voice.

[0260] In some embodiments, after recognizing the voice data, server 400 also delivers, in addition to the digital human data, requested user interface data or media asset data related to the voice data. Display device 200 displays the user interface data delivered by server 400 and the digital human data at a designated location. In some embodiments, when a user enters "What's the weather like today?", the user interface of display device 200 is shown in Figure 19.

[0261] In some embodiments, the digital human image is displayed in a user interface layer.

[0262] In some embodiments, the digital human image is displayed in a floating layer above the user interface layer.

[0263] In some embodiments, the user interface layer is located above the video layer. The digital human image is displayed in a preset area of ​​the video layer, and a target area is drawn on the user interface layer. The target area is transparent, and the preset area and the target area overlap, so that the digital human image in the video layer can be displayed to the user.

[0264] In some embodiments, the digital human interaction sequence diagram is shown in Figure 20. After receiving the voice data, the display device 200 sends the voice data to the semantic service, and the semantic service sends the semantic result to the display device 200. The display device 200 initiates a request to the push-stream control service. After the push-stream control service responds, it generates image synthesis data through speech synthesis, key point prediction, and image synthesis services, and pushes the image synthesis data and audio data to the live broadcast room. The display device 200 can obtain live broadcast data from the live broadcast room. When the push queue is empty, the push-stream control service automatically ends the push and exits the live broadcast room. The display device 200 detects that there is no action timeout, ends the live broadcast, and exits the live broadcast room.

[0265] The disclosed embodiments support the provision of high-fidelity customization of general-purpose digital humans for enterprise and individual users with small sample sizes and low resource consumption, and offer a novel anthropomorphic intelligent interaction system based on the reproduction of digital human images and voices. Digital human images can include 2D real-life images, 2D cartoon images, and 3D lifelike images. Users scan a QR code through an application to enter the terminal customization process. They customize their own digital human image by collecting the user's second-level video information or selfie images, and customize their own voice by collecting 1 to 10 sentences of audio data. After customization is complete, the display device 200 allows users to select and switch between images and voices, using the selected image and voice to provide voice and text-based interactions. During the interaction, the display device 200 receives the user's request, and a response (broadcast text) is generated by perceptual and cognitive algorithm services based on semantic understanding, voice analysis, and empathy. The response is output in video and audio formats using the digital human image and voice. The audio and video data is generated using algorithms such as speech synthesis, face recognition, and image generation, and coordinated and forwarded to the target display device by the streaming control service, completing the interaction.

[0266] In some embodiments, the present disclosure further improves some functions of the server 400. The server 400 performs the following steps, as shown in FIG21.

[0267] Step S2001: receiving voice data input by a user sent by the display device 200;

[0268] Step S2002: Recognize the voice data and obtain the recognition result;

[0269] After receiving the voice data input by the user sent by the display device 200, the server 400 recognizes the text corresponding to the voice data using voice recognition technology.

[0270] Step S2003: determining whether the recognition result includes entity data, where the entity data may include a person's name and / or a media asset name;

[0271] After obtaining the recognition result, the semantic service of server 400 performs semantic understanding on the text content. In the semantic understanding process, the recognized text is segmented and annotated to obtain segmentation information, and it is determined whether the segmentation information includes entity data.

[0272] If the recognition result does not include entity data, the recognition result is processed through semantic understanding, service distribution, vertical domain analysis, and text generation to obtain the broadcast text. Digital human data is generated based on the broadcast text, the digital human's voice characteristics, and the digital human's image, and the digital human data is sent to the display device 200 so that the display device 200 can play the digital human data.

[0273] If the recognition result includes entity data, step S2004 is executed: media data corresponding to the recognition result and digital human data corresponding to the entity data are obtained. The digital human data includes the digital human's image data and broadcast voice, and the media data includes audio and video data or interface data. Audio and video data refers to at least one of audio data and video data.

[0274] If the recognition result includes entity data, the server 400 locates the domain and intent through vertical domain classification based on the word segmentation information, and obtains media asset data corresponding to the domain and intent.

[0275] Before receiving the voice data input by the user from the display device 200, the server 400 pre-processes and standardizes the facial image, body posture, and voice of the character, and then performs model training to generate a highly realistic digital human image model.

[0276] FIG22 shows a process for generating a digital human image model according to an embodiment of the present disclosure. As shown in FIG22 , the process may include the following steps:

[0277] Step S2101: generating a painting model corresponding to at least one character name;

[0278] The step of generating a painting model corresponding to at least one character name may include: obtaining a preset number of pictures corresponding to the character names;

[0279] There is a large amount of material online that corresponds to person names. We collected photos and videos of these names from various angles and used them as the raw dataset for training. We then preprocessed and annotated the images to extract key features of the digital human, such as facial expressions and posture. The purpose of preprocessing is to remove watermarks and other factors to make the person in the photo or video clearer. Annotation involves labeling the person in the photo.

[0280] The picture is input into the Wensheng picture model to obtain a painting model corresponding to the character name.

[0281] Using the large model of Wenshengtu (Stablediffusion), a LoRA model (a smaller painting model) corresponding to the name of a person is generated based on the clear angles, scenes, etc. of the collected photos of different people (10 to 20 photos).

[0282] Step S2102: Generate at least one action model corresponding to a media asset name;

[0283] The step of generating at least one action model corresponding to a media asset name includes:

[0284] Obtaining a preset amount of sample video data, and preprocessing and labeling the sample video data;

[0285] Obtain multiple sets of video data with different themes, each set containing multiple videos of the same theme. Preprocess and normalize these videos. Video data preprocessing includes video editing, noise removal, and annotation. Normalizing the video data involves adjusting the range of motion of the characters in the video data to a uniform standard. The purpose of preprocessing and normalization is to remove irrelevant information and unify the data to facilitate subsequent model training.

[0286] Training the action generation model using the labeled sample video data;

[0287] Using preprocessed and standardized video data, skeleton keypoints are annotated and a deep learning algorithm is used to train an action generation model to learn typical actions and action sequences in the video. During the training process, the model needs to be annotated multiple times to optimize the model's action realism.

[0288] The video data corresponding to the media asset name is input into the trained action generation model to generate an action model corresponding to the media asset name.

[0289] Step S2103: generating a speech synthesis model based on tone and rhythm corresponding to at least one character name;

[0290] In some embodiments, a preset number of sample audio data are obtained; wherein the sample audio data may include audio data corresponding to the character name and audio data corresponding to the media asset name;

[0291] Preprocess and label sample audio data;

[0292] The preprocessing of the audio data corresponding to the character name includes removing noise and labeling the character name.

[0293] The step of preprocessing the audio data corresponding to the media asset name may include:

[0294] 1) Audio processing: Process the audio data, such as separating the singing voice and the accompaniment. You can use audio processing software such as Audacity for processing.

[0295] 2) Song Analysis: Use audio processing software or music analysis tools, such as Sonic Visualizer, to analyze the singing voice and extract the pitch and rhythm information of the singing voice.

[0296] 3) Lyrics conversion: Use a lyrics conversion tool to convert the song's lyrics into text format. You can use an online lyrics conversion tool, such as an LRC (lyrics) file to text tool.

[0297] The audio data is a representative section of audio data in the entire song, and the audio data is labeled with its corresponding media asset name and the corresponding lyrics in the sample audio data.

[0298] The speech synthesis model is trained using the labeled sample audio data to obtain a speech synthesis model based on pitch and rhythm corresponding to the character name.

[0299] A text-to-speech (TTS) model is trained using a deep learning algorithm to learn the song's pitch and rhythmic information, as well as the character's timbre, and convert the lyrics into speech. During training, the model undergoes multiple iterations to continuously optimize its generation capabilities. The trained TTS model is then used to generate speech that matches the character's timbre based on pitch and rhythm.

[0300] In some embodiments, audio data corresponding to a preset number of character names is obtained, and a speech synthesis model is generated based on the audio data of the character using voice cloning technology. After inputting text data, the speech synthesis model can generate speech corresponding to the text data that matches the character's timbre.

[0301] Obtain audio data of a preset number of songs, and preprocess and annotate the audio data;

[0302] The annotated audio data is used to further train the speech synthesis model corresponding to the character to obtain a speech synthesis model based on tone and rhythm corresponding to the character's name.

[0303] In some embodiments, audio data of a preset number of songs is obtained, and the audio data is preprocessed and labeled. The audio data of the labeled songs is used to train a TTS model to obtain a speech synthesis model based on pitch and rhythm. After inputting text data, the speech synthesis model can generate speech with pitch and rhythm corresponding to the text data.

[0304] Audio data corresponding to a preset number of character names are obtained, and a speech synthesis model based on pitch and rhythm is further trained using the audio data corresponding to the character names to obtain a speech synthesis model based on pitch and rhythm corresponding to the character names.

[0305] Step S2104: constructing and training a conditional adversarial network;

[0306] Step S2105: Input the painting model, action model and speech synthesis model into the trained conditional adversarial network to obtain the digital human data to be stored.

[0307] The disclosed embodiment uses technologies such as conditional generative adversarial networks (GANs), variational autoencoders, and deep reinforcement learning to generate an integrated model. The specific steps of the integrated model are as follows:

[0308] 1) Construction of a Conditional Generative Adversarial Network: A conditional generative adversarial network (CGN) is constructed, consisting of two modules: a generator and a discriminator. The generator accepts as input the LoRA image model (painting model) corresponding to the character name, the action model corresponding to the media asset name, and the TTS model, and generates a complete digital human image model. The discriminator accepts as input the complete digital human image model and the real digital human image model, and evaluates the difference between the two.

[0309] 2) Model Training: Using a large number of LoRA image models corresponding to character names, action models corresponding to media asset names, and TTS models, we primarily need to annotate and adjust parameters for actions and sounds, and train the conditional generative adversarial network. During training, we continuously optimize the parameters of the generator and discriminator to achieve highly realistic and lifelike digital human image generation.

[0310] 3) Digital Human Image Model Generation: Use the trained conditional generative adversarial network to generate a complete digital human image model. Different digital human image model effects can be obtained by inputting different LoRA image models corresponding to different character names, action models corresponding to media asset names, and TTS models.

[0311] 4) Optimization and Adjustment: Based on the actual needs of the digital human image, the digital human image model is optimized and adjusted to improve the realism and fidelity of the digital human image. For example, the digital human image model can be optimized for facial expressions and body posture to achieve a more realistic and lifelike digital human image effect.

[0312] 5) Rendering and Animation: Render and animate the digital human to achieve a more realistic and lifelike digital human image. Use rendering algorithms such as Nerf to render the digital human, and use animation software to animate the digital human.

[0313] In some embodiments, the digital human image model storage step may include: performing feature annotation on the digital human data to be stored and storing it in the server 400; performing feature annotation on the digital human data to be stored and storing it in the cloud.

[0314] In some embodiments, the feature structure is as follows: [person name, media asset name, popularity]. The popularity is the number of training data, and the number of training data found on the network is also a reflection of the popularity of the person and media asset.

[0315] In some embodiments, the storage feature structure is as follows: [character name (including basic attributes such as gender, age, etc.), media asset name, popularity].

[0316] In some embodiments, all digital human data to be stored may be feature-labeled and then stored in the server 400 .

[0317] In some embodiments, part of the digital human data to be stored (highly popular digital human data to be stored) may be feature-labeled and then stored in the server 400 .

[0318] The step of annotating the stored digital human data with features and storing it on the server may include annotating the stored digital human data with character information, media asset name, and popularity. The character information may include basic attributes such as the character's name, gender, and age. Basic attributes such as gender and age facilitate filtering user requests. For example, if a user requests videos of female singers between the ages of 20 and 40, if the age cannot be determined from the name alone, basic character attributes may be further configured.

[0319] Obtaining a first popularity and a second popularity; wherein the first popularity is the highest popularity corresponding to the character name in the stored digital human data, and the second popularity is the highest popularity corresponding to the media asset name in the stored digital human data;

[0320] Determining whether the popularity of the digital human data to be stored is less than a first popularity;

[0321] If the popularity of the digital human data to be stored is not less than the first popularity, the marked digital human data to be stored is stored in the server 400.

[0322] If the popularity of the digital human data to be stored is less than the first popularity, determining whether the popularity of the digital human data to be stored is less than the second popularity;

[0323] If the popularity of the digital human data to be stored is not less than the second popularity, the marked digital human data to be stored is stored in the server 400.

[0324] If the popularity of the digital human data to be stored is less than the second popularity, the marked digital human data to be stored is not stored in the server 400 .

[0325] In some embodiments, the annotated information for the digital human data to be stored is that the character name is Xiao A, the video name is XX, and the popularity is 3000. If the highest popularity corresponding to Xiao A (character Xiao A - video YY) in the stored digital human data is 4000, and the highest popularity corresponding to XX (character Xiao B - video XX) in the stored digital human data is 4000, then the digital human data to be stored does not need to be stored on the server 400. If the highest popularity corresponding to Xiao A (character Xiao A - video YY) in the stored digital human data is 2000, or the highest popularity corresponding to XX (character Xiao B - video XX) in the stored digital human data is 2000, then the digital human data to be stored does not need to be stored on the server 400.

[0326] In some embodiments, the digital human data stored in server 400 can be regularly updated. This can include periodically acquiring a large amount of the latest data to generate the digital human data. Updating the stored digital human data can also include recording the digital human's generation time. If the current time exceeds the generation time by a certain amount, the popularity of the digital human data can be appropriately reduced to prevent previously popular characters or videos from constantly occupying the digital human data resources, preventing the recently updated and more popular digital human data from being pushed to users.

[0327] In some embodiments, if the recognition result includes entity data, the step of obtaining digital human data corresponding to the entity data may include: if the recognition result includes a person's name, determining whether there is digital human data with features marked as corresponding to the person's name in the stored digital human data;

[0328] If no digital human data corresponding to the character name is present in the stored digital human data, the recognition result is processed through semantic understanding, service distribution, vertical domain analysis, and text generation to obtain the broadcast text. Digital human data is generated based on the broadcast text, the selected digital human's voice features, and the digital human's image, and the digital human data is sent to the display device 200 so that the display device 200 can play the digital human data.

[0329] If the stored digital human data contains digital human data with features labeled as corresponding to the character name, the digital human data with features labeled as corresponding to the character name in the stored digital human data is obtained, wherein the digital human data is video data with the character name corresponding to the character's image and voice.

[0330] In some embodiments, voice data of a user inputting "I want to watch Xiao A's video" is received. After the voice data is recognized and segmented, it is determined that the recognition result includes the entity data of Xiao A. The digital human data marked as corresponding to Xiao A is obtained on the server 400, and at the same time, the media resource data corresponding to Xiao A is obtained.

[0331] In some embodiments, when there is more than one digital human data corresponding to the character name, the step of obtaining the digital human data corresponding to the character name may include: obtaining the digital human data with the highest popularity among the stored digital human data with the feature labelled as corresponding to the character name.

[0332] In some embodiments, voice data of a user inputting "I want to watch Xiao A's video" is received. After the voice data is recognized and segmented, it is determined that the recognition result includes the entity data of Xiao A. The server 400 may include that the highest popularity corresponding to Xiao A (character Xiao A-video YY) is 4000, and the highest popularity corresponding to video XX (character Xiao A-video XX) is 3000. Then, the digital human data corresponding to the character Xiao A-video YY is obtained (the image and voice are Xiao A, and the movements and lyrics are video YY). At the same time, the media data corresponding to Xiao A is obtained.

[0333] In some embodiments, if the recognition result includes entity data, the step of obtaining the digital human data corresponding to the entity data may include: if the recognition result includes a media asset name, determining whether there is digital human data with features marked as corresponding to the media asset name in the stored digital human data; if there is no digital human data with features marked as corresponding to the media asset name in the stored digital human data, generating digital human data based on the broadcast text, the selected digital human voice features and the digital human image.

[0334] If the stored digital human data has digital human data with features marked as corresponding to the media asset name, the digital human data with features marked as corresponding to the media asset name in the stored digital human data is obtained, wherein the digital human data is the video data corresponding to the media asset name.

[0335] In some embodiments, voice data of a user inputting "I want to watch XX video" is received. After the voice data is recognized and segmented, it is determined that the recognition result includes the entity data XX. The digital human data corresponding to XX marked as XX is obtained on the server 400, and at the same time, the media resource data corresponding to XX is obtained.

[0336] In some embodiments, when there is more than one digital human data corresponding to the media asset name, the step of obtaining the digital human data corresponding to the media asset name may include: obtaining the digital human data with the highest popularity among the stored digital human data with the feature marked as corresponding to the media asset name.

[0337] In some embodiments, voice data of a user inputting "I want to watch XX video" is received. After recognizing and segmenting the voice data, it is determined that the recognition result includes the entity data XX. The server 400 may include that the highest popularity corresponding to Xiao A (character Xiao A-video YY) is 1000, and the highest popularity corresponding to video XX (character Xiao B-video XX) is 3000. Then, the digital human data corresponding to character Xiao B-video XX (the image and voice are Xiao B, and the lyrics and actions are video XX) marked as corresponding are obtained, and at the same time, the media data corresponding to XX is obtained.

[0338] In some embodiments, if the recognition result includes entity data, the step of obtaining the digital human data corresponding to the entity data includes:

[0339] If the recognition result includes a person's name and a media asset name, determine whether there is digital human data with features marked as corresponding to the media asset name in the stored digital human data;

[0340] If there is no digital human data with features marked as corresponding to the media asset name in the stored digital human data, digital human data is generated based on the broadcast text, the selected digital human voice features and the digital human image.

[0341] If there is digital human data with a feature label corresponding to the media asset name in the stored digital human data, determine whether there is digital human data with a feature label corresponding to the character name in the stored digital human data;

[0342] If there is no digital human data corresponding to the character name with features marked as the character name in the stored digital human data, the digital human data corresponding to the media asset name with features marked as the media asset name in the stored digital human data and an error message can be obtained, or digital human data can be generated based on the broadcast text, the selected digital human voice features and the digital human image.

[0343] If there is digital human data with a feature label corresponding to the character name in the stored digital human data, determine whether the character name and the media asset name match the feature labels in the stored digital human data;

[0344] If the character name and the media asset name match the feature annotations in the stored digital human data, then obtain the digital human data corresponding to the character name and the media asset name;

[0345] If the character name and the media asset name do not match in the stored digital human data feature annotations, the painting model corresponding to the media asset name is replaced with the painting model corresponding to the character name, and the voice data corresponding to the media asset name is replaced with the voice data corresponding to the character name, to generate replacement digital human data;

[0346] Determine to replace the digital human data with the digital human data corresponding to the character name and media asset name.

[0347] In some embodiments, voice data of a user inputting "I want to watch Xiao A's XX video" is received. After the voice data is recognized and segmented, it is determined that the recognition result includes two entity data, Xiao A and XX. The character corresponding to video XX in server 400 is marked as Xiao B, that is, only the digital human data of the character Xiao B-video XX is stored, the LoRA image model of video XX is replaced with the image of Xiao A, and the voice is replaced with the TTS model of Xiao A to generate the replaced digital human data. At the same time, the media data corresponding to Xiao A's XX video is obtained.

[0348] In some embodiments, the character name can be an individual name or a combined name. When the character name is a combined name, multiple character images can be embodied in one digital human data.

[0349] Step S2005: Sending the digital human data and media data to the display device 200, so that the display device 200 plays the audio and video data or displays the interface data, and plays the digital human's image and voice according to the digital human data.

[0350] In some embodiments, the digital human image data is a sequence of image frames, and the server 400 sends the sequence of image frames and the announcement voice to the display device 200 in a live streaming manner. The display device 200 displays the image corresponding to the image frame and plays the announcement voice.

[0351] In some embodiments, the digital human image data is a digital human parameter sequence, and the server 400 sends the digital human parameter sequence and the announcement voice to the display device 200. The display device 200 displays the digital human image and plays the announcement voice based on the digital human parameters and the basic model.

[0352] If the media asset data is interface data, the display device 200 displays the user interface based on the interface data while playing the image and voice of the digital human according to the digital human data.

[0353] If the media asset data is audio and video data, the display device 200 plays the image and voice of the digital human according to the digital human data before playing the audio and video data.

[0354] In some embodiments, after receiving user input of "I want to watch Xiao A's XX video," Xiao A's XX video data and digital human data are transmitted to the display device 200. The display device 200 uses Xiao A's corresponding digital human image, XX video movements, and Xiao A's singing voice to make an entertaining announcement: "xxx, xxxxx" (singing voice), Xiao A brings you XX video, as shown in FIG23. After the announcement is completed, the XX video data is displayed.

[0355] After collecting photos and video information of celebrities or popular online memes from different angles, the disclosed embodiment generates a basic image of the character and a specific action image, then uses AIGC (Artificial Intelligence Generated Content) to beautify the character's image, drives the image and action based on each key point to generate a complete video image, adds specific broadcast synthesis for personalized voice broadcast display, and displays the digital human image, action and sound in three dimensions in the search scene of the display device 200, thereby increasing the connection between search and voice feedback and improving the fun of voice interaction.

[0356] In some embodiments, the present disclosure further improves some functions of the server 400. The server 400 performs the following steps, as shown in FIG24 .

[0357] Step S2301: receiving voice data input by the user sent by the display device 200;

[0358] Step S2302: Recognize voice data to obtain voice text;

[0359] After receiving the voice data input by the user sent by the display device 200, the server 400 recognizes the voice text corresponding to the voice data using voice recognition technology.

[0360] Step S2303: Perform semantic understanding on the speech text to obtain the domain intent corresponding to the speech data;

[0361] The steps of semantically understanding the speech text to obtain the domain intent corresponding to the speech data include:

[0362] 1) Preprocess the speech text. Preprocessing includes filtering sensitive words, formatting text, and normalizing word segmentation.

[0363] 2) Calling the three-category model service to determine the specific type of the pre-processed speech text, that is, determining whether the pre-processed speech text belongs to the chat type, the question and answer (QA) type, or the task type. There is no restriction on the three-category algorithm.

[0364] 3) If the specific type of the pre-processed voice text is determined to be a chat type, the chat service is called to parse the chat intent, that is, to determine that the domain and intent corresponding to the voice data are chat.

[0365] 4) If the specific type of the pre-processed speech text is determined to be a question-and-answer type, the question-and-answer service is called to determine whether a question-and-answer pair is found;

[0366] If a question-answer pair is found, the domain and intent corresponding to the voice data are determined to be question-answering;

[0367] If no question-answer pair is found, the chat service is called to parse the chat intent, that is, to determine whether the domain and intent corresponding to the voice data are chat.

[0368] 5) If the pre-processed speech text is determined to be a task type, the intent is parsed, and the strong rule algorithm is invoked to determine whether a strong rule is matched. The strong rule algorithm includes regular expression matching and ABNF (augmented Backus-Naur Form) rule matching.

[0369] If a strong rule is hit, the corresponding field, intent, and slot are returned.

[0370] If no strong rule is matched, the method is resolved. The multi-classification model service is called to obtain the corresponding domain. The slots and grammar are parsed within the corresponding domain to match the corresponding intent. The domain, intent, and slots are then output.

[0371] Step S2304: determining the broadcast voice based on the domain intent, and determining the digital human image parameters based on the domain intent, the digital human image parameters being used to generate the digital human image and / or generate the digital human's movements;

[0372] The steps of determining the broadcast voice based on the domain intent include:

[0373] The broadcast text is determined based on the domain intent; among them, different business systems are called according to the domain intent to obtain the business results, namely the broadcast text.

[0374] The speech synthesis technology is used to generate the corresponding broadcast voice of the broadcast text. The broadcast voice is synthesized based on the voice characteristics of the digital human selected by the user and the broadcast text.

[0375] The steps for determining the parameters of a digital human image based on domain intent include:

[0376] The digital human image identifier corresponding to the domain intent is searched in the digital human image mapping table; wherein the digital human image mapping table is used to represent the correspondence between the domain intent and the digital human image identifier.

[0377] In some embodiments, the digital human image mapping table is shown in Table 1.

[0378] Table 1

[0379] The digital human image parameters corresponding to the digital human image identifier are searched in the digital human definition table. The digital human definition table is used to represent the correspondence between the digital human image identifier and the digital human image parameters. The digital human image parameters include decoration parameters and action parameters. Decoration parameters include digital human resource parameters, clothing resource parameters, hair resource parameters, prop resource parameters, makeup resource parameters, and special effects resource parameters. Clothing resource parameters include top resource parameters, bottom resource parameters, shoe resource parameters, and accessory resource parameters. Action parameters include arm swing angle, knee flexion angle, facial expression parameters, and the like.

[0380] In some embodiments, the digital human definition table is shown in Table 2.

[0381] Table 2

[0382] Different digital human images can be formed based on different clothing, hair, accessories, shoes and props.

[0383] Step S2305: Generate digital human data based on the digital human image parameters and the broadcast voice;

[0384] In some embodiments, the digital human image can be determined by the digital human resource identifier in the digital human image parameters. The digital human resource identifier is used to identify the selected basic model, or the basic model and basic parameters. The basic parameters are used to characterize the feature offset of key points of the face, so that the digital human image can be customized.

[0385] In some embodiments, the digital human image can be determined by a digital human identifier uploaded by the display device 200, where the digital human identifier is the digital human identifier corresponding to the customized digital human selected by the user.

[0386] In some embodiments, the digital human model can use Unity's digital human model. Unity's digital human model is usually driven by action parameters. Unity's digital human model is mainly implemented through Unity's animation system, especially Animator Controller and Blend Trees. Among them, Animator Controller is the core part of Unity's animation system, allowing the creation and management of animation states and transitions. Action parameters (such as speed, direction, whether to jump, etc.) can be defined in Animator Controller, and then the playback of the animation can be controlled according to these parameters. Blend Trees is an important feature of Animator Controller, which allows different animations to be mixed and transitioned according to action parameters. For example, create a Blend Tree to mix walking and running animations according to speed parameters. In this way, very complex and smooth animation effects can be created. For example, you can create a digital human model that will naturally transition from walking to running when the speed parameters are changed.

[0387] In some embodiments, the step of generating digital human data based on digital human image parameters and broadcast voice includes:

[0388] Inputting the digital human image parameters and the broadcast voice into the digital human driving system generates digital human data; the digital human data includes digital human decoration parameters, motion parameters, lip shape parameters, and the broadcast voice. When input into the digital human driving system, the lip shape parameters can be derived using the digital human lip shape driving algorithm based on the broadcast voice. When input into the digital human driving system, the digital human's specific image parameters can also be derived based on the digital human decoration parameters. The digital human data then includes the final digital human image parameter sequence, motion parameter sequence, lip shape parameter sequence, and the broadcast voice.

[0389] The digital human's lip shape driving algorithm is mainly used to synchronize the character's lip shape with the speech it produces, so that the character's lip shape movements match the pronunciation, increasing the character's realism and vividness.

[0390] In some embodiments, the lip-shaping algorithm is a rule-based approach. This approach primarily creates a set of predefined lip-shaping rules based on speech characteristics, such as phonemes and syllables. When speech is input, the corresponding lip-shaping rules are generated based on these rules.

[0391] In some embodiments, the lip-activated algorithm is based on a data-driven approach. Data-driven approaches primarily use machine learning algorithms to learn a model from a large amount of speech and lip movement data, and then use this model to predict lip movements for new speech. Common machine learning algorithms include deep learning and support vector machines (SVMs).

[0392] In some embodiments, the lip shape driving algorithm is a hybrid method that combines rule-based methods and data-driven methods, taking advantage of both the clarity of rules and the flexibility of data-driven methods.

[0393] In some embodiments, the step of generating digital human data based on digital human image parameters and broadcast voice includes:

[0394] Predict key point sequences based on broadcast speech;

[0395] Synthesize a digital human image frame sequence based on the predicted key point sequence, the digital human image selected by the user, and the digital human image parameters;

[0396] Digital human data refers to the live audio and video data of the digital human, namely, the digital human image frame sequence and broadcast voice.

[0397] Step S2306: Send the digital human data to the display device 200, so that the display device 200 plays the digital human image and voice according to the digital human data.

[0398] In some embodiments, when a Unity digital human model is selected, the digital human decoration parameters (or the digital human's final image parameters), action parameters, lip shape parameters and broadcast voice are sent to the display device 200. The display device 200 can draw the image of the Unity digital human model based on the digital human decoration parameters (or the digital human's final image parameters), and use the action parameters and lip shape parameters to drive the digital human model to make corresponding action expressions when playing the broadcast voice.

[0399] In some embodiments, the digital human data (digital human image data and broadcast voice) is sent to the display device 200 by live streaming. The display device 200 displays the digital human image based on the digital human image data and plays the broadcast voice.

[0400] In some embodiments, when the domain intent is determined to be music, a prop with headphones can be configured on the digital human avatar, as shown in Figure 25. When the domain intent is determined to be a football game, the digital human avatar can be dressed in a jersey, the prop can be a football, and a kicking action can be configured, as shown in Figure 26.

[0401] In some embodiments, after receiving voice data or obtaining voice text input from the display device 200, the server 400 further determines the user emotion type corresponding to the voice data. User emotion types are divided into three categories: Optimistic (like, happy, praise, and thankful), Pessimistic (angry, disgusting, fearful, and sad), and Neutral.

[0402] Emotion recognition technology is a technology that identifies and understands human emotions by analyzing language, voice, facial expressions, body language, and other information. It can help computer systems better understand and respond to human emotions, thereby achieving a more intelligent and humanized interactive experience.

[0403] In some embodiments, after receiving the voice data input by the user sent by the display device 200, the step of determining the user emotion type corresponding to the voice data includes:

[0404] Determine the user emotion type corresponding to the voice data based on the voice data.

[0405] The disclosed embodiments primarily identify the speaker's emotional state by analyzing voice data, including intonation, audio features, and speech content. For example, by analyzing voice data features such as pitch, volume, and speaking rate, it is possible to determine whether the speaker is angry, happy, sad, or neutral.

[0406] In some embodiments, after obtaining the voice text, the step of determining the user emotion type corresponding to the voice data includes:

[0407] Determine the user emotion type corresponding to the voice data based on the voice text.

[0408] The present disclosure identifies the user's emotional state by analyzing the vocabulary, grammar, and semantics of the voice text. For example, by analyzing the emotional vocabulary, emotional intensity, and emotional polarity in the voice text, it can be determined whether the user is positive, negative, or neutral.

[0409] In some embodiments, the step of determining the user emotion type corresponding to the voice data includes:

[0410] When receiving the voice data input by the user sent by the display device 200, the user video collected and uploaded by the display device 200 is also received, and the user video includes the user's facial image;

[0411] After receiving the digital human's wake-up voice, the display device 200 activates its image collector and simultaneously collects user video data while receiving the user's voice input. After sending the user video data to the server 400, if the server 400 detects a facial image in the user's video, it analyzes the user's facial image. If no facial image is detected in the user's video, the user's emotion type can be determined to be neutral.

[0412] Analyze the user's facial image to determine the user's emotion type corresponding to the voice data.

[0413] The disclosed embodiments identify a person's emotional state by analyzing facial expression features in facial images or videos. For example, by analyzing the movement and changes in facial features such as the eyes, eyebrows, and mouth, it is possible to determine whether a person's emotional state is anger, happiness, sadness, or surprise.

[0414] In some embodiments, the step of determining the user emotion type corresponding to the voice data includes:

[0415] When receiving the voice data input by the user sent by the display device 200, the user's physiological signals collected and uploaded by the display device are also received. The user's physiological signals include heart rate, skin conductance package and / or brain wave;

[0416] In some embodiments, after receiving the wake-up voice of the digital human, the display device 200 turns on the infrared camera of the display device 200 and collects the user's body temperature while receiving the user's input voice data.

[0417] In some embodiments, while receiving user input voice data, the display device 200 also obtains heart rate and other information collected by a smart device such as a wristband associated with the display device 200. The smart device must be within a certain distance from the display device 200. If the server 400 does not receive the user's physiological signals uploaded by the display device, the user's emotion type can be determined to be neutral.

[0418] Determine the user emotion type corresponding to the voice data based on the user's physiological signals.

[0419] The disclosed embodiments identify a person's emotional state by analyzing physiological signals of the human body, such as heart rate, skin conductance, brain waves, etc. For example, by monitoring changes in heart rate, it is possible to determine whether a person is nervous, relaxed, or excited.

[0420] The step of determining parameters of the digital human image based on the domain intent includes:

[0421] Determine the digital human image parameters based on the user's emotional type and domain intention.

[0422] The step of determining the parameters of the digital human image based on the user's emotion type and the domain intention includes:

[0423] Searching for the digital human image identifier corresponding to the user's emotion type and domain intent in the digital human image mapping table, where the digital human image mapping table is used to represent the correspondence between domain intent, user emotion type, and digital human image identifier;

[0424] In some embodiments, the digital human image mapping table is shown in Table 3.

[0425] Table 3

[0426] The digital human image parameters corresponding to the digital human image identifier are searched in the digital human definition table. The digital human definition table is used to represent the correspondence between the digital human image identifier and the digital human image parameters. The digital human image parameters include decoration parameters and action parameters.

[0427] In some embodiments, the digital human definition table is shown in Table 4.

[0428] Table 4

[0429] When the intention is in the same field, digital human images for different users and different emotions can be formed by changing the color of clothing.

[0430] In some embodiments, if the user's emotion type in the chat mode is joy, a joyful digital human image is used, as shown in FIG27 ; if the user's emotion type is like, a like image is used, as shown in FIG28 .

[0431] In some embodiments, when the domain intent is weather search, if the user's emotion type is recognized as joy, the display device 200 displays a digital human wearing a weatherman's uniform in bright colors (such as red or yellow); if the user's emotion type is recognized as sadness, the display device 200 displays a digital human wearing a weatherman's uniform in dark colors (such as dark blue or gray).

[0432] In some embodiments, the server 400 may also execute: receiving voice data input by a user sent by a display device; recognizing the voice data to obtain voice text; determining the user emotion type corresponding to the voice data; performing semantic understanding on the voice text to obtain the domain intent corresponding to the voice data; determining the broadcast voice based on the domain intent, and determining the digital human image parameters based on the user emotion type; generating digital human data based on the digital human image parameters and the broadcast voice; and sending the digital human data to the display device so that the display device plays the digital human data.

[0433] The disclosed embodiments can adapt the digital human's clothing, props, and body movements to the current display device scene (domain intent), enhancing the fun interactive experience and emotional resonance. Simultaneously, the digital human's clothing color, expression, and body movements can be changed in a timely manner according to the user's emotional tendencies to enhance the atmosphere and soothe negative emotions.

[0434] In some embodiments, the present disclosure further improves some functions of the server 400. The server 400 performs the following steps, as shown in FIG29 .

[0435] Step S2801: receiving voice data input by the user sent by the display device 200;

[0436] Step S2802: Inputting the speech data into the emotional speech model to obtain the emotion type and emotion intensity;

[0437] Among them, the speech emotion model is trained based on sample speech data from different groups of people for multiple semantic scenarios.

[0438] We collect sample speech data from groups of varying ages, genders, speaking speeds, timbre, dialects, and other dimensions across multiple semantic scenarios, and label the sample speech data accordingly. This sample speech data is then fed into an emotional speech model for training, and the relevant model parameters are adjusted. With more training sample speech data, we can obtain stable and accurate emotion types and intensities.

[0439] In some embodiments, as shown in FIG30 , after inputting speech data into the emotional speech model, user speech features, semantic scenarios, and speech segment sequences are obtained. A user speech feature vector, a semantic scenario feature vector, a speech sequence feature vector, and an emotional feature vector are then determined. Feature processing is then performed using a multi-level neural network, and feature classification is performed using a Soft-Max classifier to obtain emotional classification and emotional intensity of the speech data.

[0440] FIG31 shows a specific process of inputting speech data into the emotional speech model in step S2802 to obtain the emotion type and emotion intensity. As shown in FIG31 , the process includes the following steps:

[0441] Step S3001: Recognize voice data to obtain voice text and user voice features;

[0442] The speech recognition service uses speech recognition technology (Automatic Speech Recognition, ASR) to parse speech data into speech text. Speech text refers to the text content expressed by the user's voice.

[0443] Using voiceprint recognition technology to analyze information such as the voiceprint, rhythm, intensity, and characteristics of voice data to determine the user's voice features. The user's voice features include age, gender, speech rate, timbre, dialect, etc. Among them, the age can be children, adults, and the elderly. The speech rate can be fast, medium, and slow. The dialect can be Minnan dialect, Beijing dialect, Northeast dialect, etc.

[0444] Step S3002: Conduct semantic understanding on the voice text to obtain the semantic scenario corresponding to the voice data;

[0445] Among them, the step of conducting semantic understanding on the voice text to obtain the semantic scenario corresponding to the voice data includes:

[0446] Perform word segmentation and annotation processing on the voice text to obtain word segmentation information;

[0447] In some embodiments, the voice text is "Liu Dehua's songs". Performing word segmentation and annotation processing on "Liu Dehua's songs" obtains the word segmentation information as [{Liu Dehua-Liu Dehua[actor-1.0, singer-0.8, roleFeeble-1.0, officialAccount-1.0]}, {of-the[funcwordStructuralParticle-1.0]}, {songs-songs[musicKey-1.0]}].

[0448] Perform syntactic analysis and semantic analysis on the word segmentation information to obtain slot information;

[0449] In some embodiments, performing syntactic analysis and semantic analysis on the word segmentation information, the central word is obtained as "songs", the modifier is "Liu Dehua", and the relationship is an adjective modification relationship. In semantic analysis, it is known that there is a strong semantic relationship between songs musicKey and singer. Therefore, the parsed semantic slot result is: integrated word segmentation information: [{Liu Dehua-Liu Dehua[singer-1.0]}, {songs-songs[musicKey-1.0]}].

[0450] Locate the semantic scenario corresponding to the slot information through vertical domain classification; among them, the semantic scenario can be technically called the domain intention;

[0451] The central control system combines various business scores, obtains the optimal vertical domain business, and assigns it to the specific vertical domain business.

[0452] In some embodiments, the music search intent is located in the music field through vertical domain classification. The central control intent set only contains MUSIC_TOPIC (music theme), and the obtained score is 0.9999393, score:{topicSet=[MUSIC_TOPIC], 'query':['Andy Lau's songs'], 'task':0.9999393}, so the optimal service is music.

[0453] Step S3003: converting the user speech feature into a user speech feature vector;

[0454] The group characteristics are converted into feature vector representation and recorded as user feature vector.

[0455] Step S3004: converting the semantic scene into a semantic scene feature vector;

[0456] The semantic scene is represented by a feature vector, which is recorded as the semantic scene feature vector.

[0457] Step S3005: Divide the speech data into frames to obtain at least one speech segment sequence;

[0458] Step S3006: determining a speech sequence feature vector and an emotion feature vector based on the speech segment sequence;

[0459] In some embodiments, the step of determining a speech sequence feature vector and an emotion feature vector based on a speech segment sequence includes:

[0460] Perform feature extraction on the speech segment sequence to obtain a speech sequence feature vector;

[0461] Based on the Mel spectrum feature extraction technology, the emotional feature vector corresponding to the speech segment sequence is obtained.

[0462] In some embodiments, text sentiment analysis technology is used to analyze the input voice text to determine the emotional state to be expressed. It can identify emotional vocabulary, emotional intensity, and emotional tendency through natural language processing and emotion recognition algorithms.

[0463] Step S3007: Inputting the user speech feature vector, semantic scene feature vector, speech sequence feature vector and emotion feature vector into a multi-level neural network to obtain an emotion speech vector;

[0464] Among them, the multi-level neural network includes a two-dimensional convolutional network, a recurrent neural network and two fully connected networks. The parameters of the multi-level neural network have been determined after the training is completed.

[0465] Convolutional neural networks (CNNs) are a type of feedforward neural network with a deep structure that incorporates convolutional computations. They are a representative algorithm for deep learning. CNNs possess the ability to learn representations and perform translation-invariant classification of input information based on their hierarchical structure.

[0466] A recurrent neural network is a type of recurrent neural network that takes sequence data as input, performs recursion in the direction of sequence evolution, and all nodes (recurrent units) are connected in a chain.

[0467] A fully connected neural network, also known as a multilayer perceptron, is the most fundamental artificial neural network structure. In a fully connected neural network, each neuron is connected to all neurons in the previous and next layers, forming a densely connected structure. Fully connected neural networks are capable of learning complex features of input data and performing tasks such as classification and regression.

[0468] Step S3008: Determine the emotion type and emotion intensity based on the emotion speech vector.

[0469] The emotional speech vector is processed by a soft-max (normalized exponential function) classifier to obtain the emotion classification and emotion intensity.

[0470] The disclosed embodiments combine semantic scenarios, the user's gender and age characteristics, and the emotional characteristics of the user's voice to comprehensively output emotional intervention in speech synthesis, thereby making the voice interaction process more natural, enhancing the personality characteristics of the voice assistant, and improving the user's voice interaction experience.

[0471] In some embodiments, by inputting voice data into an emotional speech model to obtain the emotion type and emotion intensity, the influence of the emotion of the user input voice data on the broadcast voice emotion can also be ignored. For example, as shown in Figure 32, after the voice data is input into the emotional speech model, the user voice features and semantic scene are obtained. Then, the user voice feature vector and the semantic scene feature vector are determined. Next, feature processing is performed through a multi-level neural network, and feature classification is performed using a Soft-Max classifier to obtain the emotion classification and emotion intensity of the voice data.

[0472] In the above process, the specific process of inputting voice data into the trained emotional speech model to obtain the emotional type and emotional intensity includes: recognizing voice data to obtain voice text and user voice features; performing semantic understanding on the voice text to obtain the semantic scene corresponding to the voice data; converting the user voice features into a user voice feature vector; converting the semantic scene into a semantic scene feature vector; inputting the user voice feature vector and the semantic scene feature vector into a multi-level neural network to obtain an emotional speech vector, the multi-level neural network includes a two-dimensional convolutional network, a recurrent neural network and two fully connected networks; determining the emotional type and emotional intensity based on the emotional speech vector.

[0473] Step S2803: Obtain the broadcast text corresponding to the voice data;

[0474] In some embodiments, the step of obtaining the broadcast text corresponding to the voice data includes:

[0475] Recognize voice data to obtain voice text;

[0476] The voice text is processed through semantic understanding, business distribution, vertical domain analysis and text generation to obtain semantic business scenarios and broadcast text.

[0477] Perform semantic understanding on the speech text to obtain the slot information and semantic scene corresponding to the speech data;

[0478] The service corresponding to the semantic scene is called to determine the broadcast text corresponding to the slot information.

[0479] The service corresponding to the semantic scenario performs slot analysis and gives the business processing command results. Combined with the processing results, a broadcast text that conforms to the semantic execution results is synthesized.

[0480] In some embodiments, vertical classification locates the music domain, music search intent, and determines that the optimal service is music. This is then distributed to the music microservice for processing. The music microservice will parse the slot for Andy Lau and encapsulate the music information, retrieve third-party music media information for the search, and obtain third-party feedback results. For example, information about 20 songs by Andy Lau is generated based on the music service scenario: "We found 20 songs for you, including Forget Love Water. Come and listen!"

[0481] In some embodiments, the step of obtaining the broadcast text corresponding to the voice data includes:

[0482] Obtain the slot information and semantic scene corresponding to the speech data from the emotional speech model;

[0483] The service corresponding to the semantic scene is called to determine the broadcast text corresponding to the slot information.

[0484] Step S2804: synthesizing the announcement voice based on the announcement text, emotion type and emotion intensity;

[0485] In some embodiments, the step of synthesizing the announcement voice based on the announcement text, emotion type, and emotion intensity includes:

[0486] Determining a phoneme sequence corresponding to the broadcast text;

[0487] Phonemes are the smallest speech units divided according to the natural properties of speech. They are analyzed based on the pronunciation actions in syllables, and one action constitutes a phoneme.

[0488] Generate an audio feature vector sequence corresponding to the phoneme sequence;

[0489] Calculate audio feature emotion based on emotion type and emotion intensity;

[0490] Based on the audio feature vector sequence and audio feature emotion, a broadcast voice with intonation, tone and volume corresponding to the emotion type and emotion intensity is generated.

[0491] The disclosed embodiments utilize speech synthesis technology to generate announcements. Speech synthesis technology converts text into natural, fluent speech. It generates speech by synthesizing phonemes, words, or sentences, and adjusts speech characteristics such as intonation, speech rate, and volume based on the output of a sentiment model to convey specific emotional states.

[0492] In some embodiments, the step of synthesizing the announcement voice based on the announcement text, emotion type, and emotion intensity includes:

[0493] Input the emotion type and emotion intensity into the emotion model to obtain the emotional speech features;

[0494] The emotion model can generate corresponding speech expressions based on emotion classification and emotion intensity. The emotion model is a trained machine learning model that maps emotion type and emotion intensity to corresponding speech features.

[0495] The announcement speech is generated using speech synthesis technology based on emotional speech features and announcement text.

[0496] Step S2805: Send the announcement voice to the display device 200 so that the display device 200 plays the announcement voice.

[0497] In some embodiments, the display device 200 sends a voice interaction identifier while sending the user input voice data. The voice interaction identifier is used to determine the voice program used by the display device 200. The voice program includes a voice assistant and a digital human.

[0498] If the voice interaction identifier is detected as a voice assistant, after generating the broadcast voice, the broadcast voice is sent to the display device 200 so that the display device 200 plays the broadcast voice. The broadcast text can also be sent to the display device 200 together with the broadcast voice, and the broadcast text is displayed on the user interface of the display device 200.

[0499] If the voice interaction identifier is detected as a digital human, then after generating the announcement voice, the server 400 executes:

[0500] Predict key point sequences based on broadcast speech;

[0501] synthesizing digital human image data according to the key point sequence and the digital human image data;

[0502] In some embodiments, the digital human image data is image data corresponding to the digital human selected by the user, wherein the image selected by the user may be determined by receiving a digital human identification sent by the display device 200.

[0503] In some embodiments, the digital human image data is an image or digital human parameters adjusted based on a user-selected or default image. The digital human image data is a sequence of digital human image frames or a sequence of digital human parameters. The digital human image parameters are determined based on the scene and / or the user's emotion type.

[0504] The digital human image data and the announcement voice are sent to the display device 200, so that the display device 200 displays the digital human image based on the digital human image data and plays the announcement voice.

[0505] In some embodiments, after receiving the voice data input by the display device 200, the voice data is recognized to obtain voice text and user voice features. The voice text is semantically understood to obtain semantic scenes and broadcast text. The user voice features and semantic scenes (voice data may also be added) are input into the emotional speech model to obtain the emotion type and emotion intensity. The broadcast voice is synthesized based on the broadcast text, emotion type and emotion intensity, and the broadcast voice is sent to the display device so that the display device plays the broadcast voice. It should be noted that the input of the emotional speech model of the embodiment of the present disclosure during training is user voice features and semantic scenes (voice data may also be added), and the output is emotion type and emotion intensity. Please refer to the above for the internal processing method of the model and will not be repeated here.

[0506] In some embodiments, after receiving the voice data input by the display device 200, the voice data is recognized to obtain voice text and user voice features. The voice text is semantically understood to obtain semantic scenes and broadcast texts. The user voice features, semantic scenes, and broadcast texts (voice data may also be added) are input into the emotional speech model to obtain broadcast voice, and the broadcast voice is sent to the display device so that the display device plays the broadcast voice. It should be noted that the emotional speech model of the embodiment of the present disclosure takes user voice features, semantic scenes, and broadcast texts (voice data may also be added) as input during training, and the output is broadcast voice. The internal processing method of the model is referred to above and will not be repeated here.

[0507] The disclosed embodiment combines semantic scenarios, user voice features and other aspects to train the emotional speech model, fully taps into user interaction features, improves the naturalness of emotional speech synthesis, enhances user experience and emotional communication effects, and enables users to interact with the display device 200 more naturally.

[0508] In some embodiments, the present disclosure further improves some functions of the server 400. The server 400 performs the following steps, as shown in FIG33 .

[0509] Step S3201: receiving the digital human identification and the voice data input by the user sent by the display device 200;

[0510] Among them, the digital human identifier is used to represent the digital human image and voice characteristics selected by the user;

[0511] Before receiving the digital human identification and the voice data input by the user from the display device 200, a digital human selection or customization (registration) process needs to be completed. The digital human that you want to use can be selected from the registered digital humans.

[0512] The digital human registration process includes the following steps:

[0513] 1) Image recording:

[0514] Support users to record videos, take photos or select album pictures for virtual human image generation. After receiving a video or photo recorded by the user, the server 400 generates a digital human image through a series of operations such as cutting out the picture, beautifying the face, and image generation.

[0515] 2) Tone customization:

[0516] Voice customization uses voice cloning technology to replicate or reproduce the user's voice by recording a few basic text passages. This provides a personalized playback tone for digital humans during voice interaction.

[0517] 3) Set a nickname (digital person naming):

[0518] After completing the image recording and voice customization, create a nickname for the virtual digital person as the digital person identification. Under the same account, the nickname of the virtual digital person cannot be repeated.

[0519] The above steps have been described in detail above and will not be repeated here.

[0520] It should be added that after setting the nickname, there is an additional step: 4) Setting members (for example, family members):

[0521] Select the member nickname corresponding to the digital human recording user to establish an association.

[0522] In some embodiments, during the process of setting up family members, a nickname of the family member may be filled in, and the relationship between the family member and the head of household may be set, thereby constructing a family relationship diagram.

[0523] In some embodiments, a creation entry for adding family members is provided on the display device, which is freely entered by the user. Family member information includes: family member nicknames (real names may not be used to protect user privacy), relationship with the head of household (used to build family relationships), and serial number (indicating the child's number, used to build relationships between children).

[0524] In some embodiments, after a family member is created, the family member information can be viewed in the user's personal center, as shown in Figure 34. Based on the family member information, a family relationship diagram can be constructed, as shown in Figure 35. For clarity, the disclosed embodiments only draw a single-line relationship, but a double-line relationship should actually be drawn.

[0525] After confirming the family member information, during the process of setting up family members, you can fill in the family member nickname to determine the relationship between the digital human recording user and the household head.

[0526] After setting up family members, a 3-5 minute algorithm training process is performed to generate a virtual digital person for the user, who can then choose to use that person for voice interaction.

[0527] The digital human data storage is shown in Table 5:

[0528] Table 5

[0529] Step S3202: determining user identity information corresponding to the voice data, and recognizing the voice data to obtain voice text;

[0530] Before determining the user identity information corresponding to the voice data, voiceprint registration is required.

[0531] In some embodiments, voiceprint registration can be performed without a user's knowledge, meaning that the user's voiceprint information is automatically recognized while the user is speaking, completing the voiceprint registration. Specifically, after receiving the voice data input by the user, the voiceprint information of the voice data is extracted. If the voiceprint information does not match any of the registered voiceprint information in the personal voiceprint database, a prompt message will pop up, prompting the user whether to register as a new member. If the user inputs an instruction to choose not to register, the registration process will not be executed. If the user inputs an instruction to choose to register, the user will be required to set a voiceprint nickname and family member nicknames, thereby establishing an association between the voiceprint account and family members. To improve the accuracy of the voiceprint information, audio of the basic text being read aloud can also be recorded.

[0532] The data storage of voiceprint information is shown in Table 6:

[0533] Table 6

[0534] In some embodiments, voiceprint registration can be guided. The voiceprint registration function can be found in the voice zone. It generally guides the user through reading three basic text paragraphs, setting a voiceprint nickname and family member nicknames, and completing voiceprint registration, thereby establishing an association between the voiceprint account and the family member.

[0535] The disclosed embodiments use voiceprint recognition technology to analyze and compare individual voice features for identity verification or identification. As shown in Figure 36, the speaker's identity is confirmed through a series of operations, including user input voice detection, preprocessing (such as denoising), feature extraction, voiceprint comparison, and result determination. If the current speaker's voiceprint has a high degree of similarity (greater than a set threshold) with the registered voiceprint information, it is considered to be the same person. The extracted voiceprint features can be used for voiceprint registration to obtain a voiceprint model, which is then stored in a voiceprint library for subsequent voiceprint comparison.

[0536] The step of determining the user identity information corresponding to the voice data includes:

[0537] Extract voiceprint information from voice data;

[0538] In some embodiments, the step of extracting voiceprint information from speech data includes:

[0539] Splitting the voice data into at least one audio data of a preset length;

[0540] Pre-emphasize, frame, and window the sound signal time course of the audio data to obtain a windowed sound signal time course;

[0541] Perform fast Fourier transform on the windowed sound signal time history to obtain spectrum distribution information;

[0542] determining an energy spectrum based on the spectrum distribution information;

[0543] The energy spectrum is passed through a set of triangular filter banks to obtain the logarithmic energy of the filter output;

[0544] The logarithmic energy is subjected to discrete chord transform to obtain the Mel-frequency cepstral coefficient, the derivative corresponding to the Mel-frequency cepstral coefficient, and the second-order derivative;

[0545] The mel-frequency cepstral coefficients, the derivatives corresponding to the mel-frequency cepstral coefficients, and the second-order derivatives are determined as voiceprint information.

[0546] Determine whether the voiceprint information matches the registered voiceprint information in the voiceprint database;

[0547] In some embodiments, the step of determining whether the voiceprint information matches the registered voiceprint information in the voiceprint database includes:

[0548] Calculate the similarity between the voiceprint feature information and the registered voiceprint information;

[0549] The maximum number of statistical similarities greater than the similarity threshold;

[0550] If the maximum number is greater than the preset number, it is determined that the voiceprint information matches the registered voiceprint information in the voiceprint database;

[0551] If the maximum number is not greater than the preset number, it is determined that the voiceprint information does not match the registered voiceprint information in the voiceprint library.

[0552] If the voiceprint information matches the registered voiceprint information in the voiceprint database, the user identity information is determined based on the registered voiceprint information, that is, the voiceprint nickname and family member nicknames of the registered voiceprint information are obtained.

[0553] Use speech recognition technology to convert voice data into speech text.

[0554] Step S3203: Determine the relationship between the digital human and the user based on the digital human identifier and the user identity information;

[0555] The user identity information includes the nicknames of the speaker's family members.

[0556] Obtain the nickname of the family member corresponding to the digital human identity;

[0557] The relationship between the digital person and the user is determined in the family relationship graph based on the speaker's family member nickname and the family member nickname corresponding to the digital person identifier.

[0558] In some embodiments, the speaker's family member nickname is Zhang cc, and the family member nickname corresponding to the digital human identifier is Zhang aa, then the relationship between the digital human and the user is determined to be a child-father relationship.

[0559] It should be noted that both the user and the digital person need to have a family member nickname in order to determine the relationship between the digital person and the user.

[0560] Step S3204: determining a basic text according to the speech text;

[0561] The speech text is processed through natural language processing (NLP) to determine the base text. The base text refers to the text normally fed back to the speech data. Natural language processing (NLP) is a technology that uses computer technology to analyze, understand, and process natural language. Natural language processing consists of two parts: natural language understanding (NLU) and natural language generation (NLG). Natural language understanding is used to understand the meaning of natural language text, while natural language generation is used to express given intentions, ideas, etc. in natural language text.

[0562] The steps of determining the basic text according to the speech text include:

[0563] Perform word segmentation and annotation processing on the speech text to obtain word segmentation information;

[0564] Perform syntactic and semantic analysis on the word segmentation information to obtain slot information;

[0565] Use vertical domain classification to locate the domain intention corresponding to the slot information;

[0566] Determine the basic text based on domain intent and slot information.

[0567] The steps for determining the basic text based on the voice text have been described in detail above and will not be repeated here.

[0568] It should be noted that each voice domain service has a default basic text. This default basic text can be generated in real time within the service or pre-configured (data in the announcement configuration). For example, for "Today's weather," the basic text sentence structure is {area}{date}{condition}, {temperature}, {winddir}{windlevel}. For example, Beijing is cloudy today, with a north wind of 22 to 29 degrees Celsius and a north wind of 3 to 4. Alternatively, the data in the announcement configuration can be used: "Weather information found for you."

[0569] Step S3205: Generate a broadcast text based on the basic text and the relationship;

[0570] Among them, the broadcast text generation methods include pre-splicing, post-splicing, pre-splicing + post-splicing and replacing the default basic text.

[0571] In some embodiments, the step of generating the announcement text based on the basic text and the relationship includes:

[0572] Acquire splicing information corresponding to the relationship, the splicing information including a splicing position and splicing content, the splicing position including a pre-splice, and the splicing content corresponding to the pre-splice being a title set according to the relationship;

[0573] The title set according to the relationship may be randomly selected by the server or set by the user.

[0574] You can set the title of the speaker based on the relationship. For example, the title of "father" can be set as father, daddy, daddy, daddy, old man, and old man. You can also set adjectives that express intimacy, such as dear, respected, and respected.

[0575] Generate broadcast text based on splicing information and basic text.

[0576] The splicing content is spliced ​​into the splicing position of the basic text to generate the broadcast text.

[0577] In some embodiments, when a user inputs "What's the weather like today?", the semantic analysis of domain intent and slots yields the basic text "Beijing is cloudy today, 22 to 29 degrees Celsius, north wind 3 to 4." After determining that the relationship between the digital human and the user is a child-parent relationship, the splicing information is pre-splice (splice position) - Dad (splice content), and the generated broadcast text is "Dad, Beijing is cloudy today, 22 to 29 degrees Celsius, north wind 3 to 4."

[0578] In some embodiments, if the base text includes special text content, the base text can be replaced with text specific to the special text content. For example, when querying the weather, if you want to highlight a particular weather condition, such as a weather warning or a large temperature difference, you can create the desired report text based on the weather information and then replace the base text to generate the report text.

[0579] In some embodiments, if the basic text includes special text content, some reminder-related text can be configured for the special text content and added to the end of the basic text. Some related words can be configured according to the weather conditions and combined with the basic text through post-splicing.

[0580] In some embodiments, the splicing position further includes post-splicing, and the step of generating the broadcast text based on the basic text and the relationship includes:

[0581] Get the user's age;

[0582] In some embodiments, the step of obtaining the user's age includes: determining the user's age using voice recognition technology.

[0583] In some embodiments, when registering a voiceprint, an option to add age can be added, and the user's age can be directly obtained from the voiceprint registration information.

[0584] Based on age and basic text, determine the corresponding splicing content for post-splicing.

[0585] The basic text includes special text content, and some reminder-related text will be configured for the special text content, with different splicing contents set for different age groups.

[0586] In some embodiments, when the user voice input is "How's the weather today?", and the basic text obtained through semantic analysis of domain intent and slots contains stormy weather, the basic text can be replaced with "Today there is a blue warning for heavy rain, and 6-8 level winds". When it is determined that the speaker is elderly, the corresponding splicing content of the post-splicing is "Don't go out if you have nothing to do", and the generated broadcast text is "Today there is a blue warning for heavy rain, and 6-8 level winds, so don't go out if you have nothing to do". When it is determined that the speaker is middle-aged, the corresponding splicing content of the post-splicing is "Remember to take precautions when going out", and the generated broadcast text is "Today there is a blue warning for heavy rain, and 6-8 level winds, so remember to take precautions when going out". Titles can be added to the final broadcast text based on relationships, such as "Dad, today there is a blue warning for heavy rain, and 6-8 level winds, so don't go out if you have nothing to do".

[0587] In some embodiments, the step of generating the announcement text based on the basic text and the relationship includes:

[0588] Check whether the current date is the target date, where the target date is a holiday and / or anniversary. Holidays include Father's Day, Mother's Day, Children's Day, Valentine's Day, etc. Anniversaries include birthdays and wedding anniversaries, etc. Anniversaries can be written and stored by the user.

[0589] If the current date is detected as the target date, it is determined whether the target date is relevant to the relationship;

[0590] In some embodiments, the current date is Father's Day. If the relationship between the digital person and the user is a son-father relationship, then Father's Day is related to the son-father relationship. If the relationship between the digital person and the user is a grandfather-grandson relationship, then Father's Day is not related to the grandfather-grandson relationship.

[0591] If the target date is related to a relationship, determining a target text according to the relationship, wherein the target text includes a blessing text and / or a reminder text;

[0592] If the user is determined to be the recipient of a blessing based on the relationship and the target date, the target text is determined to be the blessing text;

[0593] If the user is determined to be the well-wisher based on the relationship and the target date, the target text is determined to be the prompt text.

[0594] In some embodiments, if the current date is Father's Day and the relationship between the digital human and the user is a child-father relationship, the target text is determined to be a blessing text, and the blessing text is "Dad, Happy Father's Day! I wish you happiness and happiness every year." If the relationship between the digital human and the user is a father-son relationship, the target text is determined to be a prompt text, and the prompt text is "Today is Father's Day, remember to send your dad a blessing."

[0595] In some embodiments, the step of generating the announcement text based on the basic text and the relationship includes:

[0596] Check whether the current date is the target date;

[0597] If the current date is detected as the target date, determine whether the target date is relevant to the user;

[0598] In some embodiments, the current date is Children's Day, and if the user is a child, Children's Day is relevant to the user. If the user is an adult, Children's Day is not relevant to the user.

[0599] If the target date is relevant to the user, generate the target text.

[0600] In some embodiments, the broadcast text is "Happy Children's Day, baby."

[0601] In some embodiments, the step of generating the announcement text based on the basic text and the relationship includes:

[0602] Check whether the target date is included in the preset range of dates. The preset range of dates can be the current date or three days after the current date.

[0603] If the target date is included in the preset range of dates, then determine whether the target date is relevant to the user or relationship;

[0604] If the target date is related to the user or relationship, the target text is generated. If the target date is not today, the target text is a prompt text indicating how many days are left until the target date.

[0605] In some embodiments, if the intent obtained by parsing the voice text is a holiday or anniversary query intent, the access query interface is called to obtain the name of the holiday or anniversary, the corresponding target text is queried in the broadcast text configuration, and then spliced ​​with the title to generate the broadcast text.

[0606] In some embodiments, if the intent obtained by parsing the voice text is not a holiday or anniversary query intent, obtaining a holiday query identifier corresponding to the user;

[0607] If the holiday query flag is 1, the query interface is called to retrieve the base text corresponding to the intent, check whether the current date is the target date, and set the user's holiday query flag to 0. At a fixed time each day, such as 00:00, the holiday query flag is reset to 1 to ensure that each user only queries the holiday query once a day. If the target text is obtained, it is added to the base text, that is, the target text is spliced ​​before or after the base text to obtain the broadcast text.

[0608] The above method can be used to generate announcement text for all voice application scenarios. There are slight differences in the announcement language in different business fields, but the overall idea is to obtain key business information, obtain the corresponding business information (basic text), and then combine the speaker's age and holiday information to generate the final announcement text.

[0609] Step S3206: Generate digital human data based on the voice features and image data corresponding to the digital human identifier and the broadcast text;

[0610] The algorithm for generating digital humans is a generative adversarial network, a neural network model consisting of a generator and a discriminator. The generator is responsible for producing realistic digital human images, while the discriminator is responsible for determining whether the generated images are real or forged. Through continuous confrontation and learning, the generator is able to gradually produce more realistic digital human images.

[0611] The steps of generating digital human data based on the voice features and image data corresponding to the digital human identifier and the broadcast text include:

[0612] Synthesize the broadcast voice according to the voice features corresponding to the digital human identifier and the broadcast text;

[0613] Predict key point sequences based on broadcast speech;

[0614] Digital human image data is synthesized according to the key point sequence and the image data corresponding to the digital human identification; the digital human data includes the digital human image data and the broadcast voice.

[0615] In some embodiments, the digital human image data may be decorated according to domain intent and / or user emotion type.

[0616] Step S3207: Send the digital human data to the display device 200, so that the display device 200 plays the digital human image and voice according to the digital human data.

[0617] In some embodiments, the digital human image data is a sequence of image frames, and the server 400 sends the sequence of image frames and the announcement voice to the display device 200 in a live streaming manner. The display device 200 displays the image corresponding to the image frame and plays the announcement voice.

[0618] In some embodiments, the digital human image data is a digital human parameter sequence, and the server 400 sends the digital human parameter sequence and the announcement voice to the display device 200. The display device 200 displays the digital human image and plays the announcement voice based on the digital human parameters and the basic model.

[0619] In some embodiments, after detecting that the duration of entering the target scene exceeds a preset duration, the display device 200 sends a timeout message to the server 400. The timeout message includes the target scene.

[0620] After receiving the timeout message, the server 400 generates a prompt text based on the relationship and the target scenario;

[0621] Generate digital human data based on the voice features and image data corresponding to the digital human identifier and the prompt text;

[0622] The digital human data is sent to the display device so that the display device plays the digital human image and voice according to the digital human data.

[0623] In some embodiments, when a user says "I want to play mahjong," the digital human announces, "Dad, come show them your superb winning skills." If it detects that the user has stayed in the mahjong interface for more than one hour, a timeout message is uploaded to the server 400, and digital human data is generated and sent to the display device 200. The display device 200 displays the digital human announcing, "Dad, you've been playing for a long time. End the game and take a break," as shown in Figure 37.

[0624] In the disclosed embodiment, the digital human is generated through real-life video and audio recording. By establishing a family relationship diagram, the relationship between the speaker and the digital human is obtained based on voiceprint information and virtual digital human information, and interesting broadcast content like a family chat is generated, making the user feel like they are accompanied by their family when using voice, thereby improving the user experience.

[0625] Furthermore, in practical applications, display devices may experience issues such as lag or inability to operate a digital human due to resource, network, and concurrency constraints. For example, a display device might simultaneously play high-definition video and engage in real-time remote chat while operating a digital human. These tasks consume significant display device resources, leading to lag or inability to operate the digital human. This, in turn, affects the timeliness and stability of the digital human when interacting with it, resulting in a poor user experience. To address this issue, the disclosed embodiments, based on the aforementioned digital human processing method, further add a digital human driving process.

[0626] The digital human driving process of the embodiment of the present disclosure can be executed by the aforementioned display device 200 or the aforementioned server 400, or can be executed by the display device 200 and the server 400 together.

[0627] For example, when the display device 200 and server 400 jointly execute the digital human driving process of the disclosed embodiment, the process is as follows: the display device 200 obtains the text to be driven and determines the resource utilization rate. Then, if the resource utilization rate is determined to be less than or equal to the utilization rate threshold and a primary driving solution is matched in the first solution library based on the text to be driven, the digital human is driven using the primary driving solution. If the resource utilization rate is determined to be less than or equal to the utilization rate threshold and no primary driving solution is matched in the first solution library based on the text to be driven, a driving request is sent. The first solution library includes driving texts and driving solutions, with each driving text corresponding to one driving solution. The driving request includes a desired level and the data to be driven, which instructs the real-time driving of the digital human. Upon receiving the driving request, the server 400 obtains the current concurrency number and determines the cloud level based on the current concurrency number. Then, the server 400 determines the actual level based on the desired level and the cloud level, and determines the target driving solution based on the actual level and the data to be driven. Finally, the server 400 controls the display device 200 to drive the digital human using the target driving solution. This avoids the situation where the digital human is stuck or cannot be driven due to resource and concurrent factors, thus improving the user experience.

[0628] To facilitate the description of the digital human driving process in the embodiment of the present disclosure, the subject that executes the digital human driving process will be referred to as a digital human driving device. The digital human driving device in the embodiment of the present disclosure stores data such as the application program of the first application and the operating system of the digital human driving device.

[0629] Next, the digital human driving process provided by the embodiment of the present disclosure is described using the aforementioned display device 200 as a digital human driving device. As shown in FIG38 , the digital human driving process may include the following steps:

[0630] Step S11: Acquire the text to be driven and determine the resource occupancy rate.

[0631] The text to be driven is converted from the data to be driven. The data to be driven is the instruction input by the user to drive the digital human, for example, the data to be driven can be text, voice, or other instructions.

[0632] First, obtain the text to be driven.

[0633] In some embodiments, the method for obtaining the text to be driven can be that when the digital human driving device receives the digital human driving instruction, it first obtains the voice content input by the user, and then converts the voice content into text to obtain the text to be driven.

[0634] In some other embodiments, the method for obtaining the text to be driven may also be that after the digital human driving device receives the text to be driven, it determines that the user needs to drive the digital human, that is, the digital human driving device directly receives the text to be driven without the need for a voice-to-text conversion process.

[0635] Of course, the digital human driving device may also receive data to be driven in other forms, which can be converted into text to be driven in the digital human driving device, and this disclosure does not limit this.

[0636] Secondly, after obtaining the text to be driven, determine the resource usage.

[0637] The resource utilization rate may be the resource utilization rate of the central processing unit (CPU), the resource utilization rate of the graphics processing unit (GPU), the average of the resource utilization rate of the CPU and the resource utilization rate of the GPU, the resource utilization rate of other processing units, or a combination thereof, which is not limited in the present disclosure.

[0638] Specifically, the method for determining the resource occupancy rate may be to count the real-time resource occupancy rate at a fixed period, and after obtaining the text to be driven, take the average of multiple real-time resource occupancy rates within a preset time period as the final resource occupancy rate. For example, the method for determining the resource occupancy rate may be to count the real-time resource occupancy rate every 500 milliseconds (ms). After obtaining the text to be driven, read the real-time resource occupancy rate 3 seconds (s) before this moment (i.e., the moment when the text to be driven is obtained), and calculate the average value to obtain the final resource occupancy rate.

[0639] Step S12: When it is determined that the resource occupancy rate is less than or equal to the occupancy rate threshold, and a primary driving solution is matched in the first solution library according to the text to be driven, the digital human is driven using the primary driving solution.

[0640] First, it should be noted that the digital humans provided in the embodiments of this disclosure are all 3D digital humans, and the digital human driving process in this disclosure is used to drive the head of a 3D digital human. The first solution library includes driving text and driving solutions, with one driving text corresponding to one driving solution. Each driving solution includes at least one blendshape component, and each blendshape component corresponds to a weight value. Each blendshape component is used to display a portion of the digital human's head (for example, eyes, eyebrows, mouth, etc.). Changing the weight value corresponding to the blendshape component can control the degree of influence of the blendshape component in the animation. That is, by changing the weight value of the blendshape component, expression or deformation can be generated in the portion of the digital human corresponding to the blendshape component.

[0641] Specifically, the blendshape algorithm in the disclosed embodiments is a technology used in computer animation and 3D modeling, commonly used to create realistic facial expressions and character transformations. Using the blendshape algorithm to control the expression changes and animation of a digital human can be as follows: First, the digital human needs to be modeled, that is, a complete body model of the digital human needs to be created, including the skeletal structure and skin surface geometry. Next, a head blendshape model is created (the head blendshape model is composed of multiple blendshape components, each blendshape component representing a portion of the digital human's head, such as the eyes, eyebrows, mouth, etc.) to control the changes in the digital human's facial expressions. Next, weight values ​​are set for each blendshape component to control the influence of each blendshape component in the animation. The weight values ​​can be adjusted through programming or the control panel in the animation software. Then, traditional skeletal animation techniques are used to control the digital human's posture, movements, and motion. Furthermore, based on skeletal animation, the weights of the blendshape components are used to control the changes in the digital human's facial expressions. In other words, by changing the weight values ​​of the facial blendshape components, expression changes can be achieved to suit specific action requirements. Finally, the digital human, whose facial expressions have been modified using blendshape components, is rendered in real time in a rendering engine, presenting it as a realistic full-body animation. The rendering engine interpolates and deforms the geometric shapes of the digital human model's face based on the weight values ​​of the blendshape components, producing smooth transitions and natural animation effects.

[0642] In some embodiments, the occupancy threshold in the embodiments of the present disclosure is preset, for example, the occupancy threshold is a default value, or the occupancy threshold is a value determined by relevant personnel based on the actual situation of the digital human driving the device.

[0643] Secondly, when it is determined that the resource occupancy rate is less than or equal to the occupancy rate threshold, and a primary driving solution is matched in the first solution library according to the text to be driven, the primary driving solution is used to drive the digital human.

[0644] Specifically, as shown in FIG39 , when it is determined that the resource occupancy rate is less than or equal to the occupancy rate threshold, and a primary driving solution is matched in the first solution library according to the text to be driven, the method of driving the digital human using the primary driving solution may include the following steps:

[0645] Step S121: Determine whether the resource occupancy rate is less than or equal to the occupancy rate threshold, and when the resource occupancy rate is less than or equal to the occupancy rate threshold, execute step S122; and when the resource occupancy rate is greater than the occupancy rate threshold, execute step S124.

[0646] Step S122: Determine whether a primary driving solution is matched in the first solution library according to the text to be driven. If a primary driving solution is matched in the first solution library according to the text to be driven, execute step S123; if no primary driving solution is matched in the first solution library according to the text to be driven, execute step S13.

[0647] In some embodiments, before matching the primary driving solution in the first solution library based on the text to be driven, the digital human driving process may further include creating a first solution library. The first solution library may be created by pre-setting multiple driving texts and driving solutions corresponding to the multiple driving texts in the solution library to obtain the first solution library.

[0648] Step S123: driving the digital human using the primary driving solution.

[0649] Specifically, the primary drive scheme in this disclosure includes at least one blendshape component, and each blendshape component corresponds to a weight value. When driving a digital human using the primary drive scheme, the weight value of each blendshape component in the digital human model is adjusted based on the weight value corresponding to each blendshape component in the primary drive scheme to achieve changes in the digital human's facial expressions.

[0650] Step S124: Use a graphics interchange format (GIF) dynamic image to drive the digital human.

[0651] Specifically, when the resource utilization rate exceeds the utilization threshold, it is determined that the digital human driving device is running too many tasks and resources are tight. In this case, a GIF dynamic image is used to drive the digital human for visual display. This driving method does not support switching viewing angles, meeting the minimum resource display setting.

[0652] In the above scheme, after obtaining the text to be driven, the digital human driving device determines the need to drive the digital human and determines the resource utilization rate. Subsequently, if the resource utilization rate is determined to be less than or equal to the utilization rate threshold and a primary driving solution is matched in the first solution library based on the text to be driven, the primary driving solution is used to drive the digital human. In this way, when the digital human driving device determines the need to drive the digital human, it first determines its own resource utilization rate and uses different driving solutions when the resource utilization rate falls within different ranges. This avoids lags and inoperability caused by resource constraints when driving the digital human, thus improving the user experience. Furthermore, if the resource utilization rate is determined to be within the appropriate range and a primary driving solution is matched in the first solution library, the primary driving solution is directly used to drive the digital human, saving the resource loss of real-time cloud-based driving and the time loss of network transmission.

[0653] Step S13: When it is determined that the resource occupancy rate is less than or equal to the occupancy rate threshold and no primary driving solution is matched in the first solution library according to the text to be driven, a driving application is sent.

[0654] The drive request includes the desired level and pending drive data, which are used to instruct the real-time driving of the digital human. The desired level is determined based on the desired level, drive time, actual time, and actual level corresponding to the last drive of the digital human. The initial value of the desired level can be intermediate. For example, if the highest desired level is level 3, the initial value of the desired level can be level 2. If the highest desired level is level 4, the initial value of the desired level can be level 2 or level 3.

[0655] Specifically, if it is determined that the resource utilization rate is less than or equal to the utilization rate threshold, and no primary driving solution is matched in the first solution library based on the to-be-driven text, it is determined that there is currently no suitable driving solution for driving the digital human. It is necessary to determine a driving solution based on the to-be-driven data in real time. At this point, a driving request is sent to the server, requesting the server to drive the digital human in real time based on the to-be-driven text.

[0656] In the above scheme, when it is determined that the resource occupancy rate is less than or equal to the occupancy rate threshold, and no primary driving scheme is matched in the first scheme library according to the text to be driven, a driving request is sent to the server to request the server to drive the digital human in real time according to the text to be driven. Different driving entities can be adaptively switched according to resource conditions, pursuing the optimal driving effect while ensuring timeliness, thereby improving user experience.

[0657] Step S10: When receiving the target driving solution, drive the digital human using the target driving solution.

[0658] Specifically, according to the weight value corresponding to the blendshape component in the target driving scheme, the weight value of the blendshape component of the digital human is adjusted to achieve the change of the digital human's facial expression, so that the digital human can achieve the specific expression change corresponding to the driving feature.

[0659] In some embodiments, as shown in FIG40 , after sending the driving application, the digital human driving process further includes the following steps:

[0660] Step S14: Receive the application result and determine the actual time consumed.

[0661] The application result includes the actual level and driving time. The actual time is the time between sending the driving application and receiving the application result.

[0662] Specifically, the digital human driving device records the time taken between sending the driving application and receiving the application result, and determines it as the actual time taken.

[0663] Step S15: Determine the next expected level according to the driving time, the actual time, and the actual level.

[0664] The next expected level is the expected level corresponding to the next time the digital human is driven.

[0665] In some embodiments, as shown in FIG41 , the method of determining the next expected level based on the driving time, the actual time, and the actual level may include the following steps:

[0666] Step S151: Calculate the network time consumption according to the driving time consumption and the actual time consumption.

[0667] Specifically, the network time consumption can be calculated according to the following formula: Network_T=Total_T-Driver_T, where Network_T is used to represent the network time consumption, Total_T is used to represent the actual time consumption, and Driver_T is used to represent the driver time consumption.

[0668] Step S152: Determine whether the network time consumption is greater than the first time threshold, and if the network time consumption is greater than the first time threshold, execute step S1521; if the network time consumption is less than or equal to the first time threshold, execute step S153.

[0669] Step S1521: Determine the next expected level as the actual level minus one.

[0670] Step S153: Determine whether the network time consumption is greater than the second time threshold, and if the network time consumption is greater than the second time threshold, execute step S1531; if the network time consumption is less than or equal to the second time threshold, execute step S1532.

[0671] The second time threshold is smaller than the first time threshold.

[0672] Step S1531: Determine that the next expected level is the actual level.

[0673] Step S1532: Determine the next expected level as the actual level plus one.

[0674] Specifically, when Network_T > Thr_T1, the next expected level is determined to be the actual level minus one; when Thr_T1 ≥ Network_T > Thr_T2, the next expected level is determined to be the actual level; and when Network_T ≤ Thr_T2, the next expected level is determined to be the actual level plus one. Where Network_T represents network time, Thr_T1 represents the first time threshold, and Thr_T2 represents the second time threshold.

[0675] In this solution, the next desired level is adaptively adjusted based on the application result and actual time consumption. This enables a cyclical decision-making strategy to be formed through real-time information exchange between the display device and the server. This allows for high-speed sub-level switching, which can promptly alleviate issues such as network congestion or sudden increases in concurrent access, ensuring a smoother user experience. Furthermore, determining the actual level based on network time consumption avoids network-related issues that can cause lag or inoperability when driving the digital human, thus improving the user experience.

[0676] In the following embodiments, the method of the embodiments of the present disclosure is described by taking the digital human driving device on the server side as an example of the execution subject of the digital human driving process provided by the embodiments of the present disclosure.

[0677] Next, the digital human driving process provided by the embodiment of the present disclosure is described using the aforementioned server 400 as a digital human driving device. As shown in FIG42 , the digital human driving process may include the following steps:

[0678] Step S16: Receive the driver application and obtain the current concurrent number.

[0679] The driving application includes the expected level and the data to be driven.

[0680] When a driving application is received, it is determined that the digital human needs to be driven in real time and the current concurrent number is obtained.

[0681] In some embodiments, the current concurrent number can be obtained by counting the number of requests in a period of time in real time and using the number of requests as the current concurrent number. For example, the number of requests in the last 1 second can be counted in real time and used as the current concurrent number.

[0682] Step S17: Determine the cloud level based on the current number of concurrent connections.

[0683] In some embodiments, as shown in FIG43 , determining the cloud level based on the current number of concurrent users may include the following steps:

[0684] Step S171: Obtain the initial cloud level.

[0685] The initial cloud level is preset, for example, the initial cloud level is the highest level.

[0686] Step S172: Determine whether the current number of concurrent requests is less than or equal to the concurrent threshold, and if so, execute step S1721; if so, execute step S1722.

[0687] The concurrency threshold is a positive integer.

[0688] Step S1721: Determine that the cloud level is the initial cloud level.

[0689] Step S1722: Determine the cloud level as the initial cloud level minus a preset threshold.

[0690] The current concurrency number is identified by N1, and the concurrency threshold is represented by n. During implementation, when N1≤n, the cloud level is determined to be the initial cloud level; when N1>n, the cloud level is determined to be the initial cloud level minus the preset threshold; wherein the preset threshold is a preset positive integer, which can be a default value or a value set by relevant personnel based on actual conditions.

[0691] In some embodiments, the concurrency threshold includes a first concurrency threshold and a second concurrency threshold, a first preset threshold and a second preset threshold, the first concurrency threshold is less than the second concurrency threshold, the first preset threshold is greater than the second preset threshold, and the first concurrency threshold, the second concurrency threshold, the first preset threshold, and the second preset threshold are all positive integers. Determining the cloud level based on the current concurrency number can also be: when N1≤n1, the cloud level is determined to be the initial cloud level; when n2≥N1>n1, the cloud level is determined to be the initial cloud level minus the first preset threshold; when N1>n2, the cloud level is determined to be the initial cloud level minus the second preset threshold. Wherein, N1 is used to represent the current concurrency number, n1 is used to represent the first concurrency threshold, and n2 is used to represent the second concurrency threshold.

[0692] Of course, the number of concurrent thresholds and the number of preset thresholds can be set according to the computing power of the hardware device, and this disclosure does not limit this.

[0693] Step S18: Determine the actual level based on the expected level and the cloud level.

[0694] In some embodiments, the actual level may be determined based on the expected level and the cloud level by taking the smallest value between the expected level and the cloud level as the actual level.

[0695] Step S19: Determine a target driving scheme according to the actual level and the data to be driven, and send the target driving scheme to instruct the digital human to be driven using the target driving scheme.

[0696] In some embodiments, the target driving scheme includes at least one blendshape component, and one blendshape component corresponds to one weight value.

[0697] In some embodiments, a target driving solution can be determined based on the actual level and the target driving data by inputting the target driving data and the actual level into a driving network model for solution extraction to obtain the target driving solution. The driving network model is trained using preset driving data and preset driving levels as inputs and a preset driving solution as output.

[0698] Specifically, before inputting the data to be driven and the actual level into the driving network model for scheme extraction processing to obtain the target driving scheme, the digital human driving process also includes training and generating a driving network model based on the preset driving data, the preset driving level, and the preset driving scheme.

[0699] As shown in FIG44 , the method of training and generating a driving network model according to preset driving data, preset driving levels, and preset driving schemes may include the following steps:

[0700] Step S01: Obtain preset driving data, preset driving level, and preset driving scheme, and perform feature extraction on the driving data to obtain driving features.

[0701] First, the method for obtaining preset driving data and preset driving schemes can be to call the historical driving data and its corresponding driving schemes input by the user within a certain historical time period, or to simulate the driving data and its corresponding driving schemes through preset devices. This disclosure does not limit this.

[0702] The method for obtaining the preset drive level may be to determine it according to the preset rules and the number of blendshape components in the preset drive scheme. The preset rules include the corresponding relationship between the drive level and the number of blendshape components in the drive scheme. For example, the preset rule may be: when P≤n, the preset drive level is determined to be level one, when n<P≤m, the preset drive level is determined to be level two, and when m<P, the preset drive level is determined to be level three; wherein P is used to represent the number of blendshape components in the drive scheme, and m>n. For another example, the preset rule may be: when P≤n, the preset drive level is determined to be level one, when n<P≤mi, the preset drive level is determined to be level two, when mi<P≤m, the preset drive level is determined to be level three, and when m<P, the preset drive level is determined to be level four; wherein, n<mi<m. The present disclosure does not limit the number of drive levels in the preset rules.

[0703] Afterwards, feature extraction is performed on the driving data to obtain driving features.

[0704] In some embodiments, feature extraction of the driving data may be performed using a feature extraction algorithm to extract features from the driving data to obtain driving features. For example, the feature extraction algorithm may be a Mel frequency cepstral coefficients (MFCC) algorithm or a filter bank (fbank) algorithm.

[0705] Step S02: determining the number of driver sub-networks in the driver network model according to a preset rule, and fixing the level of each driver sub-network.

[0706] Specifically, the number of driver sub-networks in the driver network model is determined according to the number of driver levels in the preset rule. For example, if the number of driver levels in the preset rule is 3, then the number of driver sub-networks is also 3.

[0707] The level of each driver sub-network may be fixed in such a manner that one driver sub-network corresponds to one driver level, and the driver levels of any two driver sub-networks are different.

[0708] Step S03: Perform the following training operation on each driving sub-network to obtain a preset number of driving sub-network models, and form a driving network model with the preset number of sub-network models.

[0709] The training operation includes: for the target driving sub-network, using the preset driving level that is the same as the driving level of the target driving sub-network and the corresponding driving features as the input of the target driving sub-network, and the corresponding preset driving scheme as the output, training the target driving sub-network n times until the loss function of the target driving sub-network converges, and obtaining the driving sub-network model corresponding to the target driving sub-network, wherein the target driving sub-network is any driving sub-network.

[0710] After the driving network model is trained and generated, the data to be driven and the actual level are input into the driving network model for scheme extraction processing to obtain the target driving scheme.

[0711] In the above scheme, after receiving the driver request, the digital human driver device obtains the current concurrency and determines the cloud level based on the current concurrency. It then determines the actual level based on the expected level and cloud level in the driver request, and determines the target driver solution based on the actual level and the text to be driven. Finally, the target driver solution is used to drive the digital human. In this way, when the display device determines that its resource utilization is within the appropriate range but cannot match a primary driver solution in the first solution library, it sends a driver request to the server, allowing the server to control the display device to drive the digital human in real time. After receiving the driver request, the server first determines its own concurrency and, when the concurrency meets the conditions, controls the display device to drive the digital human in real time. This avoids stalls and inoperability when driving the digital human due to concurrency factors, thereby improving the user experience.

[0712] In some embodiments, after sending the target driving solution, the digital human driving process further includes: returning an application result, wherein the application result includes the actual level and driving time, which is used to indicate the next expected level.

[0713] In some embodiments, after sending the target driving solution, the digital human driving device needs to return an application result; wherein, the application result is used to indicate the next expected level, including the actual level and driving time.

[0714] Therefore, the next expected level can be adaptively adjusted according to the application results and actual time consumption. A cyclic decision-making strategy can be formed through real-time information interaction between the display device and the server, and sub-level switching can be achieved in real time. Network congestion or sudden increase in concurrent access can be alleviated in time, ensuring that users have a smoother experience.

[0715] Next, the digital human driving process provided by the embodiment of the present disclosure is described with the aforementioned display device 200 and server 400 serving as digital human driving devices. As shown in FIG45 , the digital human driving process may include the following steps:

[0716] S31. The display device obtains the text to be driven and determines the resource occupancy rate.

[0717] The text to be driven is obtained by converting the data to be driven.

[0718] S32: When the display device determines that the resource occupancy rate is less than or equal to the occupancy rate threshold and a primary driving scheme is matched in the first scheme library according to the text to be driven, the display device drives the digital human using the primary driving scheme.

[0719] The first solution library includes driving texts and driving solutions, and one driving text corresponds to one driving solution.

[0720] S33: When the display device determines that the resource occupancy rate is less than or equal to the occupancy rate threshold and no primary driving solution is matched in the first solution library according to the text to be driven, the display device sends a driving request.

[0721] The driving application includes the expected level and the data to be driven, which are used to instruct the real-time driving of the digital human.

[0722] S34. The server receives the driver application and obtains the current number of concurrent requests.

[0723] The driving application includes the expected level and the data to be driven.

[0724] S35. The server determines the cloud level based on the current number of concurrent connections.

[0725] S36. The server determines the actual level based on the expected level and the cloud level.

[0726] S37. The server determines a target driving scheme according to the actual level and the data to be driven, and sends the target driving scheme to instruct the digital human to be driven using the target driving scheme.

[0727] S38. When receiving the target driving solution, the display device drives the digital human using the target driving solution.

[0728] The specific implementation of the embodiment of the present disclosure is the same as the specific implementation in the digital human driving process executed by the digital human driving device on the display device side and the digital human driving device on the server side. Therefore, the specific implementation method can refer to the specific implementation in the digital human driving process executed by the digital human driving device on the display device side and the digital human driving device on the server side, and will not be repeated here.

[0729] In the above process, after obtaining the text to be driven, the display device determines that it needs to drive the digital human and determines the resource utilization rate. Subsequently, if the resource utilization rate is determined to be less than or equal to the utilization rate threshold and a primary drive solution is matched in the first solution library based on the text to be driven, the primary drive solution is used to drive the digital human. In this way, when the display device determines that it needs to drive the digital human, it first determines its own resource utilization rate and drives the digital human when the resource utilization rate is within the appropriate range. This avoids the situation where the digital human is stuck or cannot be driven due to resource factors, thus improving the user experience. Furthermore, if the resource utilization rate is determined to be within the appropriate range and a primary drive solution can be matched in the first solution library, the primary drive solution is directly used to drive the digital human, saving the resource loss of real-time cloud-based driving and the time loss of network transmission.

[0730] Furthermore, if the display device fails to match a primary drive solution in the first solution library based on the text to be driven, it sends a drive request. Upon receiving the drive request, the server obtains the current concurrency and determines the cloud level based on the current concurrency. It then determines the actual level based on the desired level in the drive request and the cloud level, and determines the target drive solution based on the actual level and the data to be driven. Finally, the server sends the target drive solution to the display device, which then drives the digital human using the target drive solution. In this way, if the display device determines that its resource utilization is within the appropriate range but fails to match a primary drive solution in the first solution library, it sends a drive request to the server, allowing the server to control the display device to drive the digital human in real time. Upon receiving the drive request, the server first determines its own concurrency and, if the concurrency meets the requirements, controls the display device to drive the digital human in real time. This avoids lags and inoperability caused by concurrency issues when driving the digital human, improving the user experience.

[0731] The disclosed embodiments can divide the digital human driving device into functional modules based on the above-described method examples. For example, individual functional modules can be divided according to their functions, or two or more functions can be integrated into a single processing unit. These integrated modules can be implemented as either hardware or software functional modules. It should be noted that the module division in the disclosed embodiments is illustrative and represents only one logical functional division. In actual implementation, other division methods may be employed.

[0732] As shown in Figure 46, an embodiment of the present disclosure further provides a chip system that can be applied to the digital human driver device on the display device side or the digital human driver device on the server side in the aforementioned embodiments. The chip system includes at least one processor 1501 and at least one interface circuit 1502. The processor 1501 and the interface circuit 1502 are interconnected via a circuit. The processor 1501 can receive and execute computer instructions from the digital human driver device on the display device side or the digital human driver device on the server side through the interface circuit 1502. When the computer instructions are executed by the processor 1501, the digital human driver device on the display device side or the digital human driver device on the server side can perform the various steps performed by the digital human driver device on the display device side or the digital human driver device on the server side in the aforementioned embodiments. Of course, the chip system can also include other discrete components, which are not specifically limited by the embodiments of the present disclosure.

Claims

1. A server configured to: Receiving voice data input by a user sent by a display device; Recognize the voice data to obtain a recognition result; If the recognition result includes entity data, obtain media data corresponding to the recognition result and digital human data corresponding to the entity data; wherein the entity data includes a character name and / or a media name, the digital human data includes image data and broadcast voice of the digital human, and the media data includes audio and video data or interface data; The digital human data and the media resource data are sent to the display device, so that the display device plays the audio and video data or displays the interface data, and plays the image and voice of the digital human according to the digital human data.

2. The server according to claim 1, wherein the server is further configured to: Before receiving the voice data input by the user sent by the display device, generating at least one painting model corresponding to the character name, generating at least one action model corresponding to the media asset name, and generating at least one speech synthesis model based on pitch and rhythm corresponding to the character name; Inputting the painting model, the action model and the speech synthesis model into a trained conditional adversarial network to obtain digital human data to be stored; The digital human data to be stored is feature-labeled and stored in a server.

3. The server according to claim 2, wherein the step of generating a drawing model corresponding to at least one character name is performed, wherein the server is specifically configured to: Obtain a preset number of pictures corresponding to the character names; The picture is input into the Wensheng picture model to obtain the painting model corresponding to the character name.

4. The server according to claim 2, wherein the step of generating the action model corresponding to at least one media asset name is performed, wherein the server is specifically configured to: Acquire a preset amount of sample video data, and preprocess and annotate the sample video data; Using the labeled sample video data to train the action generation model; The video data corresponding to the media asset name is input into the trained action generation model to generate the action model corresponding to the media asset name.

5. The server according to claim 2, wherein the step of generating a speech synthesis model based on pitch and prosody corresponding to at least one character name is specifically configured as follows: Obtain a preset amount of sample audio data, and preprocess and annotate the sample audio data; wherein, The sample audio data includes audio data corresponding to the character name and audio data corresponding to the media asset name; The speech synthesis model is trained using the labeled sample audio data to obtain a speech synthesis model based on pitch and rhythm corresponding to the character name.

6. The server according to claim 2, wherein the step of performing feature labeling on the digital human data to be stored and storing the data in the server is specifically configured as follows: The character information, media name and popularity of the digital human data to be stored are marked; The character information includes the character name, and the popularity is the number of training data; Obtaining a first heat and a second heat; wherein the first heat is the highest heat corresponding to the character name in the stored digital human data, and the second heat is the highest heat corresponding to the media asset name in the stored digital human data; If the heat of the digital human data to be stored is not less than the first heat or the second heat, the marked digital human data to be stored is stored in the server.

7. The server according to claim 1, wherein the step of acquiring the digital human data corresponding to the entity data is specifically configured as follows: If the recognition result includes a character name or a media asset name, then the digital human data with features marked as corresponding to the character name or the media asset name in the stored digital human data is obtained.

8. The server according to claim 1, performing the step of acquiring the digital human data corresponding to the entity data, wherein the server is specifically configured to: If the recognition result includes a character name and a media asset name, and the character name and the media asset name do not match in the feature annotation of the stored digital human data, the painting model corresponding to the media asset name is replaced with the painting model corresponding to the character name, and the voice data corresponding to the media asset name is replaced with the voice data corresponding to the character name, to generate replacement digital human data; The replacement digital human data is determined to be the digital human data corresponding to the character name and the media asset name.

9. The server according to claim 1, wherein the server is further configured to: After receiving the voice data sent by the display device, obtaining voice text by recognizing the voice data; Performing semantic understanding on the voice text to obtain the domain intent corresponding to the voice data; Determine the broadcast voice based on the domain intention, and determine the digital human image parameters based on the domain intention; wherein, The digital human image parameters are used to generate the digital human image and / or generate the digital human action; Digital human data is generated based on the digital human image parameters and the broadcast voice; the digital human data is sent to the display device, so that the display device plays the digital human image and voice according to the digital human data.

10. The server according to claim 9, wherein the server is further configured to: Determining a user emotion type corresponding to the voice data; To perform the step of determining the digital human image parameters based on the domain intention, the server may be specifically configured as follows: The digital human image parameters are determined based on the user emotion type and the domain intention.

11. The server according to claim 10, wherein the server determines the user emotion type corresponding to the voice data, and is further configured to: A user emotion type corresponding to the voice data is determined based on the voice data.

12. The server according to claim 10, wherein the step of determining the user emotion type corresponding to the voice data is performed by the server being specifically configured as follows: The user emotion type corresponding to the voice data is determined based on the voice text.

13. The server according to claim 10, wherein the step of determining the user emotion type corresponding to the voice data is performed by the server being specifically configured as follows: Receive user videos uploaded and collected by the display device; The user video includes a user facial image; The user facial image is analyzed to determine the user emotion type corresponding to the voice data.

14. The server according to claim 10, wherein the step of determining the user emotion type corresponding to the voice data is performed by the server being specifically configured as follows: Receive the user's physiological signals collected and uploaded by the display device; wherein, The user's physiological signals include heart rate, skin conductance and / or brain waves; The user emotion type corresponding to the voice data is determined based on the user physiological signal.

15. The server according to claim 9, wherein the step of determining the digital human image parameters based on the domain intention is performed, and the server may be specifically configured as follows: Search the digital human image identification corresponding to the domain intention in the digital human image mapping table; wherein, The digital human image mapping table is used to represent the corresponding relationship between the domain intention and the digital human image identifier; The digital human image parameters corresponding to the digital human image identifier are searched in the digital human definition table; wherein the digital human definition table is used to characterize the corresponding relationship between the digital human image identifier and the digital human image parameters, and the digital human image parameters include decoration parameters and action parameters.

16. The server according to claim 10, wherein the step of determining the digital human image parameters based on the user emotion type and the domain intention is performed, and the server can be specifically configured as follows: Searching the digital human image identification corresponding to the user emotion type and the domain intention in the digital human image mapping table; wherein, The digital human image mapping table represents the correspondence between domain intention, user emotion type and digital human image identifier; Search the digital human image parameters corresponding to the digital human image identifier in the digital human definition table; wherein the digital human definition table is used to characterize the corresponding relationship between the digital human image identifier and the digital human image parameters, and the digital human image parameters include Includes decoration parameters and action parameters.

17. The server according to claim 1, wherein the server is further configured to: After receiving the voice data sent by the display device, the voice data is input into the emotional voice model to obtain the emotional type and emotional intensity; wherein, The speech emotion model is obtained by training based on sample speech data of different groups of people for multiple semantic scenarios; Obtaining the broadcast text corresponding to the voice data; Synthesize the announcement voice based on the announcement text, the emotion type and the emotion intensity; The announcement voice is sent to the display device so that the display device plays the announcement voice.

18. The server according to claim 17, wherein the step of inputting the speech data into a trained emotion speech model to obtain an emotion type and an emotion intensity is performed by: Recognize the voice data to obtain voice text and user voice features; Performing semantic understanding on the speech text to obtain the semantic scene corresponding to the speech data; Converting the user speech feature into a user speech feature vector; Converting the semantic scene into a semantic scene feature vector; Divide the speech data into frames to obtain at least one speech segment sequence; Determine a speech sequence feature vector and an emotion feature vector based on the speech segment sequence; Inputting the user speech feature vector, the semantic scene feature vector, the speech sequence feature vector and the emotion feature vector into a multi-level neural network to obtain an emotion speech vector; wherein the multi-level neural network includes a two-dimensional convolutional network, a recurrent neural network and two fully connected networks; An emotion type and an emotion intensity are determined based on the emotion speech vector.

19. The server according to claim 18, wherein the step of performing semantic understanding on the speech text to obtain a semantic scenario corresponding to the speech data is specifically configured as follows: Performing word segmentation and tagging processing on the speech text to obtain word segmentation information; Performing syntactic analysis and semantic analysis on the word segmentation information to obtain slot information; The semantic scene corresponding to the slot information is located through vertical domain classification.

20. The server according to claim 18, wherein the step of determining a speech sequence feature vector and an emotion feature vector based on the speech segment sequence is performed, wherein the server is specifically configured to: Performing feature extraction on the speech segment sequence to obtain a speech sequence feature vector; Based on the Mel spectrum feature extraction technology, the emotion feature vector corresponding to the speech segment sequence is obtained.

21. The server according to claim 17, wherein the step of acquiring the broadcast text corresponding to the voice data is performed, wherein the server is specifically configured to: Recognize the voice data to obtain voice text; Perform semantic understanding on the speech text to obtain the slot information and semantic scene corresponding to the speech data; The service corresponding to the semantic scene is called to determine the broadcast text corresponding to the slot information.

22. The server according to claim 17, wherein the server is configured to perform the synthesis of the announcement voice based on the announcement text, the emotion type and the emotion intensity: Determining a phoneme sequence corresponding to the broadcast text; Generate an audio feature vector sequence corresponding to the phoneme sequence; Calculate audio feature emotion based on the emotion type and the emotion intensity; Generate a broadcast voice based on the audio feature vector sequence and the audio feature emotion.

23. The server according to claim 17, wherein the server is further configured to: After synthesizing the announcement voice based on the announcement text, the emotion type and the emotion intensity, predicting a key point sequence according to the announcement voice; The digital human image data is synthesized according to the key point sequence and the digital human image data.

24. The server according to claim 23, wherein the step of sending the announcement voice to the display device so that the display device plays the announcement voice, the server is specifically configured as follows: The digital human image data and the broadcast voice are sent to the display device, so that the display device displays the digital human image based on the digital human image data and plays the broadcast voice.

25. The server according to claim 1, wherein the server is further configured to: receiving the voice data of the display device and the digital person identification; wherein, The digital human identifier is used to represent the image and voice features of the digital human selected by the user; Determine user identity information corresponding to the voice data, and obtain voice text by identifying the voice data; Determining the relationship between the digital human and the user based on the digital human identifier and the user identity information; Determine a basic text according to the speech text, where the basic text is obtained by natural language processing of the speech text; Generate a broadcast text based on the basic text and the relationship; Generate digital human data based on the voice features and image data corresponding to the digital human identification and the broadcast text; The digital human data is sent to the display device, so that the display device plays the image and voice of the digital human according to the digital human data.

26. The server according to claim 25, wherein the step of determining the user identity information corresponding to the voice data is performed by: Extracting voiceprint information of the voice data; If the voiceprint information matches the registered voiceprint information in the voiceprint database, the user identity information is determined according to the registered voiceprint information.

27. The server according to claim 25, wherein the step of determining the basic text according to the speech text is performed, wherein the server is specifically configured to: Performing word segmentation and tagging processing on the speech text to obtain word segmentation information; Performing syntactic analysis and semantic analysis on the word segmentation information to obtain slot information; Locate the domain intention corresponding to the slot information through vertical domain classification; A base text is determined based on the domain intent and the slot information.

28. The server according to claim 25, wherein the step of generating the announcement text based on the basic text and the relationship is specifically configured as follows: Obtaining the splicing information corresponding to the relationship; wherein, The splicing information includes a splicing position and a splicing content, the splicing position includes a pre-splicing, and the splicing content corresponding to the pre-splicing is a title set according to the relationship; Generate a broadcast text based on the splicing information and the basic text.

29. The server according to claim 28, wherein the splicing position further comprises post-splicing, and the server is specifically configured to: Get the user's age; Based on the age and the basic text, splicing content corresponding to the post-splicing is determined.

30. The server according to claim 25, wherein the server generates a broadcast text based on the basic text and the relationship, and is configured to: If the detected date is a target date and the target date is related to the relationship, the target text is determined according to the relationship; wherein, The target date is a holiday and / or anniversary, and the target text includes a blessing text and / or a reminder text; The target text is added to the basic text to obtain the broadcast text.

31. The server according to claim 25, wherein the step of generating the announcement text based on the basic text and the relationship is specifically configured as follows: If the date is detected to be a target date and the target date is related to the user, a target text is generated; wherein, The target date is a holiday and / or anniversary; The target text is added to the basic text to obtain the broadcast text.

32. The server according to claim 25, wherein the server is further configured to: After receiving the upload timeout message of the display device, a prompt text is generated based on the relationship and the target scene; wherein, Place The timeout message is sent to the server by the display device after detecting that the time of entering the target scene exceeds the preset time; Generate digital human data based on the voice features and image data corresponding to the digital human identification and the prompt text; The digital human data is sent to the display device, so that the display device plays the image and data of the digital human according to the digital human data.

33. The server according to claim 1, wherein the server is further configured to: Establishing connection relationships with a display device and a terminal respectively, so that the display device is associated with the terminal; After receiving the image data and audio data uploaded by the terminal, determining the image data of the digital human based on the image data, and determining the voice features of the digital human based on the audio data; Sending the digital human image data to a display device associated with the terminal, so that the display device displays a digital human image based on the digital human image data; After the digital human image is selected by the user, receiving the voice data input by the user sent by the display device; Generate a broadcast text according to the voice data; Generate digital human data based on the broadcast text, the digital human voice features and the digital human image data; The digital human data is sent to the display device, so that the display device plays the image and voice of the digital human according to the digital human data.

34. The server according to claim 33, wherein the server is further configured to: Establish a long connection with the display device; receiving request data sent by the display device; wherein, The request data includes a device identification of the display device; If the identification code corresponding to the device identification exists in the database, the identification code is sent to the display device so that the display device displays the identification code.

35. The server according to claim 34, wherein the server is further configured to: receiving the identification code uploaded by the terminal; wherein, The identification code is obtained from the display interface of the display device; If there is a display device corresponding to the identification code, an association relationship between the terminal and the display device is established so that the data uploaded by the terminal is processed by the server and then sent to the display device.

36. The server according to claim 33, wherein the digital human data comprises digital human image data and broadcast voice; and the server is specifically configured to: Synthesize the broadcast voice according to the digital human voice features and the broadcast text; Determine a key point sequence according to the broadcast voice; The digital human image data is synthesized according to the key point sequence and the digital human image data.

37. The server according to claim 33, executing the image data uploaded by the receiving terminal, the server is specifically configured to: Receiving image data uploaded by the terminal; If it is detected that the face points in the image data are qualified, a message of image detection qualification is sent to the terminal; If it is detected that the face points in the image data are unqualified, an image detection unqualified message is sent to the terminal, so that the terminal prompts the user to upload again.

38. The server according to claim 33, wherein the server is further configured to: Before receiving the audio data uploaded by the terminal, receiving the ambient recorded sound uploaded by the terminal; If it is detected that the environmental sound recording is qualified, a message of environmental sound qualification is sent to the terminal; If it is detected that the environmental sound recording is unqualified, a message that the environmental sound is unqualified is sent to the terminal, so that the terminal prompts the user to re-record.

39. The server according to claim 38, executing the audio data uploaded by the receiving terminal, the server is specifically configured to: Receiving audio of a user reading a target text, and identifying the user text corresponding to the audio; Calculating a pass rate based on the target text and the user text; If the pass rate is less than a preset value, a voice upload failure message is sent to the terminal, so that the terminal prompts the user to re-record the audio of reading the target text; If the pass rate is not less than a preset value, a voice upload success message is sent to the terminal so that the terminal displays the next target text or voice recording completion information.

40. The server according to any one of claims 1 to 39, wherein the server is further configured to: Receive driver application and obtain the current concurrent number; among them, The driving application includes the expected level and the data to be driven; Determine the cloud level according to the current number of concurrent connections; determining an actual level according to the desired level and the cloud level; A target driving scheme is determined according to the actual level and the data to be driven, and the target driving scheme is sent to a display device to instruct the display device to drive the digital human using the target driving scheme.

41. The server according to claim 40, wherein the driving request is sent by the display device after acquiring the text to be driven and determining the resource occupancy rate, detecting that the source occupancy rate is less than or equal to the occupancy rate threshold, and no primary driving scheme is matched in the first scheme library according to the text to be driven; in, The text to be driven is obtained by converting the data to be driven. The first solution library includes driving texts and driving solutions. One driving text corresponds to one driving solution.

42. The server according to claim 41, wherein the server is further configured to: Feedback the application result including the actual level and driving time to the display device, so that the display device determines the next expected level according to the driving time, the actual time, and the actual level; the next expected level is used to indicate real-time driving of the digital human; wherein, The actual time consumption is the time between sending the driving application and receiving the application result, and the next expected level is used to indicate the real-time driving of the digital human.

43. A display device comprising: A display configured to display an image and / or a user input interface: A user input interface configured to receive instructions from a user; A Bluetooth module, configured to perform operations related to the Bluetooth protocol; a communication device configured to communicate with an external device according to a predetermined protocol; a memory configured to store computer instructions and data associated with a display device; At least one processor, connected to the display, the user input interface, the Bluetooth module, the communication device and the memory, is configured to execute computer instructions to cause the display device to perform: Receiving voice data input by a user; sending the voice data to a server via the communication device; Receiving digital human data sent by the server based on the voice data; The image and voice of the digital human are played according to the digital human data.

44. A digital human processing method, comprising: Receiving voice data input by a user sent by a display device; Recognize the voice data to obtain a recognition result; If the recognition result includes entity data, obtain media data corresponding to the recognition result and digital human data corresponding to the entity data, wherein the entity data includes a character name and / or a media name, the digital human data includes image data and broadcast voice of the digital human, and the media data includes audio and video data or interface data; The digital human data and the media resource data are sent to the display device, so that the display device plays the audio and video data or displays the interface data, and plays the image and voice of the digital human according to the digital human data.

Citation Information

Cited By

  • Data depersonalization using on-device generative models

    US20260119700A1