A data processing method, device and electronic equipment
By receiving voice information, recognizing and generating rendered images of virtual digital humans, the problem of not being able to dynamically display response information in existing technologies is solved, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-14
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, playing response information via voice can no longer meet users' needs for dynamic display. How to use data virtual humans to broadcast response information has become an urgent problem to be solved.
By receiving voice information, recognizing the response information and inputting it into the text-driven model, the target key point set is determined, and a rendered image of a virtual digital human is generated, including key points of the mouth and eyes, to achieve dynamic broadcasting.
It enables dynamic broadcasting of responses through a virtual data persona, thus improving the user experience.
Smart Images

Figure CN115617162B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of human-computer interaction technology, and in particular to a data processing method, apparatus and electronic device. Background Technology
[0002] Currently in the field of human-computer interaction technology, users can send control, question-and-answer, and casual conversation requests to electronic devices via voice. The electronic device determines the corresponding response based on the received voice information and then plays the response aloud. However, with technological advancements, simply playing responses via voice is no longer sufficient for users. Users are increasingly seeking ways to dynamically display responses, such as using virtual human assistants to deliver them.
[0003] Therefore, how to use virtual data humans to broadcast and respond to information has become an urgent problem to be solved. Summary of the Invention
[0004] To address the aforementioned technical problems, this disclosure provides a data processing method, apparatus, and electronic device.
[0005] The technical solution disclosed herein is as follows:
[0006] In a first aspect, this disclosure provides a data processing method, comprising: receiving voice information sent by an electronic device to trigger human-computer interaction; recognizing the voice information to determine response information; inputting the response information into a text-driven model to determine a target key point set; wherein the target key point set includes at least key points for indicating the mouth and eyes; sending target information carrying the response information and the target key point set to the electronic device, so that the electronic device generates a rendered image of the virtual digital human based on a preset key point set corresponding to the face of the virtual digital human, the response information, and the target key point set, wherein the preset key point set includes at least key points for indicating the mouth and eyes.
[0007] Secondly, this disclosure provides a data processing method, comprising: in response to a selection operation of a target function, sending voice information to a server; receiving response information sent by the server; synthesizing the response information to determine the voice information corresponding to the response information; inputting the voice information into a voice-driven model to determine a specified set of key points; wherein the specified set of key points includes at least key points for indicating the mouth and eyes; generating a rendered image of the virtual digital human based on a preset set of key points corresponding to the face of a pre-configured virtual digital human, the response information, and a target set of key points; wherein the preset set of key points includes at least key points for indicating the mouth and eyes.
[0008] Thirdly, this disclosure provides a data processing apparatus, comprising: a receiving unit for receiving voice information sent by an electronic device to trigger human-computer interaction; a processing unit for recognizing the voice information received by the receiving unit and determining response information for the voice information; the processing unit is further configured to input the response information into a text-driven model to determine a target key point set; wherein the target key point set includes at least key points for indicating the mouth and eyes; the processing unit is further configured to control a sending unit to send target information carrying the response information and the target key point set to the electronic device, so that the electronic device generates a rendered image of the virtual digital human based on a preset key point set corresponding to the face of a pre-configured virtual digital human, the response information, and the target key point set, wherein the preset key point set includes at least key points for indicating the mouth and eyes.
[0009] Fourthly, this disclosure provides a data processing apparatus, comprising: a processing unit, configured to control a sending unit to send voice information to a server in response to a selection operation of a target function; a receiving unit, configured to receive response information sent by the server; a processing unit, configured to synthesize the response information received by the receiving unit and determine the voice information corresponding to the response information; the processing unit is further configured to input the voice information into a voice-driven model and determine a specified set of key points; wherein the specified set of key points includes at least key points for indicating the mouth and eyes; the processing unit is further configured to generate a rendered image of a virtual digital human based on a pre-configured set of preset key points corresponding to the face of a virtual digital human, the response information, and the specified set of key points; wherein the preset set of key points includes at least key points for indicating the mouth and eyes.
[0010] Fifthly, this disclosure provides an electronic device, including: a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to cause the electronic device to implement any of the data processing methods provided in the first aspect above when executing the computer program.
[0011] In a sixth aspect, this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a computing device, causes the computing device to implement any of the data processing methods provided in the first aspect above.
[0012] In a seventh aspect, the present invention provides a computer program product that, when run on a computer, causes the computer to perform a data processing method as provided in any of the first aspects.
[0013] Eighthly, this disclosure provides an electronic device, including: a memory and a processor, the memory being used to store a computer program; the processor being used to cause the electronic device to implement any of the data processing methods provided in the second aspect above when executing the computer program.
[0014] Ninthly, this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a computing device, causes the computing device to implement any of the data processing methods provided in the second aspect above.
[0015] In a tenth aspect, the present invention provides a computer program product that, when run on a computer, causes the computer to perform a data processing method as provided in any of the second aspects.
[0016] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on the first computer-readable storage medium. The first computer-readable storage medium may be packaged together with the processor of the data processing device, or it may be packaged separately from the processor of the data processing device; this disclosure does not impose any limitations on this.
[0017] The descriptions of the third, fifth, sixth, and seventh aspects in this disclosure can be referenced to the detailed description of the first aspect; and the beneficial effects of the descriptions of the third, fifth, sixth, and seventh aspects can be referenced to the analysis of the beneficial effects of the first aspect, which will not be repeated here.
[0018] The descriptions of aspects four, eight, nine, and ten in this disclosure can be referenced to the detailed description of aspect one; and the beneficial effects of the descriptions of aspects four, eight, nine, and ten can be referenced to the analysis of the beneficial effects of aspect two, which will not be repeated here.
[0019] In this disclosure, the names of the aforementioned data processing devices do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear under other names. As long as the functions of each device or functional module are similar to those of this disclosure, they fall within the scope of the claims of this disclosure and their equivalents.
[0020] These or other aspects of this disclosure will become more readily apparent in the following description.
[0021] The technical solution provided in this disclosure has the following advantages compared with the prior art:
[0022] The data processing method provided in this disclosure, when the target function is a voice interaction function, allows the user to select and input corresponding voice information on an electronic device when needed. Upon receiving the voice information, the electronic device sends it to a server. The server then recognizes the voice information and determines the response. The response is then input into a text-driven model to obtain a set of target key points corresponding to the response, such as key points indicating the mouth. Next, target information carrying the response, the set of target key points, and a preset set of key points corresponding to the virtual digital human's face is sent to the electronic device. This allows the electronic device to generate a rendered image of the virtual digital human based on the response, the set of target key points, and the preset set of key points. In this way, the electronic device can dynamically play the response information corresponding to the user's input voice information, thus solving the problem of how to broadcast response information using a virtual digital human. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0024] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 One of the flowcharts illustrating the data processing method provided in the embodiments of this application;
[0026] Figure 2 This is one of the structural schematic diagrams of the display device in the data processing method provided in the embodiments of this application;
[0027] Figure 3 This is a second schematic diagram of the structure of the display device in the data processing method provided in the embodiments of this application;
[0028] Figure 4 One of the flowcharts illustrating the data processing method provided in the embodiments of this application;
[0029] Figure 5 A second schematic flowchart illustrating the data processing method provided in this application embodiment;
[0030] Figure 6 The third schematic flowchart of the data processing method provided in the embodiments of this application;
[0031] Figure 7 The fourth flowchart illustrating the data processing method provided in the embodiments of this application;
[0032] Figure 8 Fifth flowchart illustrating the data processing method provided in the embodiments of this application;
[0033] Figure 9 A flowchart illustrating the data processing method provided in this application embodiment is shown in Figure 6.
[0034] Figure 10 Seventh schematic flowchart of the data processing method provided in the embodiments of this application;
[0035] Figure 11 This is a schematic diagram of the server structure provided in an embodiment of this application;
[0036] Figure 12 This is one of the schematic diagrams of a chip system provided in an embodiment of this application;
[0037] Figure 13 This is a schematic diagram of the structure of a display device provided in an embodiment of this application;
[0038] Figure 14 This is a second schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation
[0039] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0040] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0041] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0042] MoviePy, as described in this disclosure, is a Python module for video editing that can be used for basic operations such as cutting, splicing, and inserting titles, video compositing (i.e., non-linear editing), video processing, and creating advanced effects.
[0043] AudioTrack in this embodiment is a sound playback scheme in Android, which can be used to play Pulse Code Modulation (PCM) data streams.
[0044] Pyscenedetect, as described in this embodiment, is a powerful video scene detection library that can detect all scenes in a given video and save them as separate video files.
[0045] The VideoFileClip class in this embodiment is a direct subclass of VideoClip. It is a clip class created from a video file. In addition to the features and methods inherited from the parent class, VideoFileClip implements its own constructor and close method.
[0046] The Dlib tool in this embodiment refers to the dlib library.
[0047] The Transformer structure in this embodiment is a novel neural network architecture that is based solely on the attention mechanism, abandoning traditional recurrent or convolutional neural network structures. Attention is a mechanism in neural networks where the model learns to make predictions by selectively focusing on a given dataset. The number of attention points is quantified by learning weights, and the output is typically a weighted average.
[0048] In this embodiment of the disclosure, FLOAT refers to a floating-point data type, which is used to store single-precision floating-point numbers or double-precision floating-point numbers.
[0049] In this embodiment of the disclosure, deepspeech is an engine based on a deep learning framework for processing speech-to-text.
[0050] The Reference loss function in this embodiment is used to estimate the degree of inconsistency between the model's predicted value f(X) and the true value Y.
[0051] The reshape function in this embodiment is a MATLAB function that transforms a specified matrix into a matrix of a specific dimension, while keeping the number of elements in the matrix unchanged. The function can readjust the number of rows, columns, and dimensions of the matrix.
[0052] The OpenGL ES (OpenGL for Embedded Systems) in this disclosure is a subset of the OpenGL 3D graphics application programming interface (API) designed for embedded devices such as mobile phones, PDAs, and game consoles.
[0053] MediaPipe in this disclosure is a framework for building machine learning pipelines to process time-series data such as video and audio.
[0054] In this embodiment of the disclosure, Reshape is a function in a scripting language used to transform dimensions.
[0055] Figure 1 This is a schematic diagram illustrating an operational scenario between a display device and a control device according to one or more embodiments of this application, such as... Figure 1 As shown, a user can operate the display device 200 via a mobile terminal 300 and a control device 100. The control device 100 can be a remote control, and communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, wireless or other wired methods to control the display device 200. The user can input user commands through buttons on the remote control, voice input, control panel input, etc., to control the display device 200. In some embodiments, a mobile terminal, tablet computer, computer, laptop computer, and other smart devices can also be used to control the display device 200.
[0056] In some embodiments, the mobile terminal 300 can install software applications with the display device 200 to achieve connection and communication via network communication protocols, enabling one-to-one control operations and data communication. Audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 for synchronous display. The display device 200 can also communicate with the display device 200 via various communication methods. It can be allowed to communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The display device 200 can provide various content and interactive features. The display device 200 can be a liquid crystal display, an OLED display, or a projection display device. In addition to providing broadcast television reception functions, the display device 200 can also be equipped with a smart network television function that provides computer support.
[0057] In some embodiments, the electronic device provided in this application can be the aforementioned display device 200. When a user needs to use the voice interaction function, they can press the voice button on the control device 100 to control the display device 200 to activate the target function (such as the voice interaction function), and then input voice information (such as: "How's the weather today?") while pressing the voice control button. Alternatively, when a user needs to use the voice interaction function, they can select the target function through an electronic device (such as a mobile phone) that has established a communication connection with the display device 200, and then input voice information. In response to the selection of the target function (such as the voice interaction function), the display device 200 sends voice information to the server 400 to trigger human-computer interaction. After receiving the voice information sent by the display device 200, the server 400 uses Automatic Speech Recognition (ASR) technology to recognize the voice information and obtain the corresponding response information. Then, the server 400 converts the response information to obtain at least one word vector contained in the response information. Then, the server 400 performs intent understanding based on the word vector to obtain the intent recognition result. Based on the intent recognition result, server 400 determines the response content (e.g., "Today's weather is sunny, minimum temperature 19°C, maximum temperature 28°C"). Then, server 400 sends a response message carrying this content to display device 200. Upon receiving the response message from server 400, display device 200 synthesizes the response message (e.g., using Text-to-Speech (TTS)) to determine the corresponding speech information; it inputs the speech information into a speech-driven model to determine a specified set of key points; and it generates a rendered image of the virtual digital human based on a pre-configured set of key points corresponding to the virtual digital human's face, the response message, and the target set of key points. Thus, display device 200 can dynamically respond to the user using the virtual digital human: "Today's weather is sunny, minimum temperature 19°C, maximum temperature 28°C," enhancing the user experience.
[0058] Figure 2 A hardware configuration block diagram of a display device 200 according to an exemplary embodiment is shown. For example... Figure 2The display device 200 shown includes at least one of the following: a tuner / demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280. The controller includes a central processing unit, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first to nth interface for input / output. The display 260 may be a touch-enabled display, such as a touch screen display. The tuner / demodulator 210 receives broadcast television signals via wired or wireless reception and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals. The detector 230 is used to collect signals from the external environment or signals interacting with the external environment. The controller 250 and the tuner / demodulator 210 may be located in different separate devices; that is, the tuner / demodulator 210 may also be located in an external device of the main device containing the controller 250, such as an external set-top box.
[0059] In some embodiments, the controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200.
[0060] In some examples, the display device 200 of one or more embodiments is a television 1, and the operating system of the television 1 is the Android system, for example... Figure 3 As shown, TV 1 can be logically divided into an application layer (referred to as "application layer") 21, an application framework layer (referred to as "framework layer") 22, an Android runtime and system library layer (referred to as "system runtime library layer") 23, and a kernel layer 24.
[0061] The application layer 21 includes one or more applications. These applications can be system applications or third-party applications. For example, application layer 21 may include a first application that provides voice interaction functionality. The framework layer 22 provides application programming interfaces (APIs) and programming frameworks for the applications in application layer 21. The system runtime library layer 23 provides support for the upper layer, namely the framework layer 22. When the framework layer 22 is used, the Android operating system runs the C / C++ libraries contained in the system runtime library layer 23 to implement the functions required by the framework layer 22. The kernel layer 24 acts as software middleware between the hardware layer and application layer 21, managing and controlling hardware and software resources.
[0062] In some examples, kernel layer 24 includes a first driver and a second driver. The first driver is used to send user operations collected by detector 230 to a first application, and the second driver is used to control display 260 to display the display information sent by display unit 213.
[0063] The first application on television set 1 is launched. Then, the first driver sends the user operation collected by detector 230 to the first application for recognition. Subsequently, the processing unit 212 of the first application, in response to the target operation received by the receiving unit 210 (e.g., a selection operation for a voice interaction function), controls the sending unit 211 to send voice information (e.g., "How's the weather today?") to the server 400 to trigger human-computer interaction. After receiving the voice information sent by the display device 200, the receiving unit 410 of the server 400 uses Automatic Speech Recognition (ASR) to recognize the voice information received by the receiving unit 410, obtaining the corresponding voice text. Then, the processing unit 411 of the server 400 converts the voice text to obtain at least one word vector contained in the voice text. Then, the processing unit 411 of the server 400 performs intent understanding based on the word vector, obtaining the intent recognition result. Based on the intent recognition result, the processing unit 411 of server 400 determines the response content (e.g., "Today's weather is sunny, minimum temperature 19°C, maximum temperature 28°C"). Then, the processing unit 411 of server 400 controls the sending unit 412 to send a response message carrying this response content to the display device 200. After receiving the response message sent by server 400, the receiving unit 210 of display device 200 synthesizes the response message received by the receiving unit 210 (e.g., using Text-to-Speech (TTS) to synthesize the response message), determining the corresponding voice information; the processing unit 211 inputs the voice information into a voice-driven model to determine a specified set of key points; the processing unit 211 generates a rendered image of the virtual digital human based on a pre-configured set of key points corresponding to the virtual digital human's face, the response message, and the specified set of key points. Thus, display device 200 can dynamically respond to the user through the virtual digital human: "Today's weather is sunny, minimum temperature 19°C, maximum temperature 28°C," enhancing the user experience.
[0064] Specifically, the storage unit 413 in server 400 can be used to store the data to be stored in server 400 and the application program of server 400.
[0065] Specifically, the storage unit 213 in the display device 200 can be used to store the data to be stored in the display device 200 and the application program of the display device 200.
[0066] Specifically, the electronic device provided in this embodiment may be the aforementioned display device 200 or the server 400, and no limitation is made here.
[0067] All voice information involved in this application may be data authorized by the user or fully authorized by all parties.
[0068] In the following embodiments, the server 400 described above is used as the execution subject for the data processing method provided in the embodiments of this disclosure to illustrate the method of the present application.
[0069] This application provides a data processing method, such as... Figure 4 As shown, the data processing method may include S11-S14.
[0070] S11. Receive voice information sent by an electronic device to trigger human-computer interaction.
[0071] In some examples, users trigger services such as voice control, queries, question-and-answer sessions, and casual conversation by selecting a target function offered by the electronic device. In response to the selection of the target function, the electronic device sends voice information to the server, which may include any of the interactive commands such as voice control instructions, query instructions, question-and-answer instructions, or casual conversation instructions.
[0072] S12. Recognize the voice information and determine the response information.
[0073] In some examples, when server 400 recognizes voice information, it can determine the corresponding response information based on keywords in the voice information. For example, if the voice information is "How's the weather today?", server 400 will determine that the keywords in the voice information are "today", "weather", and "how". Then, based on the keywords, server 400 will determine that the voice information is a question-and-answer type voice message. Server 400 will query "weather" and look for information related to "today" within "weather". Finally, it will generate a response information related to "today" regarding "weather", such as: "Today's weather is sunny, with a low of 19°C and a high of 28°C."
[0074] Alternatively, server 400 first uses ASR to recognize the speech information and determine the corresponding speech text, such as "How's the weather today?". Then, server 400 segments the speech text according to a preset dictionary to determine at least one actual word contained in the speech text. Next, based on the probability of occurrence of each actual word in its corresponding historical segments, it determines the historical segment with the highest probability.
[0075] If the voice information corresponding to the currently logged-in user account on the electronic device contains the historical word segment, query the historical reply information corresponding to the historical word segment, and generate reply information based on the historical reply information.
[0076] If the historical word segment is not present in the voice information corresponding to the currently logged-in user account on the electronic device, query at least one response message corresponding to the historical word segment. Randomly select one response message from the responses and use it as the reply message. The preset dictionary contains at least one word segment.
[0077] Specifically, historical word segmentation is obtained by segmenting historical voice information sent by electronic devices.
[0078] S13. Input the response information into the text-driven model to determine the target keypoint set. The target keypoint set must include at least the keypoints used to indicate the mouth and eyes.
[0079] In some examples, facial expressions vary when a person speaks. For instance, the degree of mouth opening and closing differs depending on the type of speech. Therefore, by acquiring video data showing a person's face clearly and articulating their speech (such as a news broadcast video of a specific anchor), the text-driven model can be trained to identify key points in the anchor's facial image when they speak different speech messages. Then, when subsequent responses are input into the text-driven model, it can also identify these key points in the anchor's facial image when they are responding to different messages.
[0080] Furthermore, a virtual digital human can be viewed as a still image. To generate playable animation from this still image, specific parts of the virtual digital human can be controlled to move, thus achieving the animation effect. For example, the specified part could be the mouth. Since the face predicted by the text-driven model may differ from the standard face of the virtual digital human, it is necessary to map the key points predicted by the text-driven model to indicate the mouth onto the standard face of the virtual digital human. This allows the virtual digital human's mouth to read out the actual response content, improving the user experience.
[0081] S14. Send target information carrying response information and a target key point set to the electronic device, so that the electronic device can generate a rendered image of the virtual digital human based on the preset key point set corresponding to the face of the virtual digital human, the response information, and the target key point set. The preset key point set includes at least key points for indicating the mouth and eyes.
[0082] In some examples, the types of keypoints in the target keypoint set are the same as those in the preset keypoint set. For example, the data processing method provided in this embodiment uses a face recognition algorithm (such as Deep-SDM, Dlib, Face Alignment, MediaPipe) to extract N (e.g., 68, 106, 478) keypoints from a face image, and then aligns the keypoints in the spatial dimension. For example, first, the 106 keypoints are scaled to a certain value range in space according to a certain ratio, and then a standard face is selected (the midline of the eyes and nose is at x=0). By rotation and translation, the 106 keypoints are maximized to overlap with the keypoints set at the corners of the eyes and the bridge of the nose on the standard face. Finally, the coordinates of the spatially aligned face keypoints are obtained, which is the preset keypoint set.
[0083] In some examples, after obtaining the target keypoint set, the server can also calculate the difference (also known as the keypoint offset) between each keypoint in the target keypoint set and the initial feature point corresponding to the specified person's face image in a static state. Then, the keypoint offset and response information are sent to the electronic device. In this way, the electronic device can obtain the corrected keypoints based on each keypoint in the pre-configured keypoint set corresponding to the virtual digital human's face and the sum of the keypoint offsets corresponding to that keypoint. These corrected keypoints are then used as vertex coordinates to generate the rendered image of the virtual digital human.
[0084] Specifically, when displaying rendered images of virtual digital humans, electronic devices can control the speed at which audio byte stream data is transmitted and the speed at which vertex coordinates are updated by the AudioTrack player to drive image and audio alignment, thereby ensuring the realism of the virtual digital human.
[0085] Specifically, when the electronic device generates a rendered image of the virtual digital human based on a pre-configured set of preset key points corresponding to the virtual digital human's face, response information, and a target set of key points, the electronic device can determine at least one vertex coordinate based on the actual coordinates of the key points in the target set of key points. Each key point in the preset set of key points is used as a texture coordinate to determine at least three texture coordinates. The texture coordinates are then processed using the Delaunay triangulation algorithm to determine texture triangles. Based on the vertex coordinates, texture coordinates, texture triangles, and response information, the rendered image of the virtual digital human is generated.
[0086] In some feasible examples, combining Figure 4 ,like Figure 5As shown, the data processing method provided in this embodiment further includes: S15-S17.
[0087] S15. Obtain the training sample video and the labeling results of the training sample video. This includes the actual text data corresponding to the specified person speaking in the training sample video, and the facial image corresponding to the actual text data of the specified person during the speaking process. The labeling results include the actual key points of the facial image.
[0088] In some examples, the training sample videos can be news broadcasts, talk shows, or videos of online anchors commenting. These videos are required to contain clear, frontal images with visible lips and teeth.
[0089] The following example uses training sample videos as news broadcast data to illustrate the process of obtaining news broadcast data:
[0090] The Uniform Resource Locator (URL) of a specific webpage is extracted using regular expressions, and web crawling tools such as You-Get are used to crawl the corresponding URL links to videos.
[0091] Next, the crawled video data is segmented into scenes. For example, the data processing method provided in this embodiment requires frontal broadcast video data, but for news broadcasts, in addition to the presenter's reporting, other scenes are interspersed, so scene segmentation of the video is necessary. The data processing method provided in this embodiment uses PySceneDetect to obtain the start time points of different scenes (there are many types of scene segmentation packages, this is just one example). Then, based on the obtained time points, the video is segmented using the VideoFileClip toolkit in moviepy, and each segmented video is video data of the same scene.
[0092] Finally, the video data after scene segmentation is filtered. For example, using the face detector and face recognizer in Dlib, two frames are randomly selected from a video segment. The video is retained if the detected structure meets the following conditions, otherwise it is deleted:
[0093] A. Only one head was detected.
[0094] B. The center point of the head detection box in two random head detections is no more than 20 pixels.
[0095] C. The face recognition results for the same person were obtained in both tests.
[0096] After the above operations, the training sample videos are obtained.
[0097] Furthermore, since the target model cannot directly recognize the training sample videos, the training sample videos need to be processed. For example, the speech data in the training sample videos can be converted into corresponding text data. For instance, the data processing method provided in this embodiment uses MoviePy to extract audio data from the training sample videos, and then uses LibreSa to convert the audio data into mono 16kHz sampling rate speech data. Afterwards, the speech data is recognized using the Wenet speech recognition model to obtain the corresponding text data. Then, the text data is validated, and the phonemes corresponding to the validated text data are aligned with the corresponding pre-defined data to obtain the duration of each phoneme. Finally, each phoneme is encoded to obtain a feature vector.
[0098] Then, the phoneme feature vector is input into the target model, so that the target model can learn the facial key points corresponding to each frame of the training sample video based on the factor feature vector, and thus the target model can predict the facial key points.
[0099] S16. Input the initial feature points and training text data corresponding to the facial image of the specified person in a static state into the target model to determine the prediction key points predicted by the target model.
[0100] In some examples, a specified person’s face image in a static state includes the specified person’s face image in a static state with no expression, open eyes, and closed mouth.
[0101] In some examples, the target model uses a Transformer-based FastSpeech / FastSpeech2 model architecture (used for speech synthesis). When feeding training text data into the target model, the response information needs to be processed to determine the phonemes corresponding to the responses. Phonemes are encoded to determine the phoneme feature vector for each phoneme. These phoneme feature vectors are then used as input to the target model for supervised learning of the labels (e.g., the sequence of keypoint coordinates for each frame of the training sample video obtained through keypoint detection) and the facial keypoints detected in each frame of the video. Finally, the model outputs the predicted keypoints of the face.
[0102] Specifically, compared to the speech-driven model, the text-driven model replaces the MLP+BiLSTM model architecture in the speech-driven model with the FastSpeech speech synthesis model architecture; removes the speech synthesis part in the speech-driven model and replaces the Mel spectrum synthesis part in the speech-driven model with keypoint synthesis. Apart from the differences in input features and model, the other settings are the same as the speech-driven model.
[0103] Compared to speech-driven models, text-driven models use a Transformer structure, which results in better continuity and stability in the generation of key points.
[0104] Due to the large number of parameters and the complexity of the FastSpeech model, it is currently impossible to achieve edge-side deployment and needs to be deployed through the cloud.
[0105] 1. Phoneme time dimension alignment: The phoneme interval ct ∈ {"pad", " / ", "AA0"... "x", "z", "zh"}, a total of 67 English phonemes, 120 Chinese phonemes plus 4 special symbols (space, silence, padding, prosody boundary), a total of 191. For a piece of speech data with a length of T and a sampling rate of 16,000, it is first processed into Mel Frequency Cepstrum Coefficient (MFCC). The time window is 200 and the dimension is 39, obtaining a MFCC[T×80, 39] dimension. The text corresponding to the speech is discretized into a phoneme list Phonemes, e.g., "你好" -> {eos n i2 h aa3 uu3eos}. Then we use the Viterbi forced alignment of the MFCC[T×80, 39] and the corresponding Phonemes based on the Time-Delay Neural Network (TDNN) to obtain the time list Durations{n1, n2... nm} corresponding to the Phonemes. m is the same as the length of Phonemes, and n1 + n2 +... nm = T×80.
[0106] 2. Key point space alignment: The video frame rate is 25Fps. Extract the key points of the face image, e.g., use Deep-SDM to extract 106 key points of the face. Then align the key points in the space dimension. The alignment method is to first scale the key points to a certain value range in space according to a certain ratio, and then select a standard face (the mid-axis of the eyes and nose is at x = 0). All key points are rotated and translated to achieve the maximum overlap with the key points at the corners of the eyes and N (e.g., 8) groups of key points on the nose bridge of the standard face. Finally, the coordinates of the face key points after space alignment Fls[T×25, 318] are obtained.
[0107] 3. Audio-visual alignment: Through the above two steps, we obtain a. Phonemes({eos n i2 h aa3 uu3 eos}), b. Durations{n1, n2…nm}(n1+n2+…nm=T×80) and c. Fls[T×25, 318]. Before entering the model, 80×T and 25×T need to be aligned in the time dimension. The data processing method provided in this embodiment uses bilinear interpolation to map 25×T×318 keypoints to 80×T×318 keypoints to achieve alignment. Finally, we obtain the Input{Phonemes, Durations, Face Landmarks (Fls)} required for model input.
[0108] 4. Keypoint-Driven: The phoneme-driven model uses the FastSpeech model based on the Transformer structure. The training data consists of 6 hours of news broadcast video data extracted as Input{Phonemes, Durations, Fls}. The input data is: Phonemes:{ph1, ph2...phm}, Durations{n1, n2...nm}, and Fls[T×80, 318]. The output is the Delta[T×80, 318] dimension of the keypoint offset relative to the reference keypoint. The Loss is set to Delta + reference, and the Mean Square Error (MSE) is calculated using Fls.
[0109] 5. Mouth-Weighted Loss (Optional): Because closed-mouth lip shapes like "b, p, m, f" are difficult to fit in speech-driven keypoint tasks, a larger loss weight is needed for similar lip shapes. The generated Delta + reference is used to obtain the generated facial keypoints F, with dimensions [T×80, 318]. F is reshaped to [T×80, 106, 3]. In the second dimension, 101 represents the upper keypoint of the mouth, and 104 represents the lower keypoint. The distance between two keypoints is calculated to obtain the mouth distance matrix D [T×80, 1]. Since these values di are all numbers greater than 0 and less than approximately 0.2, we want the smaller di to be, the greater the weight. Therefore, the final weight W = 1 / (D×lamda + 0.01), where lambda is the coefficient, which we set to 1. Adding 0.01 ensures that there will be no division by zero. Finally, the mouth-weighted loss = MSE×W
[0110] 6. By training, we obtain a text-driven model. During the inference process, when a piece of text is input, we will find the corresponding phoneme sequence and duration sequence through the phoneme library and input them into the text-driven model to obtain the target key point set of the corresponding speech segment.
[0111] S17. Based on the predicted key points and the actual key points, adjust the network parameters of the target model until the target model converges to obtain the text-driven model.
[0112] In some feasible examples, combining Figure 5 ,like Figure 6 As shown, the above S17 can be specifically implemented through the following S170-S173.
[0113] S170. Based on the first objective loss function, predicted key points, and actual key points, determine the area difference corresponding to the key points used to indicate the mouth.
[0114] S171. Match the second objective loss function, predicted key points, and actual key points to determine the distance between two matched key points.
[0115] S172. Based on the third objective loss function, the predicted key points and the actual key points are matched to determine the gradient difference between two matched key points.
[0116] In some examples, the first, second, and third objective loss functions can be regression loss functions, such as L1 loss. S173. Based on the area difference, distance, and gradient difference, adjust the network parameters of the objective model until the objective model converges, obtaining the text-driven model.
[0117] In some feasible examples, combining Figure 4 ,like Figure 7 As shown, the above S13 can be specifically implemented through the following S130-S132.
[0118] S130. Process the reply information to determine the phoneme corresponding to the reply information and the phoneme duration.
[0119] S131. Encode the phonemes and determine the phoneme feature vector corresponding to each phoneme.
[0120] In some examples, server 400 pre-stores the correspondence between each phoneme and its feature vector. Later, when it is necessary to determine the corresponding feature vector for a phoneme, this correspondence can be queried to determine the feature vector for each phoneme.
[0121] Alternatively, when encoding phonemes, server 400 can input the phoneme into a phoneme model to determine the corresponding phoneme feature vector. The training process of the phoneme model is as follows:
[0122] Obtain the training sample data and the labeling results of the training sample data. The training sample data includes at least one phoneme, and the labeling results include the phoneme feature vectors corresponding to the phonemes.
[0123] The training sample data is input into the neural network model to obtain the prediction result of the neural network model for each phoneme in the training sample data.
[0124] If the predicted results and the labeled results are different, adjust the network parameters of the neural network model until the predicted results and the labeled results are the same. Then, the neural network model is considered to have converged, and the phoneme model is obtained.
[0125] S132. Input the phoneme feature vector and phoneme duration into the text-driven model to determine the target key point set.
[0126] In the following embodiments, the display device 200 described above is used as the execution subject for the data processing method provided in the embodiments of this disclosure to illustrate the method of the embodiments of this application.
[0127] This application provides a data processing method, such as... Figure 8 As shown, the data processing method may include S21-S25.
[0128] S21. In response to the selection operation of the target function, send voice information to the server.
[0129] S22. Receive the reply information sent by the server.
[0130] S23. Synthesize the response information to determine the corresponding voice information.
[0131] S24. Input the speech information into the speech-driven model and determine the specified set of key points. The specified set of key points shall include at least the key points used to indicate the mouth and eyes.
[0132] S25. Generate a rendered image of the virtual digital human based on a pre-configured set of preset key points corresponding to the face of the virtual digital human, response information, and a specified set of key points. The preset set of key points includes at least key points indicating the mouth and eyes.
[0133] In some examples, the display device 200 executes the data processing method provided in this disclosure embodiment by calling the edge-side central processing unit (CPU) or a specific graphics processing unit (GPU). For example, the display device 200 receives a segment of voice information input by the user. For example, if the length of the voice information is T, the sampling rate is 16k, and the data type is float32, then this segment of audio can be represented by 16000×T float numbers.
[0134] MFCC: Frame length 1024, frame shift 256, frequency range 90-7600, bin size 80, window selected "Hanning window", finally obtained 2D feature array of 16000×T / 256×80.
[0135] DeepSpeech (optional): The pre-training data is X = {(x(1), y(1)), (x(2), y(2)), ... (x(n), y(n))}, where n = 16000 × T / 256. x(n) represents the 80-dimensional Mel-spectral feature vector for position n, and y(n) represents the phoneme probability distribution P(ct|x) corresponding to position n, where ct ∈ {"pad", " / ", "AA0", ... "x", "z", "zh"}, with a total of 67 English phonemes, 120 Chinese phonemes, and 4 special symbols (space, silence, padding, and prosodic boundaries), for a total of 191. The label is generated in one-hot format, so y(n) is a one-dimensional vector with dimension 191. DeepSpeech uses a Recurrent Neural Network (RNN) structure, with an input of a 16000×T / 256×80 array and an output of a 16000×T / 256×191 feature vector.
[0136] Data Alignment: Training a speech-driven model requires 1. Mel-spectral feature vectors (16000×T / 256×80 dimensions) or DeepSpeech output vectors (16000×T / 256×191 dimensions) and 2. Keypoint coordinates (25×T×318), where 25 represents the frame rate and 318 represents the xyz coordinates of 106 keypoints. Before entering the model, 16000×T / 256 and 25×T need to be aligned in the time dimension. The data processing method provided in this embodiment uses bilinear interpolation to map 25×T×318 keypoints to 16000×T / 256×318 keypoints to achieve alignment. Finally, we obtain a speech feature vector of 16000×T / 256×80 (or 191) and 16000×T / 256×318 keypoint coordinates.
[0137] Key point spatial alignment: The data processing method provided in this embodiment uses Deep-SDM to extract 106 key points of the face, and then aligns the key points in the spatial dimension. The alignment method is to first scale the key points to a certain value range in space according to a certain ratio, and then select a standard face (the midline of the eyes and nose is at x=0). All key points are rotated and translated to achieve maximum overlap with N (e.g., 8) sets of key points on the corners of the eyes and the bridge of the nose on the standard face, and finally obtain the coordinates of the face key points after spatial alignment.
[0138] Frame Alignment: Changes in keypoint coordinates in each frame depend on contextual information, requiring temporal encoding of the feature vectors representing the audio using Mel spectrum or DeepSpeech. The data processing method provided in this disclosure employs a dynamic sliding window to dynamically extract audio features corresponding to the current keypoint coordinates. The window size is 18, and the step size is 1. For example, the audio features corresponding to the i-th keypoint coordinate are from position i-17 to position i. The final input data consists of audio features of dimensions [16000×T / 256, 18, 80] or [16000×T / 256, 18, 191] and keypoint labels of dimensions [16000×T / 256, 318]. T represents the total duration of the audio data received by the display device, and 256 represents the sliding window between the audio data and MFCC.
[0139] Keypoint-driven approach: The keypoint-driven approach uses a Multilayer Perceptron (MLP) + Long Short-Term Memory (LSTM) + MLP model. Training data consists of 6 hours of news broadcast video data, from which training audio data and labels are extracted. Input data includes audio features of dimensions [16000×T / 256, 18, 80] or [16000×T / 256, 18, 191] and keypoint labels of dimension [16000×T / 256, 318]. Output is the Delta [16000×T / 256, 318] keypoint offset relative to the reference. Loss is set to Delta + reference, calculated using the MSE (Maximum Sequence Equation).
[0140] Here, "reference" refers to the preset set of key points corresponding to the face of the virtual digital human, and "Delta" refers to the difference between each initial feature point of the face image of the specified person in a static state and each key point in the specified key point set, also known as the key point offset. For example, when extracting key points from a face image, Deep-SDM is used to extract 106 key points, resulting in 106 key points in the preset key point set and 106 initial feature points for the face image of the specified person in a static state. Then, Delta is obtained by calculating the difference between each key point in the preset key point set and the corresponding initial feature point in the face image of the specified person in a static state. Mouth-weighted loss (optional): Because closed-mouth lip shapes like "b, p, m, f" are difficult to fit in speech-driven keypoint tasks, a larger loss weight is needed for similar lip shapes. The generated Delta + reference is used to obtain the generated facial keypoints F, with dimensions [16000×T / 256, 318]. F is reshaped to [16000×T / 256, 106, 3]. In the second dimension, 101 represents the upper keypoint of the mouth, and 104 represents the lower keypoint. The distance between two keypoints is calculated to obtain the mouth distance matrix D [16000×T / 256, 1]. Since these values di are all numbers greater than 0 and less than approximately 0.2, the smaller the di, the greater the weight. Therefore, the final weight W = 1 / (D×lamda+0.01), where lambda is the coefficient, set to 1, and the addition of 0.01 ensures that division by zero is avoided. Finally, the mouth-weighted loss = MSE×W.
[0141] A speech-driven model is obtained through training. When a speech segment is input into the speech-driven model, a set of specified key points for the corresponding speech segment can be obtained.
[0142] In some feasible examples, combining Figure 8 ,like Figure 9 As shown, the data processing method provided in this embodiment of the disclosure further includes: S26-S28.
[0143] S26. Obtain training sample videos and training supervision videos. The training sample videos include the actual speech data corresponding to the designated person speaking in the training sample videos, and the training supervision videos include the actual key points of the facial images corresponding to the actual speech data of the designated person during their speech.
[0144] In some examples, since the target model cannot directly recognize the training sample video, it is necessary to process the training sample video, such as extracting audio data from the training sample video. For example, the encoder part of a speech recognition pre-trained model such as Deepspeech / Wenet can be used as a speech signal feature extractor to extract the feature vectors of the corresponding audio data (such as MFCC), or the feature vectors corresponding to commonly used spectral information such as audio feature fbank.
[0145] The encoder part of a pre-trained speech recognition model, such as Deepspeech / Wenet, is used as a speech signal feature extractor to extract the speech feature vector of the corresponding speech signal.
[0146] S27. Input the initial feature points and training speech data corresponding to the facial image of the specified person in a static state into the preset model to obtain the predicted key points predicted by the preset model.
[0147] In some examples, the preset model uses an architecture of MLP+Bilstm+MLP. When inputting speech information into the preset model, the speech information needs to be processed to extract audio features (such as MFCC or fbank spectral features) to obtain the feature vectors corresponding to the audio features. Then, this feature vector is used as input to the preset model for supervised learning of the coordinates of facial landmarks detected in each frame of video, and the predicted facial landmarks are output.
[0148] S28. Based on the predicted key points and the actual key points, adjust the network parameters of the preset model until the preset model converges to obtain the speech-driven model.
[0149] In some feasible examples, each keypoint corresponds to an actual coordinate; combined with Figure 8 ,like Figure 10 As shown, the above S25 can be specifically implemented through the following S250-S253.
[0150] S250. Determine at least one vertex coordinate based on the actual coordinates of the key points in the specified key point set. Each vertex coordinate corresponds to one key point.
[0151] In some examples, the electronic device can also calculate the difference (also known as the keypoint offset) between each keypoint in a specified set of keypoints and the initial feature point corresponding to the face image of the specified person in a static state. Then, the sum of each keypoint in the preset set of keypoints corresponding to the face of the pre-configured virtual digital human and the keypoint offset corresponding to that keypoint is used to obtain the corrected keypoint, and the corrected keypoint is used as the vertex coordinate.
[0152] S251. Use each keypoint in the preset keypoint set as a texture coordinate to determine at least three texture coordinates. Each texture coordinate corresponds to one keypoint.
[0153] S252. The texture coordinates are processed using the triangulation method to determine the texture triangle.
[0154] S253. Generate a rendered image of the virtual digital human based on vertex coordinates, texture coordinates, texture triangles, and response information.
[0155] In some examples, vertex coordinates, texture coordinates, and texture triangles can be input into OpenGL ES and rendered at a certain frame rate. At the same time, the input speed of vertex coordinates and the input speed of response information can be controlled to ensure that the rendered image and voice information are aligned, thus ensuring the display effect of the virtual digital human.
[0156] In some examples, a set of key points corresponding to the speech data is obtained through a speech-driven model or a text-driven model (e.g., key points to indicate the face, eyebrows, or nose). Since the face corresponding to the key point set predicted by the speech-driven or text-driven model differs from the standard face of the virtual digital human, it is necessary to use the key point set predicted by the speech-driven or text-driven model to instruct the virtual digital human to perform corresponding operations on the standard face, such as:
[0157] 1. Image key points: Deep-SDM is used to detect the static image corresponding to the virtual digital human to obtain the key points Fls[106,3] dimension of the static image corresponding to the virtual digital human.
[0158] 2. Delta bias: The set of predicted key points is obtained through a speech-driven or text-driven model. Then, the difference between each key point in the specified key point set and the initial feature point corresponding to the specified person's face image in a static state (also known as key point offset) is calculated to obtain the Delta [N, 106, 3] dimension.
[0159] 3. Delta Space Alignment: First, Delta needs to be mapped to the same dimensional space as Fls. Here, two parameters are set: scale and shift. Scale is a floating-point number, and shift is in the form of [x_, y_, z_]. The mapped Delta_ = Delta / scale – shift. The two parameters are manually adjusted multiple times based on the static images corresponding to different virtual digital humans.
[0160] 4. Driving Keypoint Acquisition: Fls represents the keypoints in the preset keypoint set, and Delta represents the keypoint offset sequence obtained by the driving model. By calculating the sum of each keypoint in the preset keypoint set corresponding to the face of the pre-configured virtual digital human and the keypoint offset corresponding to that keypoint, the corrected keypoint Fls_gen = Fls + Delta is obtained.
[0161] 5. Keypoint Selection: The more keypoints there are, the longer the rendering time on the display device 200 will be, and the more resources the display device 200 will consume during rendering. Therefore, keypoints that have a greater impact on the static image corresponding to the virtual digital human are generally selected, such as keypoints for indicating the mouth and keypoints for indicating the eyes, while other keypoints such as keypoints for indicating the nose and keypoints for indicating the cheeks are ignored. In the data processing method provided in this embodiment, 12 keypoints for indicating the eyes and 8 keypoints for indicating the mouth are selected. At the same time, 4 anchor points are added around the face of the static image corresponding to the virtual digital human, and 4 anchor points are added around the static image corresponding to the virtual digital human, for a total of 28 keypoints. The final keypoint coordinates are Fls_gen_select[N, 28, 3].
[0162] 6. Texture Triangles: Based on the key points selected above, obtain Fls_select[28, 3]. Use Delaunay to obtain the triangulation results of the key points of the driving image to obtain the triangulation sequence Tra[49, 3], representing 49 triangles and their corresponding key point numbers.
[0163] 7. Texture coordinates: The texture coordinate value range is 0-1, so the Fls_select[28,3] value range needs to be scaled to the 0-1 space. The scaling method is to divide by the width w of the stretched image (square image).
[0164] 8. Vertex coordinates: The vertex coordinate range is from -1 to 1, and it needs to be calculated as (Fls_gen_select[N, 28, 3] – 0.5××w) / 0.5w.
[0165] 9. Finally, the texture coordinates, texture triangles, vertex coordinates, and image are passed to the OpenGL ES renderer, and the audio-visual alignment is controlled by the speed of the passed vertex coordinates.
[0166] For example, when the video frame rate is 80Fps, the video input speed is 12.5ms / frame.
[0167] In some examples, when the frame rate of the rendered image of the virtual digital human generated by the electronic device is 80 Fps, the transmission of the rendered image is under considerable pressure. To reduce the pressure on the transmission of the rendered image, the frame rate of the rendered image can be reduced, for example, by selecting one frame from the rendered image every four frames, thereby reducing the frame rate of the rendered image of the virtual digital human to 20 Fps.
[0168] 10. (Optional) If the static image corresponding to the virtual digital human is a video with head movements, then the driving key point sequence needs to be aligned with the video key point sequence using a spatial alignment algorithm.
[0169] In some examples, OpenGL ES is set to render at a certain frame rate. The playback speed control module controls the input speed of vertex coordinates and audio information. Vertex coordinates are stored in a vertex buffer through the control module, and audio signals are stored in an audio buffer to await calls from AudioTrack. Audio and video playback are implemented asynchronously, with the audio signals and vertex coordinates aligned only by the playback speed control module, ultimately achieving synchronized audio and video playback.
[0170] Specifically, the computational power requirements of the data processing method provided in this embodiment of the present disclosure on the display device 200 can be controlled by adjusting the number of key points in vertex coordinates and texture coordinates. For example, 68 key points of a face are obtained through Deep-SDM. To reduce the computational power requirements of the data processing method provided in this embodiment of the present disclosure on the display device 200, typically only 12 key points for indicating the left and right eyes, 8 key points for indicating the mouth, and 16 key points for indicating the cheeks are retained from these 68 key points. Other key points can be replaced by anchor points.
[0171] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0172] This application embodiment can divide the data processing device into functional modules according to the above method example. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing unit. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0173] like Figure 11 As shown in the diagram, an embodiment of this application provides a schematic diagram of the structure of a server 400. It includes a communicator 101 and a processor 102.
[0174] The communicator 101 is used to receive voice information sent by an electronic device that triggers human-computer interaction; the processor 102 is used to recognize the voice information received by the communicator 101 and determine the response information of the voice information; the processor 102 is also used to input the response information into a text-driven model to determine a target key point set; wherein, the target key point set includes at least key points for indicating the mouth and eyes; the processor 102 is also used to control the communicator 101 to send target information carrying the response information and the target key point set to the electronic device, so that the electronic device generates a rendered image of the virtual digital human based on a preset key point set corresponding to the face of the virtual digital human, the response information and the target key point set, wherein the preset key point set includes at least key points for indicating the mouth and eyes.
[0175] In some feasible examples, the communicator 101 is further configured to acquire training sample videos and the labeling results of the training sample videos; wherein, the training sample videos include actual text data corresponding to a specified person speaking in the training sample videos, and facial images corresponding to the actual text data of the specified person speaking, and the labeling results include actual key points of the facial images; the processor 102 is further configured to input the initial feature points corresponding to the facial images of the specified person in a static state and the training text data acquired by the communicator 101 into the target model to determine the predicted key points predicted by the target model; the processor 102 is further configured to adjust the network parameters of the target model based on the predicted key points and the actual key points acquired by the communicator 101 until the target model converges to obtain a text-driven model.
[0176] In some implementable examples, processor 102 is specifically used to determine the area difference corresponding to the key point used to indicate the mouth based on the first target loss function and the actual key points obtained by the predictor 101; processor 102 is specifically used to match the second target loss function, the predicted key points, and the actual key points obtained by the predictor 101 to determine the distance between two matched key points; processor 102 is specifically used to match the third target loss function, the predicted key points, and the actual key points obtained by the predictor 101 to determine the gradient difference between two matched key points; processor 102 is specifically used to adjust the network parameters of the target model according to the area difference, distance, and gradient difference until the target model converges to obtain a text-driven model.
[0177] In some feasible examples, processor 102 is specifically used to process the response information to determine the phoneme corresponding to the response information and the phoneme duration; processor 102 is specifically used to encode the phoneme to determine the phoneme feature vector corresponding to each phoneme; processor 102 is specifically used to input the phoneme feature vector and phoneme duration into the text-driven model to determine the target key point set.
[0178] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and their functions will not be repeated here.
[0179] Of course, the server 400 provided in this application embodiment includes, but is not limited to, the modules described above. For example, the server 400 may also include a memory 103. The memory 103 may be used to store the program code of the server 400, and may also be used to store data generated by the server 400 during operation, such as data in write requests.
[0180] As an example, combined Figure 3 The result unit 410 and the sending unit 412 in the server 400 perform the same functions as the communicator 101, the processing unit 411 performs the same functions as the processor 102, and the storage unit 413 performs the same functions as the memory 103.
[0181] like Figure 12As shown, this application embodiment also provides a chip system that can be applied to the server 400 in the foregoing embodiments. The chip system includes at least one processor 1501 and at least one interface circuit 1502. The processor 1501 can be the processor in the server 400. The processor 1501 and the interface circuit 1502 can be interconnected via a line. The processor 1501 can receive and execute computer instructions from the memory of the server 400 through the interface circuit 1502. When the computer instructions are executed by the processor 1501, the server 400 can perform the various steps performed by the server 400 in the foregoing embodiments. Of course, the chip system may also include other discrete devices, which are not specifically limited in this application embodiment.
[0182] This application also provides a computer-readable storage medium for storing computer instructions executed by the server 400 described above.
[0183] This application also provides a computer program product, including computer instructions executed by the server 400 described above.
[0184] like Figure 13 As shown in the diagram, an embodiment of this application provides a schematic structural diagram of a display device 200. It includes a communicator 201 and a processor 202.
[0185] Processor 202 is configured to control communicator 201 to send voice information to server in response to a selection operation of a target function; communicator 201 is configured to receive response information sent by server; processor 202 is configured to synthesize the response information received by communicator 201 to determine the voice information corresponding to the response information; processor 202 is further configured to input the voice information into a voice-driven model to determine a specified set of key points; wherein the specified set of key points includes at least key points for indicating the mouth and eyes; processor 202 is further configured to generate a rendered image of the virtual digital human based on a pre-configured set of preset key points corresponding to the face of the virtual digital human, the response information, and the specified set of key points; wherein the preset set of key points includes at least key points for indicating the mouth and eyes.
[0186] In some feasible examples, the communicator 201 is further configured to acquire training sample videos and training supervision videos; wherein, the training sample videos include actual speech data corresponding to a specified person speaking in the training sample videos, and the training supervision videos include actual key points of the facial images corresponding to the actual speech data of the specified person during speaking; the processor 202 is further configured to input the initial feature points corresponding to the facial images of the specified person in a static state and the training speech data acquired by the communicator 201 into a preset model to obtain the predicted key points predicted by the preset model; the processor 202 is further configured to adjust the network parameters of the preset model based on the predicted key points and the actual key points acquired by the communicator 201 until the preset model converges to obtain a speech-driven model.
[0187] In some feasible examples, each keypoint corresponds to an actual coordinate; processor 202 is specifically used to determine at least one vertex coordinate based on the actual coordinates of the keypoints in a specified set of keypoints; wherein, one vertex coordinate corresponds to one keypoint; processor 202 is also used to use each keypoint in a preset set of keypoints as texture coordinates to determine at least three texture coordinates; wherein, one texture coordinate corresponds to one keypoint; processor 202 is also used to process the texture coordinates using triangulation to determine texture triangles; processor 202 is also used to generate a rendered image of the virtual digital human based on the vertex coordinates, texture coordinates, texture triangles, and response information.
[0188] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and their functions will not be repeated here.
[0189] Of course, the display device 200 provided in this application embodiment includes, but is not limited to, the modules described above. For example, the display device 200 may also include a memory 203. The memory 203 may be used to store the program code of the display device 200, and may also be used to store data generated by the display device 200 during operation, such as data in write requests.
[0190] As an example, combined Figure 3 The receiving unit 210 and the transmitting unit 211 in the display device 200 perform the same functions as the communicator 101, the processing unit 212 performs the same functions as the processor 102, and the storage unit 213 performs the same functions as the memory 103.
[0191] like Figure 14As shown, this application embodiment also provides a chip system that can be applied to the display device 200 in the foregoing embodiments. The chip system includes at least one processor 2501 and at least one interface circuit 2502. The processor 2501 may be the processor in the display device 200. The processor 2501 and the interface circuit 2502 are interconnected via a line. The processor 2501 can receive and execute computer instructions from the memory of the display device 200 through the interface circuit 2502. When the computer instructions are executed by the processor 2501, the display device 200 can perform the various steps performed by the display device 200 in the foregoing embodiments. Of course, the chip system may also include other discrete devices, which are not specifically limited in this application embodiment.
[0192] This application also provides a computer-readable storage medium for storing computer instructions executed by the display device 200.
[0193] This application also provides a computer program product, including computer instructions executed by the above-described display device 200.
[0194] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data processing method, characterized by, The method comprises: receiving voice information sent by an electronic device for triggering human-computer interaction; identifying the voice information to determine reply information of the voice information; inputting the reply information into a text-driven model to determine a target key point set; wherein the target key point set at least includes key points for indicating a mouth and eyes; sending target information carrying the reply information and the target key point set to the electronic device, so that the electronic device generates a rendering image of a virtual digital person according to a preset key point set corresponding to a face of the virtual digital person, the reply information and the target key point set, the preset key point set at least including key points for indicating a mouth and eyes; wherein the training process of the text-driven model comprises: obtaining a training sample video and a labeled result of the training sample video; wherein the training sample video includes actual text data corresponding to a specified person speaking in the training sample video, and a person image corresponding to the actual text data in the speaking process of the specified person, and the labeled result includes actual key points of the person image; inputting initial feature points corresponding to a person image of the specified person in a static state and training text data into a target model to determine predicted key points predicted by the target model; wherein the model architecture adopted by the target model is a FastSpeech / FastSpeech2 model structure based on a Transformer structure; adjusting network parameters of the target model based on the predicted key points and the actual key points until the target model converges to obtain the text-driven model; the adjusting of the network parameters of the target model based on the predicted key points and the actual key points until the target model converges to obtain the text-driven model comprises: determining an area difference value corresponding to a key point for indicating a mouth based on a first target loss function, the predicted key points and the actual key points; matching a second target loss function, the predicted key points and the actual key points to determine a distance between two matching key points; matching a third target loss function, the predicted key points and the actual key points to determine a gradient difference value between two matching key points; adjusting the network parameters of the target model according to the area difference value, the distance and the gradient difference value until the target model converges to obtain the text-driven model.
2. The data processing method according to claim 1, characterized in that, The inputting of the reply information into the text-driven model to determine the target key point set comprises: processing the reply information to determine phonemes corresponding to the reply information and phoneme durations of the phonemes; encoding the phonemes to determine phoneme feature vectors corresponding to each phoneme; inputting the phoneme feature vectors and the phoneme durations into the text-driven model to determine the target key point set.
3. A data processing apparatus, characterized by, The method comprises: a receiving unit configured to receive voice information sent by an electronic device for triggering human-computer interaction; The processing unit is configured to recognize the voice information received by the receiving unit, and determine reply information of the voice information. The processing unit is further configured to input the reply information into a text-driven model to determine a target key point set, wherein the target key point set at least includes key points for indicating a mouth and eyes. The processing unit is further configured to control a sending unit to send target information carrying the reply information and the target key point set to the electronic device, so that the electronic device generates a rendering image of a virtual digital person according to a preset key point set corresponding to a face of the virtual digital person, the reply information, and the target key point set, and the preset key point set at least includes key points for indicating a mouth and eyes. The training process of the text-driven model includes: obtaining a training sample video and a labeled result of the training sample video, wherein the training sample video includes actual text data corresponding to speech of a specified person in the training sample video, and a face image corresponding to the actual text data in the speech process of the specified person, and the labeled result includes actual key points of the face image; inputting initial feature points corresponding to a face image of the specified person in a static state and training text data into a target model to determine predicted key points predicted by the target model, wherein the target model adopts a FastSpeech / FastSpeech2 model structure based on a Transformer structure; adjusting network parameters of the target model based on the predicted key points and the actual key points until the target model converges to obtain the text-driven model, wherein adjusting the network parameters of the target model based on the predicted key points and the actual key points until the target model converges to obtain the text-driven model includes: determining an area difference value corresponding to a key point for indicating a mouth based on a first target loss function, the predicted key points, and the actual key points; matching a second target loss function, the predicted key points, and the actual key points to determine a distance between two matching key points; matching a third target loss function, the predicted key points, and the actual key points to determine a gradient difference value between two matching key points; adjusting the network parameters of the target model based on the area difference value, the distance, and the gradient difference value until the target model converges to obtain the text-driven model.
4. An electronic device, comprising: including: a memory and a processor, the memory being configured to store a computer program, and the processor being configured to, when executing the computer program, cause the electronic device to implement the data processing method of claim 1 or 2.
5. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and when the computer program is executed by a computing device, the computing device implements the data processing method of claim 1 or 2.
Citation Information
Patent Citations
Interaction system, method and device, electronic equipment and storage medium
CN112669846A
Action driving method and device of target object, equipment and storage medium
CN113554737A
Face key point information acquisition method and method and device for generating face animation
CN114429658A
Virtual anchor face swapping method and apparatus, electronic device, and storage medium
WO2021232878A1