Virtual human control method and device, server, and storage medium
By acquiring the shape and feature information of the target object, and using a virtual human generation engine and a distributed stream processing engine to generate the target virtual human, the problem of insufficient realism and vividness of virtual humans is solved, and the interactive experience and display real-time performance are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2026-04-07
AI Technical Summary
Existing virtual human technology suffers from poor realism and vividness, resulting in a poor interactive experience, and the appearance algorithm has insufficient ability to analyze and process feature data.
By acquiring the shape and feature information of the target object, processing it using a virtual human generation engine, generating a target virtual human corresponding to the target object, and then displaying the virtual human using a distributed stream processing engine, the image and vividness of the virtual human are enhanced.
The generated virtual humans are more lifelike and vivid, enhancing the interactive experience with virtual humans and improving the real-time performance and realism of the virtual human display.
Smart Images

Figure CN114662640B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of virtual human technology, and more specifically, to a control method, device, server, and storage medium for a virtual human. Background Technology
[0002] With the development of science and technology, the use of virtual human services is becoming increasingly widespread, with more and more functions, and certain application scenarios have emerged. Currently, the virtual humans used are still in a relatively rudimentary stage, consisting of intelligent cartoon characters with poor realism and liveliness, resulting in a poor interactive experience. Summary of the Invention
[0003] In view of the above problems, this application proposes a virtual human control method, device, server, and storage medium to solve the above problems.
[0004] In a first aspect, embodiments of this application provide a method for controlling a virtual human, the method comprising: acquiring shape information and feature information of a target object; and processing the shape information and feature information based on a virtual human generation engine to obtain a target virtual human corresponding to the target object.
[0005] Secondly, embodiments of this application provide a control device for a virtual human, the device comprising: an information acquisition module for acquiring shape information and feature information of a target object; and a virtual human acquisition module for processing the shape information and feature information based on a virtual human generation engine to obtain a target virtual human corresponding to the target object.
[0006] Thirdly, embodiments of this application provide a server, including a memory and a processor, wherein the memory is coupled to the processor, the memory stores instructions, and when the instructions are executed by the processor, the processor performs the above-described method.
[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing program code, which can be invoked by a processor to execute the above-described method.
[0008] The virtual human control method, device, server, and storage medium provided in this application embodiment acquire the appearance information and feature information of the target object, process the appearance information and feature information based on the virtual human generation engine, and obtain the target virtual human corresponding to the target object. Thus, the virtual human is generated through multi-dimensional information including appearance information and feature information, which can make the generated virtual human more vivid and lifelike, and improve the interactive experience with the virtual human. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A schematic diagram illustrating the application environment of the control method for the virtual human provided in the embodiments of this application is shown.
[0011] Figure 2 This illustration shows a schematic diagram of the interaction between the server and the client provided in an embodiment of this application;
[0012] Figure 3 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown;
[0013] Figure 4 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown;
[0014] Figure 5 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown;
[0015] Figure 6 This application shows Figure 5 A flowchart illustrating step S330 of the virtual human control method shown;
[0016] Figure 7 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown;
[0017] Figure 8 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown;
[0018] Figure 9 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown;
[0019] Figure 10 This application provides a schematic diagram of the functional data flow in an embodiment.
[0020] Figure 11 A block diagram of a virtual human control device according to an embodiment of this application is shown;
[0021] Figure 12 A block diagram of an electronic device for performing a virtual human control method according to an embodiment of this application is shown;
[0022] Figure 13An embodiment of this application shows a storage unit for storing or carrying program code that implements the control method for a virtual human according to an embodiment of this application. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0024] Currently, virtual human services have some applications. However, the virtual humans currently in use are still in a relatively rudimentary stage, resembling intelligent cartoon characters. Most solutions are still in the early stages of exploration, and the productization is still quite basic, with poor realism and vividness, resulting in a poor interactive experience. Furthermore, the ability of virtual human appearance algorithms to analyze and process feature data needs improvement, and the generated virtual human appearance is not distinct enough, leading to unnatural interactions and a poor user experience.
[0025] To address the aforementioned problems, the inventors, through long-term research, discovered and proposed the virtual human control method, device, server, and storage medium provided in the embodiments of this application. By generating virtual humans using multi-dimensional information including appearance and feature information, the generated virtual humans can be more lifelike and vivid, enhancing the interactive experience. The specific control method for the virtual human will be described in detail in subsequent embodiments.
[0026] The application environment for the control method of the virtual human provided in the embodiments of this application will be described below.
[0027] Please see Figure 1 , Figure 1 A schematic diagram illustrating an application environment for the control method of the virtual human provided in the embodiments of this application is shown. For example... Figure 1 As shown, the application environment includes a server 100 and a client 200. The server 100 and client 200 communicate to achieve data interaction between them. For example, the client 200 can receive the shape information and feature information of the target object, and then transmit the shape information and feature information of the target object to the server 100. The server 100 can generate a virtual target person corresponding to the target object based on the shape information and feature information of the target object, and send the virtual target person to the client 200 for display on the client 200.
[0028] In some implementations, the server 200 may be equipped with a virtual human generation engine and a distributed stream processing engine. The virtual human generation engine can be used to generate virtual humans, and the distributed stream processing engine can be used to display the virtual humans.
[0029] Please see Figure 2 , Figure 2 This illustration shows a schematic diagram of the interaction between the server and the client provided in an embodiment of this application. For example... Figure 2 As shown, firstly, the client can collect the personalized information needed to customize the virtual avatar, including a customized avatar photo, age, gender, hobbies, occupation, etc. Through different personalized information, it is possible to customize a virtual avatar to the desired personalized form. Among them, image upload mainly customizes the visible features of the personalized image. By uploading photos of relatives or friends, a familiar virtual avatar can be customized, and then interaction can be achieved with the virtual avatar. For example, by uploading a photo of the mother, a scene of the mother telling a story to the baby can be realized. Virtual avatar personalized information mainly involves uploading personalized non-avatar feature information, such as age, gender, preferences, and other multi-dimensional feature information, for more in-depth customization.
[0030] Secondly, the server-side customization service aggregates personalized photos and multi-dimensional information, and then transmits them to the deep learning model engine to generate a virtual human image with features.
[0031] Thirdly, the content input module mainly processes the input of text information, such as story content and broadcast speech scripts. At the same time, the content input module also inputs voice content with features, providing voice semantic functions.
[0032] Fourth, the distributed stream processing engine provides real-time virtual human video streams, which are then loaded and displayed on the client in real time.
[0033] Fifth, the control panel (Dashboard) is mainly responsible for service monitoring and early warning, including the collection of data from data points, providing key log information for system debugging and subsequent optimization, and statistical information on service calls.
[0034] Sixth, the virtual human avatar player is mainly responsible for registering Kafka topics and receiving data input from the distributed stream processing engine, and then displaying it to the user on the client side.
[0035] Please see Figure 3 , Figure 3 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown. This method generates a virtual human using multi-dimensional information including shape information and feature information, thereby making the generated virtual human more lifelike and vivid, and enhancing the interactive experience with the virtual human. In a specific embodiment, this virtual human control method is applied to, for example... Figure 8 The virtual human control device 300 and the server 100 configured with the virtual human control device 300 are shown. Figure 9 The following will use server 100 as an example to illustrate the specific process of this embodiment. Of course, it is understood that the server used in this embodiment can be a traditional server, a cloud server, etc., and is not limited here. The following will focus on Figure 3 The process shown will be described in detail. The control method for the virtual human may specifically include the following steps:
[0036] Step S110: Obtain the shape information and feature information of the target object.
[0037] In this embodiment, the shape information and feature information of the target object can be obtained. The target object may include animals, such as cats, dogs, pigs, pandas, lions, tigers, etc., and may also include people, such as the elderly, young people, children, men, women, etc., without limitation.
[0038] In some implementations, once a target object is identified, its shape information can be obtained. This shape information can be used to characterize the object's outward appearance. As one possible approach, a user can upload the target object's shape information via a client, and the server can retrieve this information from the client based on the connection.
[0039] In one approach, a user can upload a photo containing the appearance information of a target object via a client. The server then retrieves the photo, containing the appearance information of the target object, from the client based on the connection with the client. The client may include an image upload module, through which the user can upload the photo, and the server retrieves the photo, containing the appearance information of the target object, from the client based on the connection with the client.
[0040] As another method, users can upload descriptive text containing the appearance information of the target object through a client. The server then retrieves this descriptive text from the client based on the connection with the client. The client may include a personal information upload module, through which users can upload descriptive text containing the appearance information of the target object, and the server retrieves this descriptive text from the client based on the connection with the client.
[0041] In some implementations, once a target object is identified, its characteristic information can be obtained. This characteristic information can be used to characterize the target object's attributes. As one possible approach, a user can upload the target object's characteristic information via a client, and the server can retrieve this information from the client based on the connection with the client.
[0042] In one approach, users can upload descriptive text containing characteristic information of a target object via a client. The server then retrieves this descriptive text from the client based on the connection to the client. The client may include a personal information upload module, through which users can upload descriptive text containing characteristic information of the target object, and the server retrieves this descriptive text from the client based on the connection to the client.
[0043] In some implementations, the characteristic information of the target object may include one or a combination of several of the following: gender information, age information, height information, weight information, preference information, and occupation information. For example, the characteristic information of the target object may include gender information; the characteristic information of the target object may include gender information and age information; the characteristic information of the target object may include gender information, age information, preference information, occupation information, etc., without limitation.
[0044] For example, if the target object is a user's mother, then by obtaining the mother's appearance information, a virtual persona that matches the mother's appearance can be customized to interact with the user. For example, a virtual persona that matches the mother's appearance can tell stories to the user, thereby improving the user's interactive experience. By obtaining the mother's characteristic information, more in-depth customization can be carried out. For example, the virtual persona's age, height, weight, and occupation can be made more compatible with the mother, thereby improving the virtual persona's realism and vividness, and thus enhancing the user's interactive experience.
[0045] In some implementations, during the process of acquiring the appearance information and feature information of the target object, the user's identifier can be obtained to bind the user identifier to the appearance information and feature information of the target object. For example, assuming a user uploads the appearance information and feature information of the target object through a client, the user's identifier can be obtained. As one approach, when a user logs into the client, they can be prompted to log in using their user identifier to facilitate obtaining that identifier. As yet another approach, when a user uploads the appearance information and feature information of the target object through their client, they can be prompted to log in using their user identifier to facilitate obtaining that identifier. This user identifier can include, for example, a user's ID card number, phone number, or other identifier used to uniquely identify the user.
[0046] Step S120: Process the shape information and feature information based on the virtual human generation engine to obtain a target virtual human corresponding to the target object.
[0047] In this embodiment, after obtaining the shape information and feature information of the target object, the shape information and feature information of the target object can be processed to obtain a target virtual person corresponding to the target object. It is understood that the target virtual person corresponding to the target object should have the same or similar shape and features as the target object.
[0048] In some implementations, given the shape information and feature information of the target object, a virtual human generation engine can be used to process these information to obtain a target virtual human corresponding to the target object. As one possible implementation, the virtual human generation engine may include a deep learning model engine. Given the shape information and feature information of the target object, the deep learning model engine can then be used to process these information to obtain a target virtual human corresponding to the target object.
[0049] In some implementations, after obtaining the shape information and feature information of the target object, these information can be input into a trained virtual human generation model to obtain a target virtual human corresponding to the target object, output by the trained model. This trained virtual human generation model is obtained through machine learning. Specifically, a training dataset is first collected, where one type of data in the training dataset has attributes or features that distinguish it from another type. Then, the collected training dataset is used to train a neural network according to a preset algorithm, thereby summarizing patterns based on the training dataset to obtain the trained virtual human generation model. In this embodiment, the training dataset may, for example, consist of multiple shape information and feature information corresponding to multiple virtual humans. The preset algorithm may include motion algorithms, voice algorithms, facial expression algorithms, etc.
[0050] In some implementations, when processing the appearance information and feature information of a target object to obtain the target virtual person corresponding to the target object, the appearance information, feature information, and identification information of the target object can be processed together to obtain the target virtual person corresponding to the target object carrying the identification information. Based on this, the target virtual person corresponding to the user can be found and applied using the identification information, thereby improving the speed of finding and applying the target virtual person.
[0051] One embodiment of this application provides a virtual human control method that obtains the shape information and feature information of a target object, processes the shape information and feature information based on a virtual human generation engine, and obtains a target virtual human corresponding to the target object. Thus, the virtual human is generated through multi-dimensional information including shape information and feature information, which makes the generated virtual human more vivid and lifelike, and improves the interactive experience with the virtual human.
[0052] Please see Figure 4 , Figure 4 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown. The following will focus on... Figure 4 The process shown will be described in detail. The control method for the virtual human may specifically include the following steps:
[0053] Step S210: Obtain the shape information and feature information of the target object.
[0054] For a detailed description of step S210, please refer to step S110, which will not be repeated here.
[0055] Step S220: Aggregate the shape information and the feature information to obtain tag information.
[0056] In this embodiment, after obtaining the shape information and feature information of the target object, the shape information and feature information of the target object can be aggregated, and the aggregated information can be determined as the tag information of the target object. This realizes the aggregation of the originally independent shape information and feature information into a whole data, which can improve the convenience and accuracy of subsequent data processing and avoid omissions in data processing.
[0057] In some implementations, aggregating the shape information and feature information of the target object may include: packaging the shape information and feature information of the target object into a data packet and identifying the data packet as tag information; associating the shape information and feature information of the target object with the user identifier of the target object and identifying the associated shape information and feature information as tag information; updating the shape information of the target object based on the feature information of the target object to obtain the tag information of the target object.
[0058] The target object's appearance information can include a photograph of the target object, while its characteristic information can include one or more of the following: gender, age, height, weight, preferences, and occupation. Therefore, the target object's photograph can be modified or updated based on this characteristic information. For example, the photograph can be modified based on the target object's weight and height, or based on their occupation, and the modified photograph can then be used as the target object's tag information.
[0059] Step S230: Process the tag information based on the virtual human generation engine to obtain the target virtual human corresponding to the target object.
[0060] In this embodiment, if the tag information of the target object is obtained, the tag information of the target object can be processed to obtain the target virtual person corresponding to the target object.
[0061] In some implementations, if the tag information of the target object is obtained, the tag information of the target object can be processed based on the virtual human generation engine to obtain the target virtual human corresponding to the target object.
[0062] One embodiment of this application provides a virtual human control method that acquires the shape information and feature information of a target object, aggregates the shape information and feature information to obtain tag information, processes the tag information based on a virtual human generation engine, and obtains the target virtual human corresponding to the target object. Compared to Figure 3 The virtual human control method shown in this embodiment also aggregates multi-dimensional information including shape information and feature information, and generates virtual humans based on the aggregated information, so as to improve the generation speed and accuracy of virtual humans.
[0063] Please see Figure 5 , Figure 5 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown. The following will focus on... Figure 5 The process shown will be described in detail. The control method for the virtual human may specifically include the following steps:
[0064] Step S310: Obtain the shape information and feature information of the target object.
[0065] Step S320: Process the shape information and feature information based on the virtual human generation engine to obtain a target virtual human corresponding to the target object.
[0066] For a detailed description of steps S310-S320, please refer to steps S110-S120, which will not be repeated here.
[0067] Step S330: Process the target virtual human based on the distributed stream processing engine, and send the target virtual human to the client for display.
[0068] In this embodiment, once the server obtains the target virtual person corresponding to the target object, it can send the target virtual person to the communicating client for display, thereby enabling interaction between the user and the target virtual person. In this embodiment, the client is only used for displaying the target virtual person and interacting with the user, and does not need to be used for deep learning image calculations. Therefore, the client has lighter resource requirements and can be customized for various devices, such as smartphones, smart TVs, smart appliances, or smart home devices with displays; no limitation is made here.
[0069] In some implementations, the server can be equipped with a distributed stream processing engine, such as Kafka. Upon obtaining the target virtual human, the server can process the target virtual human using the distributed stream processing engine before sending it to the client for display. Alternatively, the client, upon obtaining the target virtual human, can render it and display it. The distributed stream processing engine can simultaneously support backend services for multiple different virtual humans and provides high-performance real-time virtual human backend video stream output capabilities. This allows for real-time loading and display of the virtual human on the client, improving the real-time performance of the virtual human display.
[0070] As one feasible approach, once the target virtual human is obtained, a channel Topic can be registered in the distributed stream processing engine. Subsequently, the client can obtain the data stream through the corresponding user identifier and channel Topic, such as obtaining the target virtual human or the video stream of the target virtual human, etc., without any limitations.
[0071] Please see Figure 6 , Figure 6 This application shows Figure 5 The flowchart illustrates step S330 of the virtual human control method shown. The following will focus on... Figure 6 The process shown will be described in detail, and the method may specifically include the following steps:
[0072] Step S331: Receive the input content information.
[0073] In this embodiment, input content information can be received. This content information may include user-inputted voice information or preset text information.
[0074] In some implementations, the client can run a virtual human application, through which the user can input content information on the client. The client then sends the input content information to the server's distributed stream processing engine, which in turn can receive the input content information.
[0075] In one approach, the content information can include interactive voice information. The client can run a virtual human application, allowing the user to input voice information via the virtual human application. The client then sends the input voice information to the server's distributed stream processing engine, which in turn receives the input voice information.
[0076] As another approach, the content information can include pre-defined text information. The client can run a virtual human application, allowing the user to select pre-defined text information on the client side. The client then sends the selected text information to the server's distributed stream processing engine, which in turn receives the selected text information.
[0077] In some implementations, the server may include a content input module, which is primarily used to handle the input of text information, such as story content or speech drafts. Upon receiving the input content, the server can input it into the content input module, which then transmits the content to the stream queue of the distributed stream processing engine for processing.
[0078] Step S332: Aggregate the target virtual person and the content information to obtain the video stream of the target virtual person.
[0079] In this embodiment, upon obtaining the input content information, the target virtual person and the content information can be aggregated to obtain a video stream of the target virtual person. It is understood that the video stream of the target virtual person includes at least the target virtual person and the content information. In some implementations, the server-side distributed stream processing engine, upon obtaining the content information, can aggregate the target virtual person generated by the virtual person generation engine to obtain a video stream of the target virtual person.
[0080] In some implementations, aggregating the target virtual person and content information to obtain a video stream of the target virtual person may include: video fusion of the target virtual person and content information to obtain a video stream of the target virtual person; updating the voice and posture of the target virtual person according to the time sequence of the content information to obtain a video stream of the target virtual person, etc., which are not limited here.
[0081] Step S333: Process the video stream of the target virtual human based on the distributed stream processing engine, and send the video stream of the target virtual human to the client for display.
[0082] In this scenario, once the server obtains the video stream of the target virtual person corresponding to the target object, it can send the video stream of the target virtual person to the client that is communicating with it for display, thereby enabling the display of the video stream of the target virtual person to the user.
[0083] In some implementations, the server can be equipped with a distributed stream processing engine. Upon obtaining the video stream of the target virtual human, the server can process the video stream using the distributed stream processing engine and then send it to the client for display. Alternatively, upon obtaining the video stream of the target virtual human, the client can render the video stream and display it.
[0084] An embodiment of this application provides a virtual human control method that acquires the shape information and feature information of a target object, processes the shape information and feature information based on a virtual human generation engine to obtain a target virtual human corresponding to the target object, processes the target virtual human based on a distributed stream processing engine, and sends the target virtual human to a client for display. Compared to Figure 3 The virtual human control method shown in this embodiment also uses a distributed streaming engine to process the display of the virtual human on the client side, so as to improve the real-time performance of the virtual human display.
[0085] Please see Figure 7 , Figure 7 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown. In this embodiment, the content information includes voice information, which will be discussed below. Figure 7 The process shown will be described in detail. The control method for the virtual human may specifically include the following steps:
[0086] Step S410: Obtain the shape information and feature information of the target object.
[0087] Step S420: Process the shape information and feature information based on the virtual human generation engine to obtain a target virtual human corresponding to the target object.
[0088] For a detailed description of steps S410-S420, please refer to steps S110-S120, which will not be repeated here.
[0089] Step S430: Receive the input content information.
[0090] For a detailed description of step S430, please refer to step S331, which will not be repeated here.
[0091] Step S440: Parse the speech information to obtain the first text information and timbre information corresponding to the speech information.
[0092] In this embodiment, the content information includes interactive voice information. This voice information can be preset, such as preset recordings or audio from a preset video, or it can be input from an external source, such as input by a user through a client or input from another device through a client; no limitation is made here.
[0093] In some implementations, upon receiving input voice information, the voice information can be parsed to obtain corresponding text information as first text information, and corresponding timbre information. As one approach, upon receiving input voice information, speech-to-text processing can be performed to obtain corresponding text information, which is then used as the first text information. As yet another approach, upon receiving input voice information, timbre recognition can be performed to obtain corresponding timbre information. Specifically, upon receiving input voice information, the voice information can be input into a trained timbre recognition model to obtain the timbre information corresponding to the voice information output by the trained timbre recognition model.
[0094] In some implementations, upon receiving input voice information, the voice information can be parsed to obtain corresponding text information as first text information, and corresponding voiceprint information can be obtained. As one approach, upon receiving input voice information, speech-to-text processing can be performed to obtain corresponding text information, which is then used as the first text information. As yet another approach, upon receiving input voice information, voiceprint recognition can be performed to obtain corresponding voiceprint information. Specifically, upon receiving input voice information, the voice information can be input into a trained voiceprint recognition model to obtain the voiceprint information corresponding to the voice information output by the trained voiceprint recognition model.
[0095] Step S450: Aggregate the target virtual person, the first text information, and the timbre information to obtain a video stream of the target virtual person, wherein the video stream of the target virtual person includes the target virtual person broadcasting the first text information using the timbre information.
[0096] In this embodiment, upon obtaining the first text information and timbre information corresponding to the voice information, the target virtual person, the first text information, and the timbre information can be aggregated to obtain a video stream of the target virtual person. This video stream includes the target virtual person broadcasting the first text information using the timbre information. Therefore, in the generated video stream of the target virtual person, the target virtual person can broadcast the first text information with the desired timbre, thereby creating an immersive audio environment for the user.
[0097] In some implementations, upon acquiring the first text information and timbre information, the server-side distributed stream processing engine can aggregate the target virtual human generated by the virtual human generation engine to obtain a video stream of the target virtual human. Aggregating the target virtual human, the first text information, and the timbre information to obtain the video stream of the target virtual human can include: video fusion of the target virtual human, the first text information, and the timbre information to obtain the video stream of the target virtual human; updating the voice, posture, and timbre of the target virtual human according to the chronological order of the first text information to obtain the video stream of the target virtual human, etc., without limitation.
[0098] In this embodiment, upon obtaining the first text information and voiceprint information corresponding to the voice information, the target virtual person, the first text information, and the voiceprint information can be aggregated to obtain a video stream of the target virtual person. The video stream of the target virtual person includes the target virtual person broadcasting the first text information using the voiceprint information. Based on this, in the generated video stream of the target virtual person, the target virtual person can broadcast the first text information using the desired voiceprint, thereby creating an immersive audio environment for the user.
[0099] In some implementations, upon acquiring the first text information and voiceprint information, the server-side distributed stream processing engine can aggregate the target virtual human generated by the virtual human generation engine to obtain a video stream of the target virtual human. Aggregating the target virtual human, the first text information, and the voiceprint information to obtain the video stream of the target virtual human can include: video fusion of the target virtual human, the first text information, and the voiceprint information to obtain the video stream of the target virtual human; updating the voice, posture, and voiceprint of the target virtual human according to the chronological order of the first text information to obtain the video stream of the target virtual human, etc., without limitation.
[0100] Step S460: Process the video stream of the target virtual human based on the distributed stream processing engine, and send the video stream of the target virtual human to the client for display.
[0101] For a detailed description of step S460, please refer to step S333, which will not be repeated here.
[0102] An embodiment of this application provides a virtual human control method, which acquires the shape information and feature information of a target object, processes the shape information and feature information based on a virtual human generation engine to obtain a target virtual human corresponding to the target object, receives input voice information, parses the voice information to obtain first text information and timbre information corresponding to the voice information, aggregates the target virtual human, the first text information, and the timbre information to obtain a video stream of the target virtual human, wherein the video stream of the target virtual human includes the target virtual human broadcasting the first text information using the timbre information, processes the video stream of the target virtual human based on a distributed stream processing engine, and sends the video stream of the target virtual human to a client for display. Compared to Figure 3 The virtual human control method shown in this embodiment further utilizes a distributed streaming engine to process the virtual human's display on the client side, thereby improving the real-time performance of the virtual human display. Additionally, this embodiment acquires audio information to generate a video stream for the virtual human, enhancing its realism and improving the user's interactive experience.
[0103] Please see Figure 8 , Figure 8A flowchart illustrating a virtual human control method according to an embodiment of this application is shown. In this embodiment, the content information includes voice information, which will be discussed below. Figure 8 The process shown will be described in detail. The control method for the virtual human may specifically include the following steps:
[0104] Step S510: Obtain the shape information and feature information of the target object.
[0105] Step S520: Process the shape information and feature information based on the virtual human generation engine to obtain the target virtual human corresponding to the target object.
[0106] For a detailed description of steps S510-S520, please refer to steps S110-S120, which will not be repeated here.
[0107] Step S530: Receive the input content information.
[0108] For a detailed description of step S530, please refer to step S331, which will not be repeated here.
[0109] Step S540: Parse the content information to obtain the second text information and semantic information corresponding to the content information.
[0110] In some implementations, upon receiving input content information, the content information can be parsed to obtain the corresponding text information as second text information, and the semantic information corresponding to the speech information can be obtained. As one approach, upon receiving input speech information, speech-to-text processing can be performed to obtain the corresponding text information, which is then used as the second text information. As yet another approach, upon receiving input speech information, semantic recognition can be performed to obtain the corresponding semantic information. Specifically, upon receiving input semantic information, the speech information can be input into a trained semantic recognition model to obtain the semantic information corresponding to the speech information output by the trained semantic recognition model.
[0111] Step S550: Determine facial expression information based on the semantic information.
[0112] In this embodiment, if the semantic information corresponding to the voice information is obtained, the facial expression information can be determined based on the semantic information.
[0113] In some implementations, a preset mapping relationship can be pre-set and stored. This preset mapping relationship includes multiple semantic information, multiple facial expression information, and the correspondence between the multiple semantic information and the multiple facial expression information. One semantic information can correspond to one facial expression. Therefore, in this embodiment, when the semantic information corresponding to the voice information is obtained, the semantic information corresponding to the voice information can be determined from the multiple semantic information based on the preset mapping relationship, and the facial expression information corresponding to the semantic information corresponding to the voice information can be determined from the multiple facial expression information.
[0114] For example, suppose multiple semantic information includes first semantic information, second semantic information, and third semantic information, and multiple facial expression information includes first facial expression information, second facial expression information, and third facial expression information, and the first semantic information corresponds to the first facial expression information, the second semantic information corresponds to the second facial expression information, and the third semantic information corresponds to the third facial expression information. Then, if the semantic information corresponding to the speech information is the first semantic information, then the facial expression information can be determined to be the first facial expression information; if the semantic information corresponding to the speech information is the third semantic information, then the facial expression information can be determined to be the third facial expression information.
[0115] In some implementations, if the semantic information corresponding to the voice information is to express happiness, then the facial expression information can be a smile; if the semantic information corresponding to the voice information is to express sadness, then the facial expression information can be a frown, etc., without limitation.
[0116] Step S560: Aggregate the target virtual person, the second text information, and the facial expression information to obtain a video stream of the target virtual person, wherein the video stream of the target virtual person includes the target virtual person broadcasting the second text information using the facial expression information.
[0117] In this embodiment, upon obtaining the second text information and facial expression information corresponding to the voice information, the target virtual person, the second text information, and the facial expression information can be aggregated to obtain a video stream of the target virtual person. This video stream includes the target virtual person using the facial expression information to read the second text information. Therefore, in the generated video stream of the target virtual person, the target virtual person can read the second text information using facial expressions that correspond to the semantics, thereby creating an immersive experience for the user.
[0118] In some implementations, upon acquiring the second text information and facial expression information, the server-side distributed stream processing engine can aggregate the data with the target virtual human generated by the virtual human generation engine to obtain a video stream of the target virtual human. Aggregating the target virtual human, the second text information, and the facial expression information to obtain the video stream of the target virtual human can include: video fusion of the target virtual human, the second text information, and the facial expression information to obtain a video stream of the target virtual human; updating the voice, posture, and facial expressions of the target virtual human according to the chronological order of the second text information to obtain a video stream of the target virtual human, etc., without limitation.
[0119] Step S570: Process the video stream of the target virtual human based on the distributed stream processing engine, and send the video stream of the target virtual human to the client for display.
[0120] For a detailed description of step S570, please refer to step S333, which will not be repeated here.
[0121] An embodiment of this application provides a virtual human control method, which acquires the shape information and feature information of a target object, processes the shape information and feature information based on a virtual human generation engine to obtain a target virtual human corresponding to the target object, receives input voice information, parses the voice information to obtain second text information and semantic information corresponding to the voice information, determines facial expression information based on the semantic information, aggregates the target virtual human, the second text information, and the facial expression information to obtain a video stream of the target virtual human, wherein the video stream of the target virtual human includes the target virtual human broadcasting the second text information with the facial expression information, processes the video stream of the target virtual human based on a distributed stream processing engine, and sends the video stream of the target virtual human to a client for display. Compared to Figure 3 The virtual human control method shown in this embodiment further utilizes a distributed streaming engine to process the virtual human's display on the client side, thereby improving the real-time performance of the virtual human display. Additionally, this embodiment uses facial expression information determined based on semantic information to generate a video stream for the virtual human, enhancing the realism of the virtual human and improving the user's interactive experience.
[0122] Please see Figure 9 , Figure 9 A flowchart illustrating a virtual human control method according to an embodiment of this application is shown. In this embodiment, the content information includes voice information, which will be discussed below. Figure 9 The process shown will be described in detail. The control method for the virtual human may specifically include the following steps:
[0123] Step S610: Obtain the shape information and feature information of the target object.
[0124] For a detailed description of step S610, please refer to step S110, which will not be repeated here.
[0125] Step S620: Obtain the scene information of the target object.
[0126] In this embodiment, the scene information of the target object can be obtained.
[0127] One approach is to acquire a photograph of the target object, identify the background within the photograph, and obtain scene information about the target object. Another approach is to acquire input text information containing scene information about the target object, identify the text information, and obtain the scene information about the target object.
[0128] In some implementations, the scene information of the target object may include an indoor scene or an outdoor scene. Indoor scenes may include living room scenes, bedroom scenes, kitchen scenes, etc., and are not limited thereto. Outdoor scenes may include street scenes, seaside scenes, grassland scenes, forest scenes, etc., and are not limited thereto.
[0129] Step S630: Based on the virtual human generation engine, process the appearance information, the feature information, and the scene information to obtain the target virtual human corresponding to the target object.
[0130] In this embodiment, after obtaining the appearance information, feature information, and scene information of the target object, the appearance information, feature information, and scene information of the target object can be processed to obtain a target virtual person corresponding to the target object. It is understood that the target virtual person is adapted to the scene information; for example, the target virtual person's clothing, background, and expression are adapted to the scene information, etc., but this is not limited here.
[0131] In some implementations, if the appearance information, feature information, and scene information of the target object are obtained, the appearance information, feature information, and scene information of the target object can be processed based on the virtual human generation engine to obtain a target virtual human corresponding to the target object.
[0132] An embodiment of this application provides a virtual human control method that acquires the shape information and feature information of a target object, acquires the scene information in which the target object is located, and processes the shape information, feature information, and scene information based on a virtual human generation engine to obtain a target virtual human corresponding to the target object. Compared to Figure 3The virtual human control method shown in this embodiment also acquires the target scene to generate a video stream of the virtual human, so as to improve the realism of the virtual human and enhance the user's interactive experience.
[0133] Please see Figure 10 , Figure 10 A flowchart illustrating the functional data provided in an embodiment of this application is shown. For example... Figure 10 As shown, firstly, the virtual human client inputs an image, and customized information such as age and gender is uploaded to the backend customization service. The customization service aggregates the virtual human's information using the user identifier and generates the virtual human's document information. Secondly, the generated aggregated virtual human information with specific identifiers is transmitted to the virtual human video generation engine (model generation service). The model generation engine generates a virtual human image with specific features. Thirdly, the virtual human image channel registers a channel Topic with the distributed stream processing engine. Subsequently, the client obtains the data stream through the corresponding virtual human identifier and the corresponding stream engine Topic. Fourthly, the user inputs interactive voice information or pre-set text information through the virtual human application to the stream engine. The stream engine connects to the model generated by the model generation service, aggregates the data, and then returns a real-time virtual human video stream to the client. Fifthly, the virtual human client receives the virtual human video stream and plays it, completing the real-time virtual human effect display.
[0134] Please see Figure 11 , Figure 11 A block diagram of a virtual human control device according to an embodiment of this application is shown. The following will focus on... Figure 11 The block diagram shown illustrates that the virtual human control device 300 includes: an information acquisition module 310 and a virtual human acquisition module 320, wherein:
[0135] The information acquisition module 310 is used to acquire the shape information and feature information of the target object.
[0136] The virtual human acquisition module 320 is used to process the shape information and the feature information based on the virtual human generation engine to obtain a target virtual human corresponding to the target object.
[0137] Furthermore, the virtual human acquisition module 320 includes: a tag information acquisition submodule and a first virtual human acquisition submodule, wherein:
[0138] The tag information acquisition submodule is used to aggregate the shape information and the feature information to obtain tag information.
[0139] The first virtual human acquisition submodule is used to process the tag information based on the virtual human generation engine to obtain the target virtual human corresponding to the target object.
[0140] Furthermore, the virtual human acquisition module 320 includes: a scene information acquisition submodule and a second virtual human acquisition submodule, wherein:
[0141] The scene information acquisition submodule is used to acquire the scene information of the target object.
[0142] The second virtual human acquisition submodule is used to process the appearance information, the feature information and the scene information based on the virtual human generation engine to obtain the target virtual human corresponding to the target object.
[0143] Furthermore, the virtual human control device 300 further includes: a virtual human processing module, wherein:
[0144] The virtual human processing module is used to process the target virtual human based on a distributed stream processing engine and send the target virtual human to the client for display.
[0145] Furthermore, the virtual human processing module includes: a content information receiving submodule, a video stream acquisition submodule, and a virtual human processing submodule, wherein:
[0146] The content information receiving submodule is used to receive input content information.
[0147] The video stream acquisition submodule is used to aggregate the target virtual person and the content information to obtain the video stream of the target virtual person.
[0148] Furthermore, the content information includes voice information, and the video stream acquisition submodule includes: a first parsing unit and a first aggregation unit, wherein:
[0149] The first parsing unit is used to parse the speech information to obtain the first text information and timbre information corresponding to the speech information.
[0150] The first aggregation unit is used to aggregate the target virtual person, the first text information, and the timbre information to obtain a video stream of the target virtual person, wherein the video stream of the target virtual person includes the target virtual person broadcasting the first text information using the timbre information.
[0151] Furthermore, the video stream acquisition submodule includes: a second parsing unit, an expression information determination unit, and a second aggregation unit, wherein:
[0152] The second parsing unit is used to parse the content information to obtain the second text information and semantic information corresponding to the content information.
[0153] The facial expression information determination unit is used to determine facial expression information based on the semantic information.
[0154] The second aggregation unit is used to aggregate the target virtual person, the second text information, and the facial expression information to obtain a video stream of the target virtual person, wherein the video stream of the target virtual person includes the target virtual person broadcasting the second text information using the facial expression information.
[0155] The virtual human processing submodule is used to process the video stream of the target virtual human based on the distributed stream processing engine, and send the video stream of the target virtual human to the client for display.
[0156] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0157] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0158] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0159] Please see Figure 12 This document illustrates a structural block diagram of a server 100 provided in an embodiment of this application. The server 100 in this application may include one or more of the following components: a processor 110, a memory 120, and one or more application programs. The one or more application programs may be stored in the memory 120 and configured to be executed by one or more processors 110. The one or more application programs are configured to perform the methods described in the foregoing method embodiments.
[0160] The processor 110 may include one or more processing cores. The processor 110 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120. Optionally, the processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 110 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 110 and may be implemented separately using a communication chip.
[0161] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the server 100 during use (such as phonebook data, audio and video data, chat log data, etc.).
[0162] Please see Figure 13 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 400 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0163] The computer-readable storage medium 400 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program code 410 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 410 may be compressed, for example, in a suitable form.
[0164] In summary, the virtual human control method, device, server, and storage medium provided in this application obtain the appearance information and feature information of the target object, process the appearance information and feature information based on the virtual human generation engine, and obtain the target virtual human corresponding to the target object. Thus, the virtual human is generated through multi-dimensional information including appearance information and feature information, which makes the generated virtual human more vivid and lifelike, and improves the interactive experience with the virtual human.
[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for controlling a virtual human, characterized in that, The method includes: The object's appearance information and its feature information are obtained, wherein the object's appearance information includes a photograph of the object, and the object's feature information includes the object's weight, height, and occupation. The photo is corrected based on the weight information, height information, and occupation information, and the corrected photo is identified as the tag information of the target object; The tag information is processed using a virtual human generation engine to obtain a target virtual human corresponding to the target object; Receive input content information, wherein the content information includes input voice information and preset text information; According to the time sequence of the content information, the voice and posture of the target virtual person are updated to obtain the video stream of the target virtual person; The video stream of the target virtual human is processed using a distributed stream processing engine, and then sent to the client for display.
2. The method according to claim 1, characterized in that, The content information includes voice information, and the aggregation of the target virtual person and the content information to obtain the video stream of the target virtual person includes: The speech information is parsed to obtain the first text information and timbre information corresponding to the speech information; The target virtual person, the first text information, and the timbre information are aggregated to obtain a video stream of the target virtual person, wherein the video stream of the target virtual person includes the target virtual person broadcasting the first text information using the timbre information.
3. The method according to claim 1, characterized in that, The step of aggregating the target virtual person and the content information to obtain the video stream of the target virtual person includes: The content information is parsed to obtain the second text information and semantic information corresponding to the content information; Based on the semantic information, determine the facial expression information; The target virtual person, the second text information, and the facial expression information are aggregated to obtain a video stream of the target virtual person, wherein the video stream of the target virtual person includes the target virtual person broadcasting the second text information using the facial expression information.
4. The method according to any one of claims 1-3, characterized in that, The feature information includes one or a combination of gender information, age information, height information, weight information, preference information, and occupation information.
5. The method according to any one of claims 1-3, characterized in that, The virtual human generation engine processes the shape information and feature information to obtain the target virtual human corresponding to the target object, including: Obtain the scene information of the target object; The virtual human generation engine processes the appearance information, feature information, and scene information to obtain the target virtual human corresponding to the target object.
6. A control device for a virtual human, characterized in that, The device includes: The information acquisition module is used to acquire the appearance information and feature information of the target object, wherein the appearance information of the target object includes a photograph of the target object, and the feature information of the target object includes the weight information, height information, and occupation information of the target object; The virtual human acquisition module is used to correct the photo based on the weight information, the height information, and the occupation information, and determine the corrected photo as the tag information of the target object; and to process the tag information based on the virtual human generation engine to obtain the target virtual human corresponding to the target object. The virtual human processing module is used to receive input content information, including input voice information and preset text information; update the voice and posture of the target virtual human according to the time sequence of the content information to obtain the video stream of the target virtual human; process the video stream of the target virtual human based on a distributed stream processing engine, and send the video stream of the target virtual human to the client for display.
7. A server, characterized in that, The method includes a memory and a processor, the memory being coupled to the processor, the memory storing instructions, and the processor performing the method as described in any one of claims 1-5 when the instructions are executed by the processor.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Block-chain-based virtual image interaction method, terminal and readable storage medium
CN109445579A
Information processing method and device and electronic equipment
CN111583415A
Interaction method and device, equipment and computer readable medium
CN112364144A
Virtual character generation method and device
CN114242037A
Virtual person interaction system, video generation method, and video generation program
JP2021086618A