Digital human interaction method, device, processor and storage medium
By pre-storing multiple digital humans in a resource pool, determining allocation priorities based on user request information, dynamically managing the number of digital humans, and synthesizing and displaying digital human videos, the problem of low efficiency in digital human resource management is solved, achieving efficient resource utilization and improved user experience.
Patent Information
- Application Number
- CN202211131832.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2042-09-16
AI Technical Summary
In existing technologies, digital human resource management is inefficient, users have long waiting times, resources cannot be fully utilized, and backend systems are put under pressure during peak periods.
By pre-storing multiple digital humans in the resource pool, determining the allocation priority based on user request information, dynamically managing the number of digital humans, synthesizing and displaying digital human videos, and optimizing resource allocation.
It improved the utilization rate of digital human resources, shortened the waiting time for users to access the system for the first time, reduced the allocation pressure on the backend system during peak hours, and improved the user experience.
Smart Images

Figure CN115550454B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to a digital human interaction method, a digital human interaction device, a processor and a storage medium. BACKGROUND
[0002] With the development of China's economy and society, the professional level of the service industry is continuously improved, and the division of labor of various posts is increasingly refined. It is essential to use digital humans to help customers answer questions. The use of digital humans saves time in different scenarios and provides professional consulting services for customers in terms of consulting problems.
[0003] In the prior art, if a user wants to access a digital human, a digital human room dedicated to serving the user needs to be started in the background. After the user finishes interacting with the digital human, the digital human room resources are recycled. This method takes a long time to start the digital human room in the background, and the user needs to wait for a long time to start interacting with the digital human, resulting in poor user experience. In addition, digital human resources need to be frequently applied, allocated and recycled, digital human resources cannot be fully utilized, and the demand for resources during user peak periods cannot be met, which puts a lot of pressure on the backend system. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a digital human interaction method, device, processor and storage medium.
[0005] To achieve the above-mentioned purpose, the first aspect of the present application provides a digital human interaction method, which comprises: in the case that a plurality of users send pull stream requests in the same time period, obtaining request information of each user in the plurality of users; determining the allocation priority of the corresponding user according to the request information; determining a to-be-served user from the plurality of users sending the pull stream request according to the allocation priority of the user and the number of digital humans pre-stored in the resource pool; allocating digital humans to the to-be-served user in turn according to the allocation priority of the to-be-served user; and in response to a query request sent by the to-be-served user, synthesizing a digital human video and displaying the digital human video to the corresponding to-be-served user.
[0006] In the embodiments of the present application, the method further comprises: determining the number of digital humans pre-stored in the resource pool and the number of pull stream requests received in the same time period; if the number of pull stream requests is more than the number of digital humans and the difference between them is greater than a set expansion difference value, increasing the number of digital humans in the resource pool; if the number of digital humans is more than the number of pull stream requests and the difference between them is greater than a set reduction difference value, reducing the number of digital humans in the resource pool.
[0007] In the embodiments of the present application, the request information includes: request scenario, urgency and request time.
[0008] In the embodiments of the present application, the request scenario includes a business question and answer scenario, a live broadcast scenario, and a chat scenario.
[0009] In the embodiments of the present application, the response to the query request sent by the to-be-served user includes synthesizing a digital human video according to the query request and the request information of the to-be-served user, and displaying the digital human video to the corresponding to-be-served user.
[0010] In the embodiments of the present application, the synthesizing of the digital human video according to the query request and the request information of the to-be-served user includes identifying the query request of the to-be-served user, and extracting user feature information of the to-be-served user from the query request; determining an image feature of a digital human according to the user feature information of the to-be-served user and the request information of the to-be-served user; generating reply information according to the identification result of the query request; and synthesizing the digital human video according to the image feature and the corresponding reply information.
[0011] In the embodiments of the present application, the user feature information includes tone information, timbre information, age information, and gender information.
[0012] In the embodiments of the present application, the image feature includes an external shape feature, a motion feature, and a sound feature.
[0013] In the embodiments of the present application, the generating of the reply information according to the identification result of the query request includes training a question and answer model according to a question and answer database; and inputting the identification result into the question and answer model to obtain the reply information.
[0014] In the embodiments of the present application, after obtaining the reply information, the method further includes obtaining a matching degree of the identification result of the query request and the corresponding reply information; adding the matching degree of the identification result and the corresponding reply information to the question and answer database, and retraining the question and answer model.
[0015] The second aspect of the present application provides a digital human interaction device, which includes: a digital human access module, configured to, in the case that a plurality of users send pull stream requests in the same time period, acquire request information of each user in the plurality of users; determine an allocation priority of a corresponding user according to the request information; determine a to-be-served user from the plurality of users sending the pull stream requests according to the allocation priority of the user and a number of digital humans pre-stored in a resource pool; and allocate digital humans to the to-be-served user in turn according to the allocation priority of the to-be-served user; a digital human driving module, configured to synthesize a digital human video in response to a query request sent by the to-be-served user; and the digital human access module is further configured to display the digital human video to the corresponding to-be-served user.
[0016] The third aspect of the present application provides a processor configured to execute the digital human interaction method described above.
[0017] The fourth aspect of the present application provides a machine-readable storage medium having instructions stored thereon, which, when executed by a processor, cause the processor to be configured to execute the digital human interaction method described above.
[0018] The fifth aspect of the present application provides a computer program product comprising a computer program which, when executed by a processor, implements the digital human interaction method described above.
[0019] Through the technical solutions provided by the present application, the present application has at least the following technical effects:
[0020] The digital human interaction method of the present application pre-stores a plurality of digital humans in a resource pool. In the case of receiving a plurality of pull stream requests sent by users in the same period, the request information of each user in the plurality of users is obtained, the allocation priority of the corresponding user is determined according to the request information, the user to be served is determined from the plurality of users sending the pull stream request according to the allocation priority of the user and the number of digital humans pre-stored in the resource pool, then the digital human is allocated to the user to be served in turn according to the allocation priority of the user to be served, the digital human video is synthesized corresponding to the query request sent by the user to be served, and the digital human video is displayed to the corresponding user to be served. The method provided by the present application manages a plurality of digital human push streams through a resource pool, improves the resource utilization rate of the digital human, shortens the waiting time when the user first accesses, and reduces the allocation pressure of the backend system during the user peak period.
[0021] Other features and advantages of the embodiments of the present application will be described in detail in the following specific implementation part. BRIEF DESCRIPTION OF DRAWINGS
[0022] The accompanying drawings are included to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used together with the following specific implementation to explain the embodiments of the present application, but do not constitute a limitation on the embodiments of the present application. In the drawings:
[0023] Figure 1 A flowchart of a digital human interaction method according to an embodiment of the present application is schematically shown;
[0024] Figure 2 A schematic diagram of a digital human interaction device according to an embodiment of the present application is schematically shown;
[0025] Figure 3 An internal structure diagram of a computer device according to an embodiment of the present application is schematically shown.
[0026] Reference Signs List
[0027] 200 - digital human interaction device; 201 - digital human access module; 202 - digital human driving module; 203 - AI capability module; 204 - digital human management module; 205 - auxiliary module; A01 - processor; A02 - network interface; A03 - internal memory; A04 - non-volatile storage medium; B01 - operating system; B02 - computer program. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It should be understood that the specific implementation manners described herein are only used to explain and illustrate the embodiments of the present application and should not be used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0029] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative positional relationship, movement condition, etc. between components in a certain specific posture (as shown in the drawings), and if the specific posture changes, the directional indications also change accordingly.
[0030] In addition, if the embodiments of the present application involve descriptions such as “first”, “second”, etc., the descriptions of “first”, “second”, etc. are only for description purposes and should not be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by “first”, “second” can explicitly or implicitly include at least one of the features. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can realize it, and when the combination of technical solutions contradicts each other or cannot be realized, it should be considered that the combination of technical solutions does not exist and is not within the scope of protection claimed by the present application.
[0031] Figure 1 A flowchart of a digital human interaction method according to an embodiment of the present application is schematically shown. As shown in FIG. 1, the digital human interaction method according to the embodiment of the present application includes the following steps. Figure 1As shown, in an embodiment of the present application, a digital human interaction method is provided, comprising the following steps: step 101: in the case that a plurality of users send pull stream requests in the same time period, obtaining the request information of each user in the plurality of users; step 102: determining the allocation priority of the corresponding user according to the request information; step 103: determining the to-be-served user from the plurality of users sending the pull stream request according to the allocation priority of the user and the number of digital humans pre-stored in the resource pool; step 104: allocating digital humans to the to-be-served user in turn according to the allocation priority of the to-be-served user; and step 105: synthesizing a digital human video in response to the query request sent by the to-be-served user and displaying the digital human video to the corresponding to-be-served user.
[0032] Specifically, in the embodiment of the present application, a plurality of digital humans are pre-stored in the resource pool, and in the case that a plurality of users send pull stream requests in the same time period, the request information of each user sending the pull stream request is obtained. Then the allocation priority of the corresponding user is determined according to the request information, and the number of digital humans stored in the resource pool is obtained. If the number of users sending the pull stream request is less than or equal to the number of digital humans, the users sending the pull stream request are determined as to-be-served users, and the to-be-served users are allocated digital humans in turn according to the allocation priority. If the number of users sending the pull stream request is greater than the number of digital humans, the users are sorted according to the allocation priority, and the users corresponding to the number of digital humans in the sorting are selected as to-be-served users, and the to-be-served users are allocated digital humans in turn according to the allocation priority. The to-be-served user sends a query request after being allocated a digital human, a digital human video is synthesized in response to the query request sent by the to-be-served user, and the digital human video is displayed to the corresponding to-be-served user.
[0033] The method provided by the present application manages a plurality of digital human push streams through a resource pool, improves the resource utilization rate of the digital humans, shortens the waiting time of the user when first accessing, and reduces the allocation pressure of the backend system during the user peak period.
[0034] In the embodiment of the present application, the method further comprises: determining the number of digital humans pre-stored in the resource pool and the number of pull stream requests received in the same time period; if the number of pull stream requests is more than the number of digital humans and the difference between the two is greater than a set expansion difference value, increasing the number of digital humans in the resource pool; and if the number of digital humans is more than the number of pull stream requests and the difference between the two is greater than a set reduction difference value, reducing the number of digital humans in the resource pool.
[0035] Specifically, in the embodiments of the present application, the number of digital humans pre-stored in the statistical resource pool and the number of pull stream request received in the same time period are counted, if the number of pull stream request is more than the number of digital humans and the difference between the two is greater than the set expansion difference value, the number of digital humans in the resource pool can be increased to alleviate the current digital human shortage, quickly allocate digital humans to users, shorten the user waiting time and improve the user experience. If the number of digital humans is more than the number of pull stream request and the difference between the two is greater than the set reduction difference value, the number of digital humans in the resource pool is reduced to reduce the occupation of redundant digital human resources in the backend system and improve the resource utilization rate of digital humans.
[0036] In the embodiments of the present application, the request information includes: request scene, urgency and request time.
[0037] In the embodiments of the present application, the request scene includes: business question and answer scene, live broadcast scene, chat scene.
[0038] Specifically, in the embodiments of the present application, the allocation priority of the user can be determined according to the request scene, urgency and request time, wherein the request scene at least includes: business question and answer scene, live broadcast scene, chat scene. In a possible implementation, the demand for digital humans in the business question and answer scene, live broadcast scene and chat scene is lower and lower, so the corresponding allocation priority is lowered in turn. For users with high urgency, the corresponding allocation priority is high. For users with early request time, the corresponding allocation priority is high.
[0039] In the embodiments of the present application, the digital human video is synthesized in response to the query request sent by the to-be-served user, and the digital human video is displayed to the corresponding to-be-served user, including: synthesizing the digital human video according to the query request and the request information of the to-be-served user; displaying the digital human video to the corresponding to-be-served user.
[0040] Specifically, in the embodiments of the present application, the query request can be a voice request or a text request. After receiving the query request of the user, the preference of the to-be-served user is analyzed and extracted from the query request and the request information, and the synthesized digital human video is made accordingly, and the digital human video is displayed to the corresponding to-be-served user. The digital human video can better match the user preference and further improve the user experience. In the present application, webRTC (Web Real-Time Communication) is used to support the API of real-time voice conversation or video conversation of web browsers. The user terminal device combines webRTC technology with webView to perform real-time audio and video conversation with the digital human.
[0041] In the embodiment of the present application, the synthesizing the digital human video according to the query request and the request information of the to-be-served user comprises: identifying the query request of the to-be-served user, and extracting user feature information of the to-be-served user from the query request; determining an image feature of the digital human according to the user feature information of the to-be-served user and the request information of the to-be-served user; generating reply information according to the identification result of the query request; and synthesizing the digital human video according to the image feature and the corresponding reply information.
[0042] In the embodiment of the present application, the generating the reply information according to the identification result of the query request comprises: training a question and answer model according to a question and answer database; inputting the identification result into the question and answer model to obtain the reply information.
[0043] In the embodiment of the present application, the user feature information comprises: tone information, timbre information, age information, and gender information.
[0044] In the embodiment of the present application, the image feature comprises: an appearance feature, a motion feature, and a sound feature.
[0045] Specifically, in the embodiment of the present application, first, the query request of the to-be-served user is identified, and the user feature information of the to-be-served user is extracted from the query request. For example, in the case that the query request is voice information, the tone information, the timbre information, and the gender information of the user can be extracted from the query request, and the age information and the regional information of the user are analyzed and determined. In the case that the query request is text information, the age information, the regional information, the gender information, and the current tone information of the user can be analyzed and determined from the mood adverbs and the conjunctions. Then, the image feature of the digital human is determined according to the user feature information and the request information. The image feature comprises: an appearance feature, a motion feature, and a sound feature. The appearance feature comprises: the hairstyle, the clothing, and the background image of the digital human, the motion feature comprises: the limb motion of the digital human, such as waving hands, waving, bowing, nodding, and bending, and the sound feature comprises: male voice / female voice, sweet / soft, lively / cute / serious, etc. For example, the request scene of the user can be determined according to the request information, if the current request scene is a business question and answer scene, then formal clothes are matched for the digital human, and a more serious expression and tone are selected as the image feature of the digital human. For another example, according to the voice information of the to-be-served user, it is determined that the to-be-served user is a little girl, then the appearance of the digital human is matched with the appearance of a child digital human, and a cute and sweet female voice is selected as the image feature of the digital human.
[0046] Then, the query request of the user to be served is recognized, converted into text information, and the text information is input into an intelligent question and answer model to output corresponding reply information. In the present application, intelligent question and answer models of different industries can be introduced to adapt to application scenarios of different industries, thereby reducing the cost of migrating virtual digital human customer service question and answer applications to other industries. Further, the context of the interaction between the user and the digital human is recorded, thereby facilitating the intelligent question and answer model to more accurately understand the user's question and to give reasonable reply information, so as to improve the question and answer interaction experience. Finally, a digital human video is synthesized according to the image characteristics and the corresponding reply information. The video stream is pushed to a streaming media service, and the streaming media service synchronizes the video stream to the user terminal device.
[0047] In the embodiments of the present application, after obtaining the reply information, the method further includes: obtaining a matching degree of the recognition result of the query request and the corresponding reply information; adding the matching degree of the recognition result and the corresponding reply information to the question and answer database, and retraining the question and answer model.
[0048] Specifically, in the embodiments of the present application, data backflow is introduced to determine whether the recognition result of the query request matches the corresponding reply information and to obtain a matching degree. The matching degree, the recognition result of the query request and the corresponding reply information are added to the question and answer database, and the question and answer model is retrained to improve the accuracy of the intelligent question and answer model in specific business scenarios and to improve the user experience.
[0049] The digital human interaction method of the present application pre-stores a plurality of digital humans in a resource pool. In the case where a plurality of users send pull stream requests in the same period, the request information of each user in the plurality of users is obtained, the allocation priority of the corresponding user is determined according to the request information, the number of digital humans pre-stored in the resource pool is determined according to the allocation priority of the user and the number of digital humans pre-stored in the resource pool, the user to be served is determined from the plurality of users sending the pull stream request, and then the digital human is allocated to the user to be served in turn according to the allocation priority of the user to be served. The digital human video corresponding to the query request sent by the user to be served is synthesized, and the digital human video is displayed to the corresponding user to be served. Through the method provided by the present application, the resource utilization rate of the digital human is improved by managing a plurality of digital humans in the resource pool, the waiting time of the user when first accessing is shortened, and the allocation pressure of the backend system during the user peak period is reduced.
[0050] It should be understood that, although Figure 1 The steps in the flowchart of the present application are displayed in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, Figure 1At least one of the steps in the method can comprise a plurality of sub-steps or a plurality of stages, which sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the order of the sub-steps or stages is not necessarily sequential, but can be performed alternately or in rotation with other steps or sub-steps or stages of other steps.
[0051] In one embodiment, as shown in Figure 2 In one embodiment, as shown in
[0052] Further, the digital human access module 201 comprises a webRTC (Web Real-Time Communication), which is an API supporting real-time voice conversation or video conversation of a web browser, and a user terminal device combines the webRTC technology with a webView to perform real-time audio and video conversation with a digital human; a streaming media service, which is used for collecting, caching, scheduling and transmitting and playing streaming media content (such as audio and video); a resource pool, which is used for managing digital human rooms; and a dynamic routing, which is used for routing a to-be-served user to an idle digital human room resource.
[0053] The digital human driving module 202 comprises a voice synthesis module, which is used for synthesizing output voice by inputting text, and can also control parameters such as tone and volume; a motion synthesis module, which is used for driving facial expressions of a digital human by voice; a video synthesis module, which is used for completing digital human video recording by introducing a video recorder component in a UE (Unreal Engine), and can also add some rich text elements to a digital human video to enrich the expression of the digital human; and a video coding and decoding module, which is used for coding and decoding of a digital human video, so that the digital human video can be output in a streaming manner.
[0054] The digital human interaction device 200 further comprises an AI capability module 203, a digital human management module 204, and an auxiliary module 205. The AI capability module 203 comprises a voice recognition module for recognizing a query of a user to be served and converting the query into corresponding text; an intelligent question and answer module for inputting a recognition result of the query request and outputting corresponding reply information; and a dialogue management module for recording a context of interaction between the user and the digital human, so that the intelligent question and answer model can more accurately understand a question of the user and thus give a reasonable answer text, to improve the question and answer interaction experience.
[0055] The digital human management module 204 comprises a digital human image management module for managing an image of the digital human such as a face, a hairstyle, a dress, a background, a posture, etc.; a digital human voice management module for managing a voice configuration of the digital human such as a male voice / female voice, a sweet voice / soft voice, a lively voice / lovely voice / serious voice; a digital human action management module for managing a body action of the digital human such as waving a hand, waving a hand, bowing, nodding, bending, etc.; a digital human image management module for managing an image of the digital human such as a face, a hairstyle, a dress, a background, a posture, etc.; a digital human voice management module for managing a voice configuration of the digital human such as a male voice / female voice, a sweet voice / soft voice, a lively voice / lovely voice / serious voice; and a digital human action management module for managing a body action of the digital human such as waving a hand, waving a hand, bowing, nodding, bending, etc.
[0056] The auxiliary module 205 comprises a unified gateway for forwarding different types of requests of a user terminal to corresponding services, and the introduction of the gateway can also filter some illegal requests to ensure the security of the backend services; a data backflow module for obtaining a matching degree of a recognition result of the query request and corresponding reply information; and a data storage module comprising a persistent storage mysql, a cos storage minio, and a cache database redis, wherein the mysql mainly stores user information and configuration information related to the digital human, the minio mainly stores digital human assets such as a hairstyle, a dress, a background image, an action, and a voice of the digital human, and the Redis mainly caches user login information and other data.
[0057] In the embodiment of the present application, the resource pool is further used to determine a number of digital humans of the digital humans pre-stored in the resource pool and a number of pull stream request received in the same time period; if the number of the pull stream request is more than the number of the digital humans and a difference between the two is greater than a set expansion difference value, the number of the digital humans in the resource pool is increased; and if the number of the digital humans is more than the number of the pull stream request and a difference between the two is greater than a set reduction difference value, the number of the digital humans in the resource pool is reduced.
[0058] In the embodiment of the present application, the request information comprises a request scene, an emergency degree, and a request time.
[0059] In the embodiment of the present application, the request scene includes a business question and answer scene, a live broadcast scene, and a chat scene.
[0060] In the embodiment of the present application, the digital human driving module 202 is configured to synthesize a digital human video according to the query request and the request information of the to-be-served user; and the digital human access module 201 is configured to display the digital human video to the corresponding to-be-served user.
[0061] In the embodiment of the present application, the AI capability module 203 is configured to identify the query request of the to-be-served user and extract user feature information of the to-be-served user from the query request; the digital human management module 204 is configured to determine the image feature of the digital human according to the user feature information of the to-be-served user and the request information of the to-be-served user; the AI capability module 203 is further configured to generate reply information according to the identification result of the query request; and the digital human driving module 201 is further configured to synthesize the digital human video according to the image feature and the corresponding reply information.
[0062] In the embodiment of the present application, the user feature information includes tone information, timbre information, age information, and gender information.
[0063] In the embodiment of the present application, the image feature includes an external shape feature, a motion feature, and a sound feature.
[0064] In the embodiment of the present application, the AI capability module 203 is further configured to train a question and answer model according to a question and answer database; and the identification result is input into the question and answer model to obtain the reply information.
[0065] In the embodiment of the present application, the data backflow module is configured to, after obtaining the reply information, acquire a matching degree of the identification result of the query request and the corresponding reply information; add the matching degree of the identification result and the corresponding reply information to the question and answer database; and retrain the question and answer model.
[0066] The digital human interaction device includes a processor and a memory, and the digital human access module, the digital human driving module, and the like are stored in the memory as program units, and the corresponding functions are implemented by the processor executing the program modules stored in the memory.
[0067] The processor includes a core, and the core retrieves the corresponding program unit from the memory. The core can be set to one or more, and the digital human interaction is realized by adjusting the core parameters.
[0068] The memory can include a non-permanent memory in a computer readable medium, a random access memory (RAM), and / or a non-volatile memory such as a read-only memory (ROM) or a flash memory (flash RAM), and the memory includes at least one memory chip.
[0069] An embodiment of the present application provides a processor configured to execute the digital human interaction method described above.
[0070] An embodiment of the present application provides a machine readable storage medium, which stores instructions, when executed by a processor, causes the processor to be configured to execute the digital human interaction method described above.
[0071] In one embodiment, a computer device, which can be a server, is provided, and an internal structure diagram of the computer device can be as shown in Figure 3 Figure 3 An internal structure diagram of a computer device according to an embodiment of the present application is schematically shown. The computer device includes a processor A01, a network interface A02, a memory (not shown in the figure) and a database (not shown in the figure) connected through a system bus. The processor A01 of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes an internal memory A03 and a non-volatile storage medium A04. The non-volatile storage medium A04 stores an operating system B01, a computer program B02 and a database (not shown in the figure). The internal memory A03 provides an environment for the operating system B01 and the computer program B02 in the non-volatile storage medium A04 to run. The network interface A02 of the computer device is configured to communicate with an external terminal through a network connection. The computer program B02 is executed by the processor A01 to implement a digital human interaction method.
[0072] Those skilled in the art can understand that Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0073] An embodiment of the present application further provides a computer program product, when executed on a data processing device, is adapted to execute a program that is initialized with the following method steps: in the case that a plurality of users send pull stream requests in the same time period, obtaining request information of each user in the plurality of users; determining an allocation priority of the corresponding user according to the request information; determining a to-be-served user from the plurality of users that send the pull stream requests according to the allocation priority of the user and a number of digital humans pre-stored in a resource pool; allocating digital humans to the to-be-served user in sequence according to the allocation priority of the to-be-served user; in response to a query request sent by the to-be-served user, synthesizing a digital human video, and displaying the digital human video to the corresponding to-be-served user.
[0074] Those skilled in the art will appreciate that embodiments of the application can be readily used as software, hardware, or a combination of software and hardware. In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0075] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0076] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0077] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0078] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0079] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, for storing, in general, data and / or program instructions. The memory can also include non-volatile memory, such as read-only memory (ROM), electrically programmable read-only memory (EPROM), or electrically erasable programmable read-only memory (EEPROM), for storing, in general, static data and / or program instructions. The memory can also include removable media, such as flash memory, for storing, in general, program instructions and / or data. The memory is an example of computer-readable media.
[0080] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0081] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, method, article or apparatus that includes a list of elements does not only include those elements, but also includes other elements not explicitly listed, or further includes elements inherent in such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.
[0082] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.
Claims
1. A digital human interaction method, characterized by, The digital human interaction method comprises: In the case that a plurality of users send pull stream requests in the same time period, obtaining request information of each user in the plurality of users; determining the allocation priority of the corresponding user according to the request information; determining the user to be served from the plurality of users sending the pull stream request according to the allocation priority of the user and the number of digital humans pre-stored in the resource pool, comprising: if the number of users sending the pull stream request is less than or equal to the number of digital humans, the users sending the pull stream request are determined as the user to be served; if the number of users sending the pull stream request is greater than the number of digital humans, the users are sorted according to the allocation priority, and the users corresponding to the number of digital humans in the sorting are selected as the user to be served; allocating digital humans to the user to be served in turn according to the allocation priority of the user to be served; in response to the query request sent by the user to be served, synthesizing a digital human video and displaying the digital human video to the corresponding user to be served.
2. The digital human interaction method of claim 1, wherein, The method further comprises: determining the number of digital humans pre-stored in the resource pool and the number of pull stream requests received in the same time period; if the number of pull stream requests is more than the number of digital humans and the difference between the two is greater than a set expansion difference value, increasing the number of digital humans in the resource pool; if the number of digital humans is more than the number of pull stream requests and the difference between the two is greater than a set reduction difference value, reducing the number of digital humans in the resource pool.
3. The digital human interaction method of claim 1, wherein, The request information comprises: request scenario, urgency and request time.
4. The digital human interaction method of claim 3, wherein, The request scenario comprises: business question and answer scenario, live broadcast scenario, chat scenario.
5. The digital human interaction method of claim 1, wherein, In response to the query request sent by the user to be served, synthesizing a digital human video and displaying the digital human video to the corresponding user to be served, comprising: synthesizing a digital human video according to the query request and the request information of the user to be served; displaying the digital human video to the corresponding user to be served.
6. The digital human interaction method of claim 5, wherein, The synthesis of the digital human video according to the query request and the request information of the user to be served comprises: identifying the query request of the user to be served and extracting user feature information of the user to be served from the query request; determining the image characteristics of the digital human according to the user feature information of the user to be served and the request information of the user to be served; generating reply information according to the identification result of the query request; synthesizing the digital human video according to the image characteristics and the corresponding reply information.
7. The digital human interaction method of claim 6, wherein, The user feature information comprises: tone information, timbre information, age information, gender information.
8. The digital human interaction method of claim 6, wherein, The image characteristics comprise: shape characteristics, motion characteristics, sound characteristics.
9. The digital human interaction method of claim 6, wherein, The generation of reply information according to the identification result of the query request comprises: training a question and answer model according to a question and answer database; inputting the identification result into the question and answer model to obtain the reply information.
10. The digital human interaction method of claim 9, wherein, After obtaining the reply information, the method further comprises: obtaining the matching degree of the identification result and the corresponding reply information of the query request; adding the matching degree of the identification result and the corresponding reply information to the question and answer database to retrain the question and answer model.
11. A digital human interaction device, characterized by, The digital human interaction device comprises: The digital human access module is configured to, in a case where a plurality of users send pull stream requests in the same time period, acquire request information of each of the plurality of users, determine an allocation priority of a corresponding user according to the request information, and determine a to-be-served user from the plurality of users sending the pull stream requests according to the allocation priority of the user and a quantity of digital humans pre-stored in a resource pool, including: if a quantity of users sending the pull stream requests is less than or equal to the quantity of digital humans, determining the users sending the pull stream requests as the to-be-served users; if the quantity of users sending the pull stream requests is greater than the quantity of digital humans, sorting the users according to the allocation priority, and selecting users corresponding to the quantity of digital humans in the sorting as the to-be-served users; and allocating digital humans to the to-be-served users in sequence according to the allocation priority of the to-be-served users. The digital human driving module is configured to synthesize a digital human video in response to a query request sent by a to-be-served user. The digital human access module is further configured to display the digital human video to the corresponding to-be-served user.
12. A processor, comprising: The digital human interaction method is configured to perform any one of claims 1 to 10.
13. A machine-readable storage medium having stored thereon instructions, the instructions being executable by a machine to cause the machine to: The instruction, when executed by the processor, causes the processor to be configured to perform the digital human interaction method of any one of claims 1 to 10.
14. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the digital human interaction method of any one of claims 1 to 10.
Citation Information
Patent Citations
Communication service arrangement method and device, computer equipment and storage medium
CN113361913A
Digital human video generation method and device, computer equipment and storage medium
CN114170335A