Digital human interaction method, system and device, medium and program product
Through artificial intelligence technology, the input data is detected in the 2D digital human system and the interaction strategy is determined, which solves the problem of insufficient flexibility in the existing digital human system interaction, and realizes more intelligent and flexible human-computer interaction.
Patent Information
- Application Number
- CN202510123024.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing 2D digital human system is difficult to achieve real-time two-way interaction, and the interaction mode is insufficient to meet the diverse user needs.
Through artificial intelligence technology, the input data is detected, the interaction strategy corresponding to the interaction object associated with the input data is determined, and the target digital person outputs the response data based on this strategy is controlled to achieve flexible interaction between digital persons and users.
It improves the flexibility and pertinence of interaction between digital people and users, can meet diverse user needs, and makes the human-computer interaction functions of the digital human interaction system smarter, more flexible and more humanized.
Smart Images

Figure CN120045670A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a digital human interaction method, system, device, medium and program product. Background Art
[0002] In the current development pattern of digital human technology, 2D digital human systems usually can only meet the one-way information conduction function and it is difficult to achieve real-time two-way interaction with users. At the same time, when digital humans face various questions raised by users, they can only rely on pre-set fixed scripts and databases with limited capacity to achieve fixed interaction modes. The flexibility of the above digital human interaction methods is insufficient and cannot meet the diverse interaction needs between digital humans and users in practical applications. Summary of the Invention
[0003] Based on the above technical problems, embodiments of this application provide a digital human interaction method, system, device, medium and program product.
[0004] The technical solution provided by the embodiments of this application is as follows:
[0005] Embodiments of this application first provide a digital human interaction method, and the method includes:
[0006] If the digital human interaction system detects input data, determine an interaction strategy corresponding to the interaction object associated with the input data through artificial intelligence technology;
[0007] Analyze the input data to determine response data;
[0008] Based on the interaction strategy, control the target digital human matching the interaction object to output the response data to achieve the interaction between the target digital human and the interaction object.
[0009] In some embodiments, the determining, through artificial intelligence technology, an interaction strategy corresponding to the interaction object associated with the input data includes:
[0010] Obtain the scene parameters of the current scene where the interaction object is located;
[0011] Determine the object parameters of the interaction object;
[0012] Through the artificial intelligence technology, determine the interaction strategy based on at least the scene parameters and the object parameters.
[0013] In some embodiments, the object parameters include interaction requirement data corresponding to the interaction object; the scene parameters include brightness parameters and / or noise parameters of the scene where the interaction object is currently located; determining the interaction strategy based at least on the scene parameters and the object parameters includes:
[0014] Based on the degree of difference between the interaction requirement data and the brightness parameter and / or the noise parameter, determining the interaction strategy as: adjusting the display state and / or audio output state of the target digital human based on the degree of difference.
[0015] In some embodiments, before controlling the target digital human matching the interaction object to output the response data based on the interaction strategy, it further includes:
[0016] If a start instruction is detected, obtain configuration parameters associated with the interaction object;
[0017] Create the target digital human based on the configuration parameters.
[0018] In some embodiments, before determining the interaction strategy corresponding to the interaction object associated with the input data, it further includes:
[0019] Statistically analyze the sampling point amplitudes of the audio sampling points included in the input data to obtain the first audio stream intensity;
[0020] If the first audio stream intensity is greater than or equal to the intensity threshold, continuously collect the input data.
[0021] In some embodiments, the method further includes:
[0022] Statistically analyze the sampling point amplitudes of the audio sampling points included in the input data to obtain the second audio stream intensity of the input data;
[0023] If the continuous time period during which the second audio stream intensity is less than the intensity threshold is greater than or equal to the time period threshold, stop creating the target digital human, or end the interaction between the target digital human and the interaction object.
[0024] An embodiment of the present application further provides a digital human interaction system, and the digital human interaction system includes:
[0025] A collection module, configured to collect input data;
[0026] A processing module, configured to, if the digital human interaction system detects the input data, determine an interaction strategy corresponding to the interaction object associated with the input data through artificial intelligence technology; analyze the input data to determine response data;
[0027] An interaction module, configured to control a target digital human matching with the interaction object to output the response data based on the interaction strategy, so as to implement the interaction between the target digital human and the interaction object.
[0028] An embodiment of the present application further provides an electronic device, which includes a processor and a memory; a computer program is stored in the memory; when the computer program is executed by the processor, it can implement the digital human interaction method as described in any one of the previous items.
[0029] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored; when the computer program is executed by a processor of an electronic device, it can implement the digital human interaction method as described in any one of the previous items.
[0030] An embodiment of the present application further provides a computer program product, which includes a computer program; when the computer program is executed by a processor of an electronic device, it can implement the digital human interaction method as described in any one of the previous items.
[0031] For the digital human interaction method provided by the embodiment of the present application, if the digital human interaction system detects input data, the interaction strategy corresponding to the interaction object associated with the input data is determined through AI technology. In this way, not only is the action of determining the interaction strategy strictly controlled, but also the matching degree between the interaction strategy, the input data, and the interaction object can be improved, thereby improving the pertinence of the interaction strategy; and after analyzing the input data to determine the response data, the target digital human matching with the interaction object is controlled to output the response data based on the interaction strategy. In this way, not only is the matching degree between the target digital human and the interaction object improved, but also the matching degree between the response data output by the target digital human and the interaction requirements of the interaction object is improved, thereby improving the flexibility and pertinence of the interaction between the interaction object and the target digital human, being able to meet the diverse digital human interaction needs of the interaction object, and making the human-computer interaction function of the digital human interaction system more intelligent, flexible, and user-friendly. Description of the Drawings
[0032] Figure 1 It is a schematic flowchart of the digital human interaction method provided by the embodiment of the present application;
[0033] Figure 2 It is a schematic structural diagram of the digital human interaction system provided by the embodiment of the present application;
[0034] Figure 3 It is another schematic structural diagram of the digital human interaction system provided by the embodiment of the present application;
[0035] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0037] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0038] In recent years, digital human technology has achieved remarkable development and demonstrated broad application prospects in many fields. Among them, a digital human is a virtual image with appearance features and behavior patterns produced through technological means such as modeling, motion capture, or artificial intelligence (AI), and the digital human can be presented through a display device. The value of a digital human lies in being able to break through the spatial limitations of the physical world and provide a greater sense of immersion, participation, and interaction.
[0039] In practical applications, digital humans produced through modeling and motion capture are usually 3D digital humans, while 2D digital humans comparable to real humans can be produced through AI technology. Among them, 2D digital humans can be set according to real-world characters, and their appearances can also be exactly the same as those of real-world characters. In practical applications, 2D digital humans restored and produced based on real-world characters are also called digital twins. As an important branch of digital human technology, 2D digital humans can replicate real human images 1:1 and present them on the screen to provide information and services for users. At present, 2D digital humans have gradually emerged in scenarios such as customer service, education, and live streaming.
[0040] At the same time, in the current development pattern of digital human technology, 2D digital human systems usually can only meet the one-way information conduction function and it is difficult to achieve real-time two-way interaction with users; at the same time, when digital humans face various questions raised by users, they can only rely on pre-set fixed scripts and databases with limited capacity to achieve fixed interaction modes. The flexibility of the above digital human interaction methods is insufficient to meet the diverse interaction needs between digital humans and users in practical applications.
[0041] Based on the above technical problems, the embodiments of the present application provide a digital human interaction method, system, device, medium, and program product. The digital human interaction method provided by the embodiments of the present application can be applied to a digital human interaction system.
[0042] In one implementation manner, a digital human interaction system may include multiple models or modules that are independent of each other in function but coupled in the data processing process.
[0043] Figure 1 For the flowchart of the digital human interaction method provided by the embodiments of the present application, as Figure 1As shown, the method may include the following steps:
[0044] Step 101: If the digital human interaction system detects input data, determine the interaction strategy corresponding to the interaction object associated with the input data through AI technology.
[0045] Correspondingly, if the digital human interaction system does not detect input data, the interaction strategy may not be determined.
[0046] In one implementation, the input data may include at least one of audio data, gestures, body movements, and text data.
[0047] In one implementation, the interaction strategy may include at least one of the ways, steps, links, and the logical jump sequence between each link for interacting with the interaction object.
[0048] In one implementation, the interaction object may include a person or a pet.
[0049] In one implementation, the interaction strategy may be determined in the following way:
[0050] Collect the object image of the interaction object through AI technology, and identify the object image to determine the object attributes of the interaction object. Then, determine the interaction strategy based on the object attributes and the input data to improve the fluency and efficiency of the interaction with the interaction object; Exemplarily, the object attributes may include data such as object type, object gender, and object age. For example, if the object type is a pet, the habitual interaction method of the pet can be determined through AI technology, and then the interaction strategy can be determined based on the habitual interaction method. Another example is that if the object age is an elderly person, the interaction methods such as large font, bright color, and high volume commonly used by the elderly can be determined through AI technology to determine the interaction strategy with the elderly.
[0051] Step 102: Analyze the input data to determine the response data.
[0052] In one implementation, the response data may include at least one of text data, voice data, and image data.
[0053] In one implementation, the response data may be determined in the following way:
[0054] Perform semantic recognition on the input data through AI technology to obtain the semantic recognition result. Then, determine the response data and / or reply data expected by the interaction object based on the semantic recognition result, and determine the response data and / or reply data as the response data.
[0055] Step 103: Based on the interaction strategy, control the target digital human that matches the interaction object to output response data, so as to realize the interaction between the target digital human and the interaction object.
[0056] In one implementation, the target digital human can be created in advance through AI technology.
[0057] In one implementation, the target digital humans matched by different interaction objects can be different. For example, the target digital human matched by a child object can include a 2D digital human with a child image, and the target digital human matched by a pet object can include a 2D digital human wearing a pet image costume.
[0058] In one implementation, the interaction between the target digital human and the interaction object can be realized in any of the following ways:
[0059] Control the actions of the target digital human based on the interaction actions included in the interaction strategy, and control the target digital human to output text data or audio data corresponding to the response data based on the speech rate or intonation included in the interaction strategy, so as to realize the interaction between the target digital human and the interaction object.
[0060] Control the target digital human to perform role actions corresponding to the role information based on the role information included in the interaction strategy, and control the target digital human to output response data in the role manner corresponding to the role information; wherein, the role information can include a customer service role, a rescue role, etc.
[0061] As can be seen from the above, in the digital human interaction method provided by the embodiments of the present application, if the digital human interaction system detects input data, the interaction strategy corresponding to the interaction object associated with the input data is determined through AI technology. In this way, not only is the action of determining the interaction strategy strictly controlled, but also the matching degree between the interaction strategy and the input data and the interaction object can be improved, thereby improving the pertinence of the interaction strategy; and after analyzing the input data to determine the response data, based on the interaction strategy, control the target digital human that matches the interaction object to output the response data. In this way, not only is the matching degree between the target digital human and the interaction object improved, but also the matching degree between the response data output by the target digital human and the interaction requirements of the interaction object is improved, thereby improving the flexibility and pertinence of the interaction between the interaction object and the target digital human, meeting the diverse digital human interaction needs of the interaction object, and making the human-computer interaction function of the digital human interaction system more intelligent, flexible and user-friendly.
[0062] Based on the foregoing embodiments, in the digital human interaction method provided by the embodiments of the present application, determining the interaction strategy corresponding to the interaction object associated with the input data through AI technology can be implemented through the following steps:
[0063] Step A1: Obtain the scene parameters of the scene where the interaction object is currently located.
[0064] In one implementation, the scene in which the interactive object is currently located may include an office scene, a home scene, a commuting scene, a travel scene, and the like.
[0065] In one embodiment, the scene parameters can be used to describe the environmental state of the scene in which the interactive object is currently located; for example, the environmental state can include whether the environment is crowded or empty, and can also include the temperature, humidity, altitude, wind speed, weather conditions and other states of the environment.
[0066] In one implementation, the scene parameters may be obtained in any of the following ways:
[0067] The data of the scene in which the interactive object is currently located is collected by the collection device included in the digital human interaction system, and the above data is analyzed to determine the scene parameters; illustratively, the collection device may include a camera, a temperature sensor, a humidity sensor, a gyroscope and other devices.
[0068] The digital human interaction system obtains the location information of the current location of the interactive object, and searches in the network space based on the location information, so as to obtain the current scene and scene parameters of the interactive object.
[0069] Step A2: Determine the object parameters of the interactive object.
[0070] In one embodiment, the object parameters may include data such as the object type, the object's gender, the object's clothing characteristics, and the object's height or body shape.
[0071] In one embodiment, the object parameters may be determined by:
[0072] With the permission of the interactive object, the image data containing the pixel features of the interactive object is collected by the collection device provided in the digital human interaction system, and the features of the image data are extracted and recognized to determine the object parameters.
[0073] Step A3: Determine the interaction strategy based on at least scene parameters and object parameters through AI technology.
[0074] In one embodiment, the interaction strategy may be determined by:
[0075] The scene parameters and object parameters are analyzed by AI technology to determine the degree of match between the scene parameters and the object parameters, and then the interaction strategy is determined based on the degree of match; illustratively, if the degree of match is less than the degree threshold, the interaction strategy can be determined to be a question-and-answer strategy. For example, if the scene parameters indicate that the current scene of the interactive object is an office scene, but the object parameters indicate that the interactive object is a child, then the interaction strategy can be determined to be: a question-and-answer strategy for providing question consultation for the child.
[0076] As can be seen from the above, in the digital human interaction method provided by the embodiments of the present application, obtaining the scene parameters of the current scene where the interaction object is located can achieve accurate tracking and recognition of the characteristics of the current scene where the interaction object is located; moreover, by determining the object parameters of the interaction object, accurate control of the characteristics of the interaction object can be achieved; on this basis, through AI technology, determining the interaction strategy based at least on the scene parameters and the object parameters can improve the matching degree between the interaction strategy and the scene parameters and the object parameters, so that the interaction method corresponding to the interaction strategy can meet the interaction needs of the current scene where the interaction object is located and the interaction object itself.
[0077] Based on the foregoing embodiments, in the digital human interaction method provided by the embodiments of the present application, the object parameters include the interaction requirement data corresponding to the interaction object; the scene parameters include the brightness parameter and / or the noise parameter of the current scene where the interaction object is located.
[0078] In one implementation, the interaction requirement data may include a set of characteristics that the interaction object expects the target digital human to possess when interacting with the target digital human.
[0079] In one implementation, the interaction requirement data may be input by the interaction object in advance, or may be obtained from the cyber space associated with the interaction object with the permission of the interaction object, or may be obtained by analyzing the behavior and / or expression of the interaction object over at least a period of time. The embodiments of the present application do not make any limitations thereto.
[0080] In one implementation, the brightness parameter may include the intensity of the brightness of the current scene where the interaction object is located, and may also include the spectrum or wavelength of the light in the current scene where the interaction object is located.
[0081] In one implementation, the noise parameter may include the amplitude of the noise in the current scene where the interaction object is located, and may also include the frequency components included in the noise.
[0082] Correspondingly, determining the interaction strategy based at least on the scene parameters and the object parameters can be achieved in the following manner:
[0083] Based on the degree of difference between the interaction requirement data and the brightness parameter and / or the noise parameter, determine the interaction strategy as: adjusting the display state and / or the audio output state of the target digital human based on the degree of difference.
[0084] In one implementation, the degree of difference may include a first difference and / or a second difference; wherein, the first difference may include the light difference between the brightness requirement in the interaction requirement parameters and the brightness parameter, and the second difference may include the audio difference between the ambient sound requirement in the interaction requirement parameters and the noise data.
[0085] In one embodiment, the display state may include the color and / or brightness of the physical appearance of the target digital human.
[0086] In one embodiment, the audio output state may include at least one of the frequency, timbre, and amplitude of the audio data output by the target digital human.
[0087] In one embodiment, the interaction strategy may be determined in the following manner:
[0088] If the first difference is greater than or equal to the light threshold and / or the second difference is greater than or equal to the audio threshold, determine the interaction strategy: adjust the display state of the target digital human based on the first difference to improve the recognition rate of the target digital human during the interaction, and / or, adjust the amplitude of the audio output by the target digital human based on the second difference to improve the recognition rate of the audio data output by the target digital human in the environment where the interaction object is currently located.
[0089] As can be seen from the above, in the digital human interaction method provided by the embodiments of the present application, the object parameters include the interaction requirement data corresponding to the interaction object, the scene parameters include the brightness parameter and / or noise parameter of the scene where the interaction object is currently located, and, based on the degree of difference between the interaction requirement data and the brightness parameter and / or noise parameter, determine the interaction strategy as: adjusting the display state and / or audio output state of the target digital human based on the degree of difference. In this way, through the above operations, during the interaction between the interaction object and the target digital human, the recognition rate of the target digital human and / or its audio output state can be improved, thereby improving the interaction efficiency between the interaction object and the target digital human and improving the interaction effect between the two.
[0090] Based on the foregoing embodiments, in the digital human interaction method provided by the embodiments of the present application, before controlling the target digital human matching the interaction object to output response data based on the interaction strategy, the following operations may also be performed:
[0091] If a start instruction is detected, obtain the configuration parameters associated with the interaction object; create a target digital human based on the configuration parameters.
[0092] Correspondingly, if the start instruction is not detected, the configuration parameters may not be obtained.
[0093] In one embodiment, the start instruction may be input by the interaction object or by a professional technician.
[0094] In one embodiment, the start instruction may be input in the form of voice, gesture, and action, etc.; correspondingly, the data corresponding to the start instruction may include instruction audio, instruction gesture, and instruction action, etc.
[0095] In one embodiment, when the audio acquisition module in the digital human interaction system acquires the audio stream input by the interaction object, it can judge the data in the audio stream based on a preset volume threshold, so as to screen the audio data in the audio stream and obtain valid audio data; at this time, the speech-to-text module in the digital human interaction system receives the valid audio data, recognizes and converts the valid audio data to obtain valid text data, and then matches the keywords in the valid text data to obtain a matching result. If the matching result indicates that the above keywords match the start instruction, it can be determined that the start instruction is detected. At this time, the audio acquisition module can be controlled to stop the audio acquisition operation, and the digital human system can be set to the awakened state.
[0096] In one embodiment, the configuration parameters can be input by the interaction object, or can be obtained by the digital human interaction system from the cyberspace with the permission of the interaction object.
[0097] In one embodiment, the configuration parameters can include interaction requirement parameters, and can also include object parameters of the interaction object, and can also include appearance image characteristics, motion characteristics, and audio characteristics that the interaction object expects the target digital human to have, etc.
[0098] In one embodiment, the target digital human can be realized in the following way:
[0099] Identify the types of parameters and the parameter values of each parameter in the configuration parameters through AI technology, and determine the feature set that the digital human should have according to the recognition result of the parameters in the configuration parameters, and then create the target digital human according to the feature set.
[0100] As can be seen from the above, in the digital human interaction method provided by the embodiments of the present application, if a start instruction is detected, the configuration parameters associated with the interaction object are obtained, and the target digital human is created based on the configuration parameters. In this way, through the above operations, the matching degree between the target digital human and the configuration parameters associated with the interaction object is improved, so that the probability that the target digital human meets the interaction requirements of the interaction object can be improved.
[0101] Based on the foregoing embodiments, in the digital human interaction method provided by the embodiments of the present application, before determining the interaction strategy corresponding to the interaction object associated with the input data, the following operations can also be performed:
[0102] Statistically analyze the sampling point amplitudes of the audio sampling points included in the input data to obtain the first audio stream intensity; if the first audio stream intensity is greater than or equal to the intensity threshold, continuously acquire the input data.
[0103] Correspondingly, if the first audio stream intensity is less than the intensity threshold, the acquisition of the input data can be stopped.
[0104] In one embodiment, the audio acquisition module of the digital human interaction system can acquire input data and sample the input data to obtain the sample point amplitude a corresponding to the single-channel sample point data 1 , a 2 , a 3 , … a i , where a i is the i-th audio sample point data, and the first audio stream intensity can also be calculated by Equation (1):
[0105]
[0106] where i is an integer greater than or equal to 1 and less than or equal to n, and n can be an integer greater than 3.
[0107] Exemplarily, a c can be a preset intensity threshold. If a p is greater than a t , the audio acquisition module can continuously acquire input data. Otherwise, it can be determined that the input data is invalid data, and at this time, the input data can be discarded.
[0108] As can be seen from the above, the digital human interaction method provided by the embodiments of the present application statistically obtains the first audio stream intensity from the sample point amplitudes of the audio sample points included in the input data. If the first audio stream intensity is greater than or equal to the intensity threshold, the input data is continuously acquired. In this way, through the above method, the filtering and screening of the input data are realized, so that the validity of the input data can be improved.
[0109] Based on the foregoing embodiments, the digital human interaction method provided by the embodiments of the present application can also perform the following operations:
[0110] Statistically obtain the second audio stream intensity of the input data from the sample point amplitudes of the audio sample points included in the input data; if the continuous period during which the second audio stream intensity is less than the intensity threshold is greater than or equal to the period threshold, stop creating the target digital human, or end the interaction between the target digital human and the interaction object.
[0111] Correspondingly, if the second audio stream intensity is greater than or equal to the intensity threshold, or the continuous period during which the second audio stream intensity is less than the intensity threshold is less than the period threshold, the target person can be continuously created, or the interaction between the target digital human and the interaction object can be continued.
[0112] In one embodiment, the process of obtaining the second audio stream intensity can be the same as the process of obtaining the first audio stream intensity, and will not be elaborated here.
[0113] In one embodiment, if the audio acquisition module of the digital human interaction system acquires input data, it can continuously perform sampling on the input data to obtain audio sampling points, and statistically analyze the amplitudes of the corresponding sampling points to obtain the second audio stream intensity.
[0114] In one embodiment, if the duration during which the second audio stream intensity is less than the intensity threshold is greater than or equal to the duration threshold, it may indicate that the interaction object has stopped the current interaction process. At this time, if the digital human interaction system is in the state of creating a target digital human, the creation of the target digital human can be stopped. If the digital human interaction system is in the state of the target digital human interacting with the interaction object, the interaction between the target digital human and the interaction object can be ended.
[0115] As can be seen from the above, in the digital human interaction method provided by the embodiments of the present application, after statistically analyzing the amplitudes of the audio sampling points included in the input data to obtain the second audio stream intensity, if the duration during which the second audio stream intensity is less than the intensity threshold is greater than or equal to the duration threshold, the creation of the target digital human is stopped, or the interaction between the target digital human and the interaction object is ended. In this way, through the above steps, the real-time tracking detection and judgment of the second audio stream intensity corresponding to the input data are realized, and the real-time tracking control of whether to stop creating the target digital human and whether to end the interaction between the target digital human and the interaction object is also realized.
[0116] Figure 2 FIG. is a schematic structural diagram of the digital human interaction system provided by the embodiments of the present application. As Figure 2 shown, the digital human interaction system 2 may include an audio acquisition module 201, a speech-to-text module 202, a large model module 203, a text-to-speech module 204, a digital human rendering module 205, and a video synthesis module 206.
[0117] Exemplarily, the audio acquisition module 201 is configured to acquire the audio data output by the interaction object, and identify the audio data according to the set volume threshold to obtain effective audio stream data; the speech-to-text module 202 is configured to identify the effective audio stream data, convert it into text data, and then perform a matching operation based on the keywords in the text data. For example, if the keyword match is to start or create a digital human, the digital human system can be set to the awakened state; the large model module 203 is configured to identify the text data, and thus output the corresponding text recognition result; the text-to-speech module 204 is configured to convert the text recognition result to obtain the corresponding output audio stream data; the digital human rendering module 205 is configured to create a target digital human based on the output audio stream data, and generate video stream data for the target digital human to interact; the video synthesis module 206 is configured to output the video stream data, so as to realize the interaction with the interaction object.
[0118] Exemplarily, the speech-to-text module 202 may include an AI deep learning model that can real-time recognize audio stream data into text data; the large model module 203 may include a large language model (LLM), which can identify text data through a keyword matching method combined with a matching algorithm of word meaning and pinyin. The LLM can also determine response data corresponding to the input data in combination with a knowledge base and output streaming response data to subsequent modules; the text-to-speech module 204 can also be implemented through an AI deep learning model, which can assemble the streaming response data into complete sentences and convert them into audio stream data in real time; the digital human rendering module 205 can also be implemented based on an AI deep learning model. It can, through a hardware acceleration method, synthesize target digital human video frame data in real time, and push the target digital human video frame data and audio stream data to the video synthesis module 206 for the video synthesis module 206 to encode the video frame data and audio stream data into a final digital human interaction video result by means of audio-visual coding technology.
[0119] Through the above-mentioned various modules, real-time processing of audio data output by the interaction object can be achieved, and the real-time performance, flexibility, and intelligence level of the interaction between the digital human interaction system and the interaction object can also be improved.
[0120] Figure 3 Another structural schematic diagram of the digital human interaction system provided by the embodiments of this application is shown in Figure 3 As shown, the digital human interaction system 2 may include
[0121] A collection module 301 for collecting input data;
[0122] A processing module 302, which, if the digital human interaction system detects input data, determines an interaction strategy corresponding to the interaction object associated with the input data through artificial intelligence technology; analyzes the input data to determine response data;
[0123] An interaction module 302 for controlling a target digital human matching the interaction object to output response data based on the interaction strategy, so as to achieve the interaction between the target digital human and the interaction object.
[0124] In some embodiments, the processing module 302 is used to obtain scene parameters of the current scene where the interaction object is located; determine object parameters of the interaction object; and determine an interaction strategy through artificial intelligence technology based on at least the scene parameters and the object parameters.
[0125] In some embodiments, the object parameters include interaction requirement data corresponding to the interaction object; the scene parameters include the brightness parameter and / or noise parameter of the current scene where the interaction object is located;
[0126] The processing module 302 is configured to determine an interaction strategy based on the degree of difference between the interaction requirement data and the brightness parameter and / or the noise parameter, where the interaction strategy is to adjust the display state and / or the audio output state of the target digital human based on the degree of difference.
[0127] In some embodiments, the processing module 302 is configured to, if a start instruction is detected, obtain configuration parameters associated with the interaction object, and create a target digital human based on the configuration parameters.
[0128] In some embodiments, the acquisition module 301 is configured to count the sampling point amplitudes of the audio sampling points included in the input data to obtain a first audio stream intensity. If the first audio stream intensity is greater than or equal to an intensity threshold, continuously acquire the input data.
[0129] In some embodiments, the acquisition module 301 is configured to count the sampling point amplitudes of the audio sampling points included in the input data to obtain a second audio stream intensity of the input data.
[0130] The processing module 302 is configured to, if the continuous period during which the second audio stream intensity is less than the intensity threshold is greater than or equal to a period threshold, stop creating the target digital human, or end the interaction between the target digital human and the interaction object.
[0131] The embodiments of the present application further provide an electronic device. Figure 4 As shown in the structural schematic diagram of the electronic device provided by the embodiments of the present application, Figure 4 the electronic device 4 includes a processor 401 and a memory 402. A computer program is stored in the memory 402. When the computer program is executed by the processor 401, it can implement the digital human interaction method described in any of the previous items.
[0132] The embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor of an electronic device, it can implement the digital human interaction method described in any of the previous items.
[0133] The embodiments of the present application further provide a computer program product, where the program product includes a computer program. When the computer program is executed by a processor of an electronic device, it can implement the digital human interaction method described in any of the previous items.
[0134] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated herein.
[0135] The methods disclosed in the method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.
[0136] The features disclosed in each product embodiment provided by this application can be combined arbitrarily without conflict to obtain new product embodiments.
[0137] The features disclosed in each method or device embodiment provided by this application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.
[0138] It should be noted that the above computer-readable storage medium can be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; it can also be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0139] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0140] The serial numbers of the embodiments of this application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus necessary general hardware nodes. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0142] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0145] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A digital human interaction method, characterized in that: The method is applied to a digital human interaction system; comprising: If the digital human interaction system detects input data, an interaction strategy corresponding to an interaction object associated with the input data is determined by artificial intelligence technology; Analyze the input data to determine response data; Based on the interaction strategy, the target digital human matching the interaction object is controlled to output the response data, so as to realize the interaction between the target digital human and the interaction object.
2. The method according to claim 1, characterized in that The step of determining the interaction strategy corresponding to the interaction object associated with the input data by using artificial intelligence technology includes: Obtaining scene parameters of the scene in which the interactive object is currently located; Determining object parameters of the interactive object; The interaction strategy is determined by the artificial intelligence technology at least based on the scene parameters and the object parameters.
3. The method according to claim 2, characterized in that The object parameter includes the interaction requirement data corresponding to the interactive object; the scene parameter includes the brightness parameter and / or noise parameter of the scene where the interactive object is currently located; The determining the interaction strategy based at least on the scene parameter and the object parameter comprises: Based on the difference between the interaction demand data and the brightness parameter and / or the noise parameter, the interaction strategy is determined as: adjusting the display state and / or audio output state of the target digital human based on the difference.
4. The method according to claim 1, characterized in that Before controlling the target digital person matching the interactive object to output the response data based on the interactive strategy, the method further includes: If a start instruction is detected, obtaining configuration parameters associated with the interactive object; The target digital human is created based on the configuration parameters.
5. The method according to claim 1, characterized in that Before determining the interaction strategy corresponding to the interaction object associated with the input data, the method further includes: Counting the sampling point amplitudes of the audio sampling points contained in the input data to obtain a first audio stream strength; If the intensity of the first audio stream is greater than or equal to the intensity threshold, the input data is continuously collected.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Performing statistics on sampling point amplitudes of audio sampling points contained in the input data to obtain a second audio stream intensity of the input data; If the duration during which the intensity of the second audio stream is less than the intensity threshold is greater than or equal to the period threshold, the creation of the target digital human is stopped, or the interaction between the target digital human and the interactive object is ended.
7. A digital human interaction system, characterized in that: The digital human interaction system comprises: A collection module, used for collecting input data; A processing module, configured to determine an interaction strategy corresponding to an interaction object associated with the input data by using artificial intelligence technology if the digital human interaction system detects the input data; and to analyze the input data to determine response data; The interaction module is used to control the target digital human matching the interaction object to output the response data based on the interaction strategy, so as to realize the interaction between the target digital human and the interaction object.
8. An electronic device, characterized in that: The electronic device includes a processor and a memory; a computer program is stored in the memory; when the computer program is executed by the processor, the digital human interaction method as described in any one of claims 1 to 6 can be implemented.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program; when the computer program is executed by a processor of an electronic device, the digital human interaction method as described in any one of claims 1 to 6 can be implemented.
10. A computer program product, characterized in that The program product includes a computer program; when the computer program is executed by a processor of an electronic device, it can implement the digital human interaction method as described in any one of claims 1 to 6.