Virtual human action driving method and device based on language analysis and storage medium

By obtaining virtual human movement posture data and voice data, determining the evaluation indicators of the degree of movement fluency, and generating virtual human movements, the problems of movement inconsistency and lag in the existing technology are solved, and higher movement fluency and effect are achieved.

CN120147488AActive Publication Date: 2025-06-13GUANGZHOU HAOCHUAN NETWORK TECH CO LTD

Patent Information

Application Number
CN202510621945.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

When facing a large number of voice commands, the movements are incoherent, stuttered, and the movements are smoother. This is mainly due to the linearity of the virtual human movements generated by interpolation, resulting in unstable effects during special action connections.

Method used

By obtaining the virtual person's movement posture data and the voice data at the current moment, the evaluation indicators of the movement fluency are determined based on the voice data and the action posture data, the virtual person's movement is generated, and the transition action data is used to improve the smoothness of the action.

Benefits of technology

It improves the smoothness of virtual human movements, reduces the situation of incoherence and lag, and significantly improves the action effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147488A_ABST
    Figure CN120147488A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual human action driving method and device based on language analysis and a storage medium, and belongs to the technical field of virtual human. The method comprises the following steps: acquiring motion posture data of a virtual human and voice data at the current moment, wherein the motion posture data comprises motion information of each joint of the virtual human; according to the voice data and the action posture data, determining evaluation indexes of action fluency degrees of all first action data corresponding to the voice data; virtual human actions are generated according to the first action data and second action data, the second action data are transition action data determined according to the first action data and the evaluation indexes, and the technical effect of improving the fluency degree of the virtual human actions is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual humans, and in particular to a virtual human motion driving method, device, and storage medium based on language analysis. Background Art

[0002] With the development of virtual human technology, the usage scenarios of virtual humans are becoming increasingly broad. In most scenarios, there are increasingly high requirements for the interaction capabilities of virtual humans. Currently, the interaction of virtual humans is mainly through voice. Virtual humans can complete some basic actions by receiving voice commands. However, the way virtual humans implement actions is often to select actions from an action database according to voice commands. When faced with a large number of voice commands, problems such as discontinuous actions and lags will occur. This is because the connection between two current actions is based on interpolation, and the virtual human actions generated by interpolation are often linear, resulting in unstable effects during the connection of some special actions, and thus poor action smoothness.

[0003] The above content is only used to assist in understanding the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of the present invention is to provide a virtual human motion driving method, device, and storage medium based on language analysis, aiming to improve the smoothness of virtual human actions.

[0005] To achieve the above objective, the present invention provides a virtual human motion driving method based on language analysis, characterized in that the virtual human motion driving method based on language analysis includes the following steps: Obtain the motion posture data of the virtual human and the voice data at the current moment, where the motion posture data includes: the motion information of each joint of the virtual human; Determine the evaluation index of the motion smoothness of all the first action data corresponding to the voice data according to the voice data and the motion posture data; Generate virtual human actions according to the first action data and the second action data, where the second action data is the transition action data determined according to the first action data and the evaluation index.

[0006] Optionally, the step of determining the evaluation index of the motion smoothness of all the first action data corresponding to the voice data according to the voice data and the motion posture data includes: Extract at least two action keywords and character information of the voice data; Determine first action data based on the described person information, the described action keywords, and the described action posture data, obtaining at least two pieces of first action data, where the first action data corresponds to at least one of the described action keywords; Calculate the evaluation index based on all of the first action data.

[0007] Optionally, the evaluation index includes: a first evaluation index and a second evaluation index. The first evaluation index is the first average joint position distance between two adjacent pieces of first action data, and the second evaluation index is the second average joint position distance between the joint position of the first action data and the joint position of the preset posture. The first action data includes: starting position data and ending position data corresponding to the joints. The step of calculating the evaluation index based on all of the first action data includes: Calculate the first joint distance between the ending position data of the action executed first and the starting position data of the action executed later in two adjacent first actions according to the execution order of all the first action data, obtaining a plurality of the first joint distances, and determining the first evaluation index based on the plurality of first joint distances; Calculate the second joint distances between the preset joint positions of the preset posture and the starting position data and the ending position data of each of the first action data respectively, obtaining a plurality of the second joint distances, and determining the second evaluation index based on the plurality of second joint distances; Use the first evaluation index and the second evaluation index as the evaluation index.

[0008] Optionally, the evaluation index includes: a third evaluation index. The step of calculating the evaluation index based on all of the first action data includes: Calculate the sum of the number of frames corresponding to all the first action data to obtain the total number of frames of the first action; Determine the action rate based on the total number of frames of the first action and the action time corresponding to the voice data; Use the action rate as the third evaluation index.

[0009] Optionally, before the step of determining the action rate based on the total number of frames of the first action and the action time corresponding to the voice data, it further includes: When there is an indication word indicating the action execution speed in the voice data, determine the action time according to the indication word; When there is no indication word indicating the action execution speed in the voice data, determine the action time according to the speech rate corresponding to the voice data.

[0010] Optionally, the evaluation metrics include: a first evaluation metric, a second evaluation metric, and a third evaluation metric. The first evaluation metric is the first average joint position distance between two adjacent first action data. The second evaluation metric is the second average joint position distance between the joint position of the first action data and the joint position of a preset posture. The third evaluation metric is the action rate. Before the step of generating the virtual human action from the first action data and the second action data, the method further includes: Determining a transition label of transition data according to the adjacent first action data; Determining a frame number range of the transition data according to the first evaluation metric, the second evaluation metric, and the third evaluation metric; Matching and determining the second action data in a transition action database according to the transition label and the frame number range.

[0011] Optionally, the step of generating the virtual human action from the first action data and the second action data includes: Generating an action sequence according to the action order, the first action data, and the second action data; Generating the virtual human action according to the action sequence and rendering resources.

[0012] In addition, to achieve the above object, the present invention further provides a virtual human action driving device based on language analysis. The virtual human action driving device based on language analysis includes: An acquisition module, configured to acquire action posture data of a virtual human and voice data at the current moment. The action posture data includes motion information of each joint of the virtual human; An evaluation module, configured to determine an evaluation metric for the action smoothness of all first action data corresponding to the voice data according to the voice data and the action posture data; A generation module, configured to generate a virtual human action according to the first action data and second action data. The second action data is transition action data determined according to the first action data and the evaluation metric.

[0013] In addition, to achieve the above object, the present invention further provides a virtual human action driving device based on language analysis. The virtual human action driving device based on language analysis includes: a memory, a processor, and a virtual human action driving program based on language analysis stored on the memory and executable on the processor. The virtual human action driving program based on language analysis is configured to implement the steps of the virtual human action driving method according to any one of the above.

[0014] In addition, to achieve the above object, the present invention further provides a storage medium, on which a virtual human action driving program based on language analysis is stored. When the virtual human action driving program based on language analysis is executed by a processor, the steps of the virtual human action driving method based on language analysis described in any one of the above are implemented.

[0015] The present invention proposes a virtual human action driving method based on language analysis. The method obtains the action posture data of the virtual human and the speech data at the current moment, and determines an evaluation index of the action smoothness of all the first action data corresponding to the speech data according to the speech data and the action posture data; compared with the frame filling method, it can accurately determine whether there is an unsmooth situation in the action time, rather than directly linearly filling the needles, and can generate virtual human actions according to the first action data and the second action data, so as to effectively improve the smoothness of the actions and further improve the action effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic structural diagram of a virtual human action driving device based on language analysis in the hardware operating environment involved in the embodiment solution of the present invention; Figure 2 is a schematic flowchart of the first embodiment of the virtual human action driving method based on language analysis of the present invention; Figure 3 is a schematic flowchart of the second embodiment of the virtual human action driving method based on language analysis of the present invention; Figure 4 is a schematic flowchart of the third embodiment of the virtual human action driving method based on language analysis of the present invention.

[0017] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0019] Refer to Figure 1 , Figure 1 is a schematic structural diagram of a virtual human action driving device based on language analysis in the hardware operating environment involved in the embodiment solution of the present invention.

[0020] As Figure 1As shown in the figure, the virtual human motion driving device based on language analysis may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, an interaction device 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The interaction device 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard). Optionally, the interaction device 1003 may also be connected to the communication bus through a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (WI-FI) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) memory, or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0021] Those skilled in the art can understand that Figure 1 the structure shown in does not constitute a limitation on the virtual human motion driving device based on language analysis, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0022] As Figure 1 shown, the memory 1005, as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and a virtual human motion driving program based on language analysis.

[0023] In Figure 1 the... device shown, the network interface 1004 is mainly used for data communication with other devices; the interaction device 1003 is mainly used for data interaction with users; the processor 1001 and the memory 1005 in the virtual human motion driving device based on language analysis of the present invention may be arranged in the virtual human motion driving device based on language analysis. The virtual human motion driving device based on language analysis calls the virtual human motion driving program stored in the memory 1005 through the processor 1001 and executes the virtual human motion driving method provided by the embodiments of the present invention.

[0024] The embodiments of the present invention provide a virtual human motion driving method based on language analysis. Referring to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of a virtual human motion driving method based on language analysis of the present invention.

[0025] In this embodiment, the virtual human motion driving method based on language analysis includes: Step S1: Obtain the motion pose data of the virtual human and the voice data at the current moment. The motion pose data includes the motion information of each joint of the virtual human. The motion pose data here can be in the BVH (Biovision Hierarchy) format, which can include skeleton information and motion data. The motion pose data of the virtual human here is not limited to the data of the virtual human at the current moment, but should include the motion data that the current virtual human can achieve. It should be noted that the BVH format data here can be set with identifiers for determining the actions associated with the data, such as: the name of the action, the number of frames of the action, the skeleton identifier corresponding to the action, etc., for example: the identifier of the leg-lifting action, 15 frames, and the 1st male skeleton. The type of the identifier is not limited here. The voice data at the current moment is the received voice, and this voice data is generally data for interacting with the virtual human and can generally be collected through means such as a microphone.

[0026] Step S2: Determine the evaluation index of the motion smoothness of all the first action data corresponding to the voice data according to the voice data and the motion pose data. Specifically, extract the action keywords of the voice data. The action keywords are used to jointly determine the first action data with the motion pose data. After determining the first action data, since the first action data corresponds to the voice data, it is necessary to analyze all the first action data to determine the evaluation index.

[0027] Step S3: Generate the virtual human motion according to the first action data and the second action data. The second action data is the transition action data determined according to the first action data and the evaluation index.

[0028] In this embodiment, determine the second action data corresponding to the first action data according to the result of the evaluation index, and generate the virtual human motion by combining the first action data and the second action data.

[0029] In this embodiment, by obtaining the motion pose data of the virtual human and the voice data at the current moment, and determining the evaluation index of the motion smoothness of all the first action data corresponding to the voice data according to the voice data and the motion pose data; compared with the frame interpolation method, it can accurately determine whether there is an unsmooth situation in the action time, rather than directly linearly interpolating frames, and can generate the virtual human motion according to the first action data and the second action data, thereby effectively improving the smoothness of the motion and further improving the motion effect.

[0030] Further, based on the first embodiment, a second embodiment of the virtual human motion driving method based on language analysis according to the present invention is proposed. In this embodiment, referring to Figure 3 , the step of determining the evaluation index of all the first motion data corresponding to the voice data according to the voice data and the motion posture data includes: Step S21, extracting at least two motion keywords and character information of the voice data; In this embodiment, at least two motion keywords are extracted. Optionally, when there is only one motion keyword in the voice data, a default motion can be selected as the motion keyword. The character information here can be used to screen the first motion data; Step S22, determining the first motion data according to the character information, the motion keywords and the motion posture data, obtaining at least two first motion data, and the first motion data corresponding to at least one of the motion keywords; Determining the first motion data in the motion posture data according to the character information and the keyword, specifically by matching the keyword and the character information with the identifier corresponding to the motion posture data, and determining the first motion data according to the matching result. Of course, in some embodiments, it is necessary to convert the motion keyword into a vocabulary corresponding to the identifier of the motion posture data. Optionally, one motion keyword is converted into keywords of multiple motions.

[0031] Step S23, calculating the evaluation index according to all the first motion data.

[0032] In this embodiment, the number of the evaluation indexes can be one or more, and is used to evaluate the smoothness of the motions corresponding to all the first motion data.

[0033] In this embodiment, by extracting at least two motion keywords and character information of the voice data; determining the first motion data according to the character information, the motion keywords and the motion posture data, obtaining at least two first motion data, and calculating the evaluation index according to all the first motion data, the accuracy of the obtained first motion data can be improved, and further the accuracy of the evaluation index can be improved.

[0034] Further, based on the first embodiment or the second embodiment, a third embodiment of the virtual human motion driving method based on language analysis according to the present invention is proposed. Referring to Figure 4, the evaluation indicators include: a first evaluation indicator and a second evaluation indicator. The first evaluation indicator is the first average joint position distance between two adjacent first action data. The second evaluation indicator is the second average joint position distance between the joint position of the first action data and the joint position of the preset posture. The first action data includes: the starting position data and the ending position data corresponding to the joints. The steps of calculating the evaluation indicators based on all the first action data include: Step S231, calculate the first joint distance between the ending position data of the first action executed first and the starting position data of the second action executed later in two adjacent first actions according to the execution order of all the first action data, to obtain a plurality of the first joint distances, and determine the first evaluation indicator according to the plurality of first joint distances; It should be noted that the ending position data here is specifically the position data of each joint in the last frame of the first action data, and the starting position data here is specifically the position data of each joint in the first frame of the first action data. For two adjacent first action data, the positions of the same joints in the above data can reflect the smoothness during the action process. The greater the first joint distance, the lower the smoothness. It should be noted that the first joint distance here is the distance between the same joint in different first action data, rather than calculating the joint distance using different joints. For example: the joint distance between the position of the knee and the ankle cannot be calculated. Determining the first evaluation indicator according to a plurality of first joint distances can be calculating the average first joint distance as the first evaluation indicator. Of course, for two first action data with the number of joint movements lower than the threshold, the largest first joint distance among the plurality of first joint distances can also be selected as the first evaluation indicator. For example: if the first action is only the rotation action of the arm, selecting the largest first joint distance as the first evaluation indicator can improve the accuracy of the first evaluation indicator. To avoid the influence of joints that are stationary in both adjacent first action data.

[0035] Step S232, calculate the second joint distance between the preset joint position of the preset posture and the starting position data and the ending position data of each first action data respectively, to obtain a plurality of the second joint distances, and determine the second evaluation indicator according to the plurality of second joint distances; Preferably, the preset posture here can be the standing posture, and the preset joint position is the position of the joint in the standing posture. Determining the second evaluation indicator according to the plurality of second joint distances can be calculating the average second joint distance as the second evaluation indicator. The scenario corresponding to the preset posture is often that after completing the previous first action data, the preset posture is first reached, and then the next first action data is executed.

[0036] Step S233, use the first evaluation index and the second evaluation index as the evaluation index.

[0037] Use the first evaluation index and the second evaluation index together as the evaluation index. In other embodiments, only the first evaluation index or the second evaluation index may be used as the evaluation index.

[0038] In this embodiment, the first joint distance between the end position data of the previously executed action and the start position data of the subsequently executed action in two adjacent first actions is calculated based on the execution order of all the first action data, obtaining a plurality of the first joint distances. The first evaluation index is determined according to the plurality of the first joint distances; and the second joint distance between the preset joint position of the preset posture and the start position data and the end position data of each of the first action data is calculated respectively, obtaining a plurality of the second joint distances. The second evaluation index is determined according to the plurality of the second joint distances, so as to accurately reflect the smoothness of the virtual human during the completion of at least two actions.

[0039] Further, based on any of the above embodiments, a fourth embodiment of the virtual human action driving method based on language analysis according to the present invention is proposed. The evaluation index includes: a third evaluation index. The step of calculating the evaluation index according to all the first action data includes: Calculate the sum of the number of frames corresponding to all the first action data to obtain the total number of frames of the first action; Directly calculate the total number of all frames of the first action data saved in the BVH format. In other formats, other methods can also be used to determine the total number of frames.

[0040] Determine the action rate according to the total number of frames of the first action and the action time corresponding to the voice data; The action time here can be determined according to the keywords of the voice data. And the action rate is calculated through the action time and the total number of frames of the first action. The specific calculation method can be the total number of frames of the first action divided by the action time.

[0041] Use the action rate as the third evaluation index.

[0042] In this embodiment, by calculating the sum of the number of frames corresponding to all the first action data to obtain the total number of frames of the first action and determining the action rate according to the total number of frames of the first action and the action time corresponding to the voice data, the evaluation index can be accurately determined.

[0043] Further, before the step of determining the action rate according to the total number of frames of the first action and the action time corresponding to the voice data, it further includes: When there are indicative words indicating the execution speed of the action in the voice data, determine the action time according to the indicative words; Optionally, each first action data is set with an identifier of the default action time for determining the time required by default for the action. When words such as "fast" or "high speed" appear, multiply the default action time by the coefficient corresponding to words such as "high speed" or "fast" to modify the action time. When there are indicative words in the voice data that limit the execution within a certain period of time, use the time range corresponding to the indicative words as the action time. In this embodiment, two different ways of determining the action time can be implemented for different types of indicative words, so as to more accurately determine the action time and avoid problems of too long or too short action time.

[0044] When there are no indicative words indicating the execution speed of the action in the voice data, determine the action time according to the corresponding speech rate of the voice data.

[0045] Specifically, calculate the number of words per unit time in the voice data as the speech rate data, and determine the action time according to the speech rate data and a preset conversion coefficient.

[0046] Further, based on any of the above embodiments, a fifth embodiment of the virtual human action driving method based on language analysis of the present invention is proposed. The evaluation indicators include: a first evaluation indicator, a second evaluation indicator, and a third evaluation indicator. The first evaluation indicator is the first average joint position distance between two adjacent first action data, the second evaluation indicator is the second average joint position distance between the joint position of the first action data and the joint position of the preset posture, and the third evaluation indicator is the action rate. Before the step of generating the virtual human action according to the first action data and the second action data, it further includes: Determine the transition label of the transition data according to the adjacent first action data; Determine the frame number range of the transition data according to the first evaluation indicator, the second evaluation indicator, and the third evaluation indicator; Match and determine the second action data in the transition action database according to the transition label and the frame number range.

[0047] In this embodiment, it should be noted that the transition tags here are used to screen the second action data that meet the requirements. For example, if two adjacent first action data are walking actions and running actions, directly switching from walking to running will result in relatively rigid changes in the positions of each joint. The transition action can maintain the coherence and naturalness of the actions. Therefore, corresponding transition tags are needed to screen out the required second action data. Optionally, determine the frame number range of the transition data according to the weight coefficients corresponding to the first evaluation index, the second evaluation index, and the third evaluation index, and the total number of frames of the first action; and match and determine the second action data in the transition action database according to the transition tags and the frame number range.

[0048] Further, the step of generating the virtual human action according to the first action data and the second action data includes: Generate an action sequence according to the action order, the first action data, and the second action data; Generate the virtual human action according to the action sequence and the rendering resources.

[0049] Specifically, it can be through a rendering resource adaptation unit: dynamically adjust the parameters according to the rendering characteristics of the target platform, and when it is detected that the user manually adjusts the transition parameters, update the data in real time.

[0050] In this embodiment, determine the transition tags of the transition data through the adjacent first action data, determine the frame number range of the transition data according to the first evaluation index, the second evaluation index, and the third evaluation index, and match and determine the second action data in the transition action database according to the transition tags and the frame number range, thereby improving the smoothness of the generated virtual human action.

[0051] In addition, an embodiment of the present invention also proposes a virtual human action driving device based on language analysis, which is characterized in that the virtual human action driving device based on language analysis includes: An acquisition module, configured to acquire the action posture data of the virtual human and the voice data at the current moment, where the action posture data includes: the motion information of each joint of the virtual human; An evaluation module, configured to determine the evaluation index of the action smoothness of all the first action data corresponding to the voice data according to the voice data and the action posture data; A generation module, configured to generate a virtual human action according to the first action data and the second action data, where the second action data is the transition action data determined according to the first action data and the evaluation index.

[0052] The virtual human action driving device based on language analysis can implement the steps of any of the above embodiments.

[0053] In addition, an embodiment of the present invention further provides a virtual human motion driving device based on language analysis. The virtual human motion driving device based on language analysis includes: a memory, a processor, and a virtual human motion driving program stored on the memory and executable on the processor. The virtual human motion driving program is configured to perform the steps of the virtual human motion driving method based on language analysis described in any one of the above.

[0054] In addition, an embodiment of the present invention further provides a storage medium. A virtual human motion driving program is stored on the storage medium. When the virtual human motion driving program is executed by a processor, it implements the steps of the virtual human motion driving method based on language analysis described in any one of the above.

[0055] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.

[0056] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0057] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0058] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A virtual human action driving method based on language analysis, characterized in that: The virtual human action driving method based on language analysis comprises the following steps: Acquire the action posture data of the virtual person and the voice data at the current moment, wherein the action posture data includes: the movement information of each joint of the virtual person; Determine, according to the voice data and the action posture data, an evaluation index of the action fluency of all the first action data corresponding to the voice data; A virtual human action is generated according to the first action data and the second action data, wherein the second action data is transition action data determined according to the first action data and an evaluation index.

2. The method for driving virtual human actions based on language analysis according to claim 1, characterized in that: The step of determining the evaluation index of the action fluency of all the first action data corresponding to the voice data according to the voice data and the action posture data comprises: Extracting at least two action keywords and character information from the voice data; Determine first action data according to the character information, the action keyword and the action posture data to obtain at least two first action data, wherein the first action data corresponds to at least one of the action keywords; The evaluation index is calculated based on all the first action data.

3. The method for driving virtual human actions based on language analysis as claimed in claim 2, characterized in that: The evaluation index includes: a first evaluation index and a second evaluation index, the first evaluation index is a first average joint position distance between two adjacent first motion data, the second evaluation index is a second average joint position distance between the joint position of the first motion data and the joint position of the preset posture, the first motion data includes: starting position data and ending position data corresponding to the joint, and the step of calculating the evaluation index according to all the first motion data includes: Calculate the first joint distance between the end position data of the first action and the start position data of the second action in two adjacent first actions according to the execution order of all the first action data, obtain a plurality of the first joint distances, and determine the first evaluation index according to the plurality of first joint distances; Respectively calculating the second joint distances between the preset joint position of the preset posture and the starting position data and the ending position data of each of the first motion data to obtain a plurality of the second joint distances, and determining the second evaluation index according to the plurality of the second joint distances; The first evaluation index and the second evaluation index are used as the evaluation index.

4. The method for driving virtual human actions based on language analysis as claimed in claim 2, characterized in that: The evaluation index includes: a third evaluation index, and the step of calculating the evaluation index according to all the first action data includes: Calculate the sum of the frame numbers corresponding to all the first motion data to obtain the total frame number of the first motion; Determine the motion rate according to the total number of frames of the first motion and the motion time corresponding to the voice data; The action rate is used as the third evaluation index.

5. The method for driving virtual human actions based on language analysis as claimed in claim 4, characterized in that: Before the step of determining the action rate according to the total number of frames of the first action and the action time corresponding to the voice data, the method further includes: When the voice data contains an indicative word indicating the speed of executing the action, determining the action time according to the indicative word; When the voice data does not contain an indicative word indicating the speed of executing the action, the action time is determined according to the corresponding speech speed of the voice data.

6. The method for driving virtual human actions based on language analysis according to claim 1, characterized in that: The evaluation index includes: a first evaluation index, a second evaluation index and a third evaluation index, the first evaluation index is a first average joint position distance between two adjacent first motion data, the second evaluation index is a second average joint position distance between the joint position of the first motion data and the joint position of the preset posture, and the third evaluation index is a motion rate. Before the step of generating a virtual human motion according to the first motion data and the second motion data, the step also includes: Determine a transition label of transition data according to the adjacent first action data; Determine a frame number range of transition data according to the first evaluation index, the second evaluation index, and the third evaluation index; The second action data is determined by matching the transition tag and the frame number range in a transition action database.

7. The method for driving virtual human actions based on language analysis according to claim 6, characterized in that: The step of generating a virtual human action according to the first action data and the second action data comprises: generating an action sequence according to an action order, the first action data, and the second action data; The virtual human action is generated according to the action sequence and the rendering resources.

8. A virtual human motion driving device based on language analysis, characterized in that: The virtual human action driving device based on language analysis includes: An acquisition module, used to acquire the action posture data of the virtual person and the voice data at the current moment, wherein the action posture data includes: the motion information of each joint of the virtual person; An evaluation module, used to determine an evaluation index of the action fluency of all first action data corresponding to the voice data according to the voice data and the action posture data; A generation module is used to generate virtual human actions according to the first action data and second action data, wherein the second action data is transition action data determined according to the first action data and an evaluation index.

9. A virtual human motion driving device based on language analysis, characterized in that: The language analysis-based virtual human action driving device comprises: a memory, a processor, and a language analysis-based virtual human action driving program stored in the memory and executable on the processor, wherein the language analysis-based virtual human action driving program is configured to implement the steps of the language analysis-based virtual human action driving method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores a virtual human action driving program based on language analysis, and when the virtual human action driving program based on language analysis is executed by the processor, the steps of the virtual human action driving method based on language analysis as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and system for driving virtual character to act by real-time voice

    CN111939558A

  • 3D virtual digital human interaction action generation method and system based on deep learning

    CN115797606A

  • Human body motion posture evaluation method and system

    CN119992663A

Cited By

  • Digital human quality assessment method and device

    CN120726537A