Digital human intelligent interaction control method and system in private domain live broadcast scene

By performing emotion recognition and normalization on the bullet screen text stream and calculating the group emotion change rate, the problem of the inability to perceive the trend of group emotion changes in existing technologies is solved. This achieves the accuracy of digital human interaction control and cross-session optimization, and improves the interaction effect in private domain live streaming scenarios.

CN122027840BActive Publication Date: 2026-06-19TIANJIN BAIMA PLANET INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-13
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing digital human interaction control systems cannot perceive the relative changing trends of group emotions, resulting in a disconnect between the timing of triggering interactive control commands and the emotional state of the group in the live broadcast room. Furthermore, control parameters cannot be adaptively updated across sessions, and historical user data in private domain live broadcast scenarios cannot be effectively utilized.

Method used

By performing emotion recognition on the bullet screen text stream, a multi-dimensional emotion intensity vector is obtained and normalized. The group emotion change rate is calculated using short-term and long-term control windows. Combined with the interaction trigger threshold and private domain language library, the digital human is driven to output an action sequence and the control parameters are updated within the verification time window.

Benefits of technology

It achieves accurate perception of the changing trends of the audience's emotions in the live broadcast room, can trigger proactive resolution control in a timely manner, improves the accuracy of interactive control and the ability to reuse parameters across sessions, and optimizes the control effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122027840B_ABST
    Figure CN122027840B_ABST
Patent Text Reader

Abstract

This application relates to the field of digital human interaction technology, and discloses a method and system for intelligent interactive control of digital humans in private domain live streaming scenarios. The method includes: obtaining a group emotion change rate by performing multi-dimensional emotion recognition on bullet screen text, logarithmic normalization of the number of online users, and calculation of the dual-window ratio; triggering the digital human to deconstruct speech and action sequence output based on the group emotion change rate; and adaptively updating the interaction trigger threshold and the starting value of the long-term control window through cross-session interaction control files. This application improves the accuracy of group emotion perception and the cross-session reuse capability of control parameters in digital human interactive control in private domain live streaming scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital human interaction technology, and in particular to a digital human intelligent interaction control method and system in a private domain live streaming scenario. Background Technology

[0002] With the rapid development of private domain e-commerce live streaming, digital human technology is widely used in private domain live streaming scenarios. It uses technologies such as speech recognition, natural language processing, and computer vision to drive virtual digital human avatars to interact with users in the live stream in real time. Existing digital human live streaming interactive control systems typically adopt an edge computing and cloud-based collaborative architecture, using the WebRTC protocol to achieve real-time data transmission across multiple terminals. They employ skeletal rigging algorithms to drive the digital human model's movements in synchronization, and combine this with a speech recognition module to extract semantic features and generate corresponding responses to be output to the live stream.

[0003] However, existing digital human interaction control systems use single bullet screen text as control signal input, independently classify each bullet screen for emotion, output discrete emotion tags, and trigger one-to-one passive responses accordingly. This control structure has the following technical defects: First, discrete emotion tags lose continuous value information of emotion intensity, making it impossible for subsequent control modules to perform statistical processing of bullet screen emotion signals over time, resulting in the system's inability to perceive the relative changing trend of the group's emotions in the live broadcast room over time. Second, the number of online users in the private domain live broadcast room changes dynamically over time. Existing technologies directly accumulate and statistically analyze bullet screen emotion signals, resulting in higher accumulated values ​​during periods with more online users, making the emotional statistics of different periods and sessions incomparable in terms of scale, and preventing the interaction trigger threshold from being reused across sessions. Third, existing technologies do not quantitatively evaluate the control effect of the interaction after the digital human executes the interaction output. The interaction trigger threshold is a fixed preset value in each live broadcast, and the control parameters are not adaptively updated with historical interaction effects.

[0004] Because existing technology cannot perceive the relative changing trends of group emotions and can only respond to the absolute emotional intensity of a single bullet comment, when the emotions of the group in the live broadcast room change significantly in a short period of time, the system cannot identify and trigger proactive mitigation control in a timely manner. This results in a disconnect between the timing of the triggering of digital human interaction control commands and the actual changes in the emotional state of the group in the live broadcast room. Furthermore, because the dimensions of control signals are incomparable, even if a threshold triggering mechanism is introduced, the threshold cannot be optimized and updated based on historical session data due to the lack of cross-session comparability. On this basis, because the control effect lacks quantitative feedback, the system repeatedly goes through the same exploration process in each live broadcast and cannot use the historical user emotion data accumulated in the private domain live broadcast scenario to continuously optimize the control parameters. This prevents the user historical data assets unique to the private domain live broadcast scenario from playing a role in digital human interaction control. Summary of the Invention

[0005] This application provides a digital human intelligent interaction control method and system for private domain live streaming scenarios, which solves the problems in the prior art that digital human interaction control systems cannot trigger active resolution control based on the relative change trend of group emotions and that control parameters cannot be adaptively updated across sessions. It improves the accuracy of group emotion perception and the ability of control parameters to be reused across sessions in digital human interaction control in private domain live streaming scenarios.

[0006] Firstly, this application provides a digital human intelligent interaction control method in a private domain live streaming scenario, the digital human intelligent interaction control method in a private domain live streaming scenario includes:

[0007] Step S1: Perform sentiment recognition processing on each bullet screen text in the private domain live broadcast room bullet screen text stream to obtain the multi-dimensional sentiment intensity vector corresponding to each bullet screen text.

[0008] Step S2: Divide the multidimensional emotion intensity vector by the base-2 logarithm of the sum of the number of online users at the corresponding collection time and 2 to obtain the normalized emotion control quantity; accumulate the normalized emotion control quantity according to the collection time and write it into the short-term control window and the long-term control window respectively; divide the mean of the normalized emotion control quantity in each emotion dimension in the short-term control window by the sum of the mean of the normalized emotion control quantity in each emotion dimension in the long-term control window and the preset smoothing term to obtain the group emotion change rate corresponding to each emotion dimension;

[0009] Step S3: Compare the group emotion change rate corresponding to each emotion dimension with the interaction trigger threshold corresponding to each emotion dimension. When the group emotion change rate corresponding to any emotion dimension is not lower than the interaction trigger threshold corresponding to that emotion dimension, select a decryption script from the private domain script library according to the trigger dimension and the corresponding group emotion change rate. After the decryption script is synthesized by speech, drive the digital human to output an action sequence to the private domain live broadcast room and record the control output time.

[0010] Step S4: Within the verification time window starting from the control output time, obtain the interaction effect coefficient based on the change in the normalized emotional control quantity, write the group emotion change rate and the interaction effect coefficient into the private domain interaction control file, and update the interaction trigger threshold and the starting value of the long-term control window for the next session by the private domain interaction control file.

[0011] Secondly, this application provides a digital human intelligent interaction control system for private domain live streaming scenarios, the digital human intelligent interaction control system for private domain live streaming scenarios comprising:

[0012] The recognition module is used to perform sentiment recognition processing on each bullet screen text in the private domain live broadcast room bullet screen text stream to obtain a multi-dimensional sentiment intensity vector corresponding to each bullet screen text.

[0013] The analysis module is used to divide the multidimensional emotion intensity vector by the base-2 logarithm of the sum of the number of online users at the corresponding collection time and 2 to obtain the normalized emotion control quantity; the normalized emotion control quantity is accumulated and written into the short-term control window and the long-term control window according to the collection time; the mean of the normalized emotion control quantity in each emotion dimension in the short-term control window is divided by the sum of the mean of the normalized emotion control quantity in each emotion dimension in the long-term control window and the preset smoothing term to obtain the group emotion change rate corresponding to each emotion dimension;

[0014] The comparison module is used to compare the group emotion change rate corresponding to each emotion dimension with the interaction trigger threshold corresponding to each emotion dimension. When the group emotion change rate corresponding to any emotion dimension is not lower than the interaction trigger threshold corresponding to that emotion dimension, the module selects a decryption script from the private domain script library according to the trigger dimension and the corresponding group emotion change rate. The decryption script is then synthesized into speech to drive the digital human to output an action sequence to the private domain live broadcast room, and the control output time is recorded.

[0015] The update module is used to obtain the interaction effect coefficient based on the change in the normalized emotional control quantity within the verification time window starting from the control output time, write the group emotion change rate and the interaction effect coefficient into the private domain interaction control file, and update the interaction trigger threshold and the starting value of the long-term control window for the next session by the private domain interaction control file.

[0016] Thirdly, a digital human intelligent interaction control device for a private domain live streaming scenario is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the digital human intelligent interaction control device for the private domain live streaming scenario to execute the aforementioned digital human intelligent interaction control method for the private domain live streaming scenario.

[0017] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, cause the computer to execute the aforementioned digital human intelligent interaction control method in a private domain live streaming scenario.

[0018] In the technical solution provided in this application, a multi-dimensional emotion intensity vector is obtained by performing emotion recognition processing on each bullet screen text in the private domain live broadcast room. The multi-dimensional emotion intensity vector is divided by the logarithm of the sum of the number of online users at the corresponding collection time and 2, with the base 2, to obtain the normalized emotion control quantity. The normalized emotion control quantity is accumulated and written into the short-term control window and the long-term control window respectively. The group emotion change rate corresponding to each emotion dimension is obtained by dividing the mean in the short-term control window by the sum of the mean in the long-term control window and the preset smoothing term. In the above processing, the logarithmic normalization operation eliminates the interference of the fluctuation of the number of online users on the emotion statistics, so that the normalized emotion control quantity has dimensional comparability between different time periods and different sessions. The dual-window ratio structure transforms the control signal from the absolute emotion intensity of a single bullet screen into a continuous statistical quantity reflecting the relative change trend of group emotion. This enables the digital human interaction control system to have the ability to perceive the change trend of group emotion in the live broadcast room for the first time, so as to trigger active resolution control in a timely manner when the group emotion changes significantly, rather than waiting for a single bullet screen to trigger a passive reply.

[0019] This application further compares the group emotion change rate corresponding to each emotion dimension with the interaction trigger threshold corresponding to each emotion dimension dimension. Based on the trigger dimension and the corresponding group emotion change rate, it selects and cancels dialogue from the private domain dialogue library and drives the digital human to output the action sequence corresponding to the meaning of the canceled dialogue. This achieves a multi-dimensional and accurate mapping between control signals and digital human output behavior. Within the verification time window starting from the control output moment, the interaction effect coefficient is obtained based on the change in the normalized emotion control quantity. The group emotion change rate and the interaction effect coefficient are written into the private domain interaction control file, and the file updates the interaction trigger threshold and the starting value of the long-term control window for the next session. This transforms the control system from an open-loop structure to a closed-loop structure with cross-session feedback. The interaction trigger threshold gradually converges to the actual trigger level of that dimension in the private domain live streaming scenario as historical session data accumulates. The starting value of the long-term control window carries historical emotion baseline information, enabling the group emotion change rate to form an effective ratio judgment at the beginning of the live stream. The user's historical emotion data accumulated in the private domain live streaming scenario can continue to play a role in the iterative update of control parameters. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of an embodiment of the digital human intelligent interaction control method in a private domain live streaming scenario according to the present application.

[0022] Figure 2 This is a schematic diagram of an embodiment of the digital human intelligent interaction control system in a private domain live streaming scenario, as described in this application.

[0023] Figure 3 This is a schematic block diagram of the structure of a digital human intelligent interactive control device in a private domain live streaming scenario, as described in this embodiment of the invention. Detailed Implementation

[0024] This application provides a digital human intelligent interactive control method and system for private domain live streaming scenarios. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0025] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the digital human intelligent interaction control method in a private domain live streaming scenario in this application includes:

[0026] Step S1: Perform sentiment recognition processing on each bullet screen text in the private domain live broadcast room bullet screen text stream to obtain the multi-dimensional sentiment intensity vector corresponding to each bullet screen text.

[0027] It is understood that the executing entity of this application can be a digital human intelligent interactive control system in a private domain live streaming scenario, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiment uses a server as an example for illustration.

[0028] Specifically, the multidimensional sentiment intensity vector refers to the vector obtained after performing sentiment recognition on multiple sentiment dimensions of a single bullet screen text. Each dimension corresponds to a typical sentiment type in the private domain e-commerce live streaming scenario, including price concerns, quality doubts, purchase enthusiasm, urging for purchase, and negative resistance, totaling 5 dimensions, i.e., k=5. The value of each dimension is a continuous real number from 0 to 1, with a larger value indicating a higher activation level of that sentiment dimension. The components of the multidimensional sentiment intensity vector corresponding to each bullet screen text are independent of each other, and multiple dimensions of the same bullet screen text can simultaneously have high activation intensity to represent a complex sentiment state.

[0029] Step S2: Divide the multidimensional emotion intensity vector by the base-2 logarithm of the sum of the number of online users at the corresponding collection time and 2 to obtain the normalized emotion control quantity; accumulate the normalized emotion control quantity according to the collection time and write it into the short-term control window and the long-term control window respectively; divide the mean of the normalized emotion control quantity in each emotion dimension in the short-term control window by the sum of the mean of the normalized emotion control quantity in each emotion dimension in the long-term control window and the preset smoothing term to obtain the group emotion change rate corresponding to each emotion dimension;

[0030] Specifically, the normalized emotional control quantity refers to the vector obtained by dividing the multidimensional emotional intensity vector by the base-2 logarithmic value of the sum of the number of online users at the corresponding collection time and 2. The reason for choosing a base-2 logarithmic function in the denominator is that the number of online users in a private live broadcast room usually fluctuates between tens and thousands. The base-2 logarithmic function compresses this range to between 6 and 11, effectively eliminating the impact of differences in the number of users on the emotional control quantity, while avoiding the problem of the normalized value being too large due to the logarithmic value being too small. Adding 2 to the denominator is to ensure that the denominator is not zero when the number of online users is 0, and that the base-2 logarithmic value after adding 2 is not less than 1, preventing the normalized value from being greater than the original value. The short-term control window is set to 30 seconds, and the long-term control window is set to 300 seconds, with a ratio of 1:10. The short-term control window reflects the recent emotional intensity of the live stream at the current moment, while the long-term control window reflects the historical baseline emotional level of the live stream. The ratio of the two means, i.e., the group emotional change rate, reflects the magnitude of change in current emotion relative to the historical baseline. The preset smoothing term is set to 0.001 to prevent division by zero errors when the long-term control window mean is zero. This value is much smaller than the long-term dimension mean under normal circumstances and does not affect the calculation result of the group emotional change rate.

[0031] Step S3: Compare the group emotion change rate corresponding to each emotion dimension with the interaction trigger threshold corresponding to each emotion dimension. When the group emotion change rate corresponding to any emotion dimension is not lower than the interaction trigger threshold corresponding to that emotion dimension, select the decryption script from the private domain script library according to the trigger dimension and the corresponding group emotion change rate. After the decryption script is synthesized by speech, drive the digital human to output the action sequence to the private domain live broadcast room and record the control output time.

[0032] Specifically, the initial value of the interaction trigger threshold is set based on the historical trigger experience of each emotional dimension in the private domain e-commerce live streaming scenario. The preset initial value is used in the first live stream, and then it is updated by the private domain interaction control file in step S4 for each subsequent live stream. In step S4, the verification time window is set to 30 seconds, consistent with the short-term control window, to ensure that the verification and triggering phases use the same time scale for comparing emotional intensity. The interaction effect coefficient is calculated by subtracting the mean of the normalized emotional control quantity at the end of the verification time window from the mean of the short-term dimension at the triggering dimension and then dividing by the mean of the short-term dimension at the control output moment. This coefficient reflects the relative decrease in emotional intensity of the triggering dimension after the current dissolution of the speech output. A positive value indicates a decrease in emotional intensity, and a negative value indicates an increase in emotional intensity. The private domain interaction control archive records the group emotional change rate and interaction effect coefficient corresponding to each triggering event in sequence. The interaction trigger threshold for the next session is determined by the mean of the group emotional change rate in the archive of historical sessions. The starting value of the long-term control window is determined by the mean of the normalized emotional control quantity of private domain users in each emotional dimension in the archive of historical sessions, so that the long-term control window contains historical emotional baseline information at the start of the next live broadcast, rather than accumulating from scratch.

[0033] Step S4: Within the verification time window starting from the control output moment, obtain the interaction effect coefficient based on the change in the normalized emotional control quantity, write the group emotion change rate and the interaction effect coefficient into the private domain interaction control file, and update the interaction trigger threshold and the starting value of the long-term control window for the next session by the private domain interaction control file.

[0034] Specifically, the verification time window is set to 30 seconds, consistent with the short-term control window, to ensure that the verification and triggering phases use the same time scale for comparing emotional intensity. The interaction effect coefficient is calculated as follows: the mean of the normalized emotional control quantity within the short-term control window at the end of the verification time window is subtracted from the mean within the short-term control window at the control output time. This difference is then divided by the mean within the short-term control window at the control output time to obtain the interaction effect coefficient. This coefficient reflects the relative change in the normalized emotional control quantity of the triggering dimension after the dissipation of the dialogue output. A positive value indicates a decrease in the emotional intensity of the triggering dimension, while a negative value indicates that the emotional intensity of the triggering dimension has not been dissipated. The group emotional change rate and interaction effect coefficient corresponding to this triggering event are written into the private domain interaction control archive in session order. The private domain interaction control archive records the group emotional change rate and interaction effect coefficient for each session and each dimension, indexed by the emotional dimension. The interaction trigger threshold for the next session is determined by the average of the group emotion change rate corresponding to the emotion dimension in the historical sessions in the private domain interaction control archive. This allows the interaction trigger threshold to gradually converge to the actual trigger level of that dimension in the private domain live streaming scenario as historical session data accumulates. The starting value of the long-term control window is determined by the average of the normalized emotion control quantity of private domain users in the historical sessions in the private domain interaction control archive across each emotion dimension. This ensures that the long-term control window already contains historical emotion baseline information when the next live stream begins, rather than accumulating from scratch. This allows the group emotion change rate to form an effective ratio judgment at the beginning of the live stream.

[0035] In one specific embodiment, step S1 includes:

[0036] Each bullet screen text in the private live stream is input into a sentiment classification model with the Sigmoid function as the output layer activation function. Based on the output layer of the sentiment classification model, the activation intensity of each bullet screen text in each sentiment dimension is calculated to obtain the activation intensity value of each bullet screen text in each sentiment dimension.

[0037] Arrange the activation intensity values ​​of each emotion dimension corresponding to the same bullet screen text in the order of emotion dimension to obtain a multi-dimensional emotion intensity vector;

[0038] The multidimensional sentiment intensity vector corresponding to each bullet screen text is associated with the sending time of the corresponding bullet screen text and the private domain user identifier to obtain a multidimensional sentiment intensity vector sequence with time annotation.

[0039] The sequence of multidimensional emotion intensity vectors with time stamps is written into the bullet screen emotion buffer queue in the order of sending time. The bullet screen emotion buffer queue outputs the multidimensional emotion intensity vectors corresponding to each collection time and the corresponding private domain user identifiers to step S2.

[0040] Specifically, the output layer of the sentiment classification model uses the Sigmoid function as the activation function. The output value of the Sigmoid function is a continuous real number from 0 to 1. The output values ​​of each sentiment dimension are independent of each other and are not affected by the values ​​of other dimensions. This allows the same bullet screen text to have high activation intensity on multiple sentiment dimensions simultaneously, representing the complex emotional state of users expressing multiple emotions at the same time in private domain e-commerce live streaming scenarios. The activation intensity value of each sentiment dimension refers to the single real value output by the corresponding sentiment dimension neuron in the output layer of the sentiment classification model after being calculated by the Sigmoid function. The value ranges from 0 to 1, and the closer the value is to 1, the higher the activation degree of that sentiment dimension. Arranging the activation intensity values ​​of each sentiment dimension corresponding to the same bullet screen text in the order of sentiment dimension yields a multi-dimensional sentiment intensity vector. The number of dimensions of this vector is consistent with the number of neurons in the output layer of the sentiment classification model, i.e., k=5.

[0041] The multidimensional sentiment intensity vector sequence with time stamps refers to an ordered data sequence formed by performing a ternary association between the multidimensional sentiment intensity vector corresponding to each bullet screen text and the sending time of the bullet screen text and the private domain user identifier. The sending time is a Unix timestamp with millisecond precision; the private domain user identifier is a unique code assigned to registered users by the private domain platform. When the bullet screen sender is an unregistered user, the hash value of their device fingerprint is used as a temporary identifier, which is used in step S4 to associate the normalized sentiment control quantity with the private domain user's historical data and write it into the private domain interaction control file. The bullet screen sentiment buffer queue stores the above ternary association data in the order of sending time and continuously outputs the multidimensional sentiment intensity vector and the corresponding private domain user identifier corresponding to each collection time to step S2 in a first-in-first-out manner.

[0042] In one specific embodiment, in step S2, the multidimensional emotion intensity vector is divided by the base-2 logarithm of the sum of the number of online users at the corresponding collection time and 2 to obtain the normalized emotion control quantity, including:

[0043] The j-th dimension of the emotional intensity vector is denoted as the emotional component, and the number of online users at the corresponding collection time is denoted as the number of live online users. The logarithm of the sum of the number of live online users and 2 is taken to the base 2 to obtain the logarithm of the online scale.

[0044] Divide the emotional component by the logarithm of the online scale to obtain the j-th dimension normalized emotional component; arrange the normalized emotional components corresponding to each emotional dimension in order of emotional dimension to obtain the normalized emotional control quantity.

[0045] Specifically, the emotional component refers to the single real value corresponding to the j-th dimension in the multi-dimensional emotional intensity vector, where j ranges from 1 to k, and k is the total number of emotional dimensions. In this embodiment, k=5, and each emotional dimension corresponds to price doubt, quality question, purchase enthusiasm, urging for purchase, and negative resistance, respectively. The number of live stream viewers refers to the real-time number of online users in the private domain live stream at the corresponding collection time, collected by the live stream platform interface at 1-second intervals. The logarithmic value of the online scale refers to the real value obtained by taking the logarithm of the sum of the number of live stream viewers and 2, with the base 2. The reason for choosing the logarithmic function with the base 2 in the denominator is that the number of online users in the private domain live stream usually fluctuates between tens and thousands. The logarithmic function with the base 2 compresses this value range to between 6 and 12, eliminating the influence of the difference in the number of users on the emotional control quantity while avoiding the problem of abnormally amplified value after normalization due to the logarithmic value being too small. Adding 2 to the denominator is to ensure that the denominator is not zero when the number of online users is 0, and that the logarithmic value with the base 2 after adding 2 is not less than 1, preventing the normalized emotional component from being greater than the original emotional component.

[0046] The j-th normalized sentiment component refers to the real value obtained by dividing the sentiment component by the logarithmic value of the online scale. This division operation is performed independently for each dimension, and there is no mutual influence between the normalized sentiment components of each dimension. The normalized sentiment components corresponding to the k sentiment dimensions are arranged in order of sentiment dimension to obtain a normalized sentiment control quantity with the same number of dimensions as the multi-dimensional sentiment intensity vector. The j-th component of the normalized sentiment control quantity is the j-th normalized sentiment component, and this vector serves as the data unit subsequently written into the short-term control window and the long-term control window.

[0047] In one specific embodiment, step S2 involves accumulating and writing the normalized emotion control quantity into the short-term control window and the long-term control window according to the collection time, including:

[0048] The normalized sentiment control quantity is written into the short-time control window in the order of collection time. The duration of the short-time control window is 30 seconds. The j-th dimension normalized sentiment component of all normalized sentiment control quantities in the short-time control window is calculated by exponential weighted moving average according to the collection time to obtain the short-time dimension mean.

[0049] The normalized sentiment control values ​​are written into a long-term control window in the order of collection time. The long-term control window has a duration of 300 seconds. The j-th dimension normalized sentiment component of all normalized sentiment control values ​​in the long-term control window is calculated by an exponentially weighted moving average according to the collection time to obtain the long-term dimension mean.

[0050] Specifically, the short-term control window refers to a sliding time window extending 30 seconds backward from the current acquisition time as the right endpoint, while the long-term control window refers to a sliding time window extending 300 seconds backward from the current acquisition time as the right endpoint. Both slide and update synchronously as the acquisition time progresses. The short-term control window is set to 30 seconds because short-term fluctuations in user emotions in private domain live streaming scenarios are usually concentrated within 30 seconds, and this time length can capture recent changes in the group's emotions in the live streaming room at the current moment. The long-term control window is set to 300 seconds because 300 seconds can cover the emotional baseline level of a complete topic cycle in a private domain live streaming scenario. The ratio of the two time lengths is 1:10, which allows the rate of change in group emotions to effectively reflect the magnitude of change in current emotions relative to the historical baseline.

[0051] The short-term mean refers to the real value obtained by calculating the exponentially weighted moving average of the j-th dimension normalized sentiment components of all normalized sentiment control quantities within the short-term control window, over the time of data collection. The long-term mean refers to the real value obtained by calculating the exponentially weighted moving average of the j-th dimension normalized sentiment components of all normalized sentiment control quantities within the long-term control window, over the time of data collection. The recursive calculation method of the exponentially weighted moving average is as follows: multiply the j-th dimension normalized sentiment component at the current time by a decay coefficient, and add the product of the mean at the previous time point multiplied by 1 minus the decay coefficient to obtain the mean at the current time point. The decay coefficient corresponding to the short-term control window is 0.95, and the decay coefficient corresponding to the long-term control window is 0.99. The smaller decay coefficient of the short-term control window gives higher weight to recent data, while the larger decay coefficient of the long-term control window retains historical data for a longer period of influence. The difference in the decay coefficients makes the short-term mean more sensitive to current sentiment changes, while the long-term mean is more stable to the historical sentiment baseline.

[0052] In one specific embodiment, step S2 involves dividing the mean of the normalized emotional control quantity across each emotional dimension within the short-term control window by the sum of the mean of the normalized emotional control quantity across each emotional dimension within the long-term control window and a preset smoothing term to obtain the group emotional change rate corresponding to each emotional dimension, including:

[0053] Divide the short-term mean by the sum of the long-term mean and the preset smoothing term to obtain the j-th dimension group sentiment change rate; where the preset smoothing term is 0.001, used to prevent division by zero error when the long-term mean is zero;

[0054] Perform the corresponding division operation on each emotional dimension to obtain the group emotional change rate corresponding to each emotional dimension.

[0055] Specifically, the group sentiment change rate refers to the real value obtained by dividing the short-term dimension mean by the sum of the long-term dimension mean and a preset smoothing term. This calculation is performed independently for each sentiment dimension. The group sentiment change rate of the j-th dimension is calculated only by the j-th dimension short-term mean and the j-th dimension long-term mean. The calculation of the group sentiment change rate between different sentiment dimensions does not affect each other. The physical meaning of the group sentiment change rate is the multiple relationship between the current j-th dimension sentiment intensity and the historical baseline level. A value of 1 indicates that the current sentiment intensity is the same as the historical baseline, a value greater than 1 indicates that the current sentiment intensity is higher than the historical baseline, and a larger value indicates a greater increase in the current sentiment intensity relative to the historical baseline.

[0056] The preset smoothing term is set to 0.001. This value is based on the fact that the long-term mean of a private live stream is usually not lower than 0.01 under normal circumstances. The preset smoothing term value of 0.001 is much smaller than the long-term mean under normal circumstances, so its impact on the calculation result of the group emotion change rate is negligible. Only when the long-term mean approaches zero does the preset smoothing term prevent division by zero errors, ensuring that the calculation of the group emotion change rate can still be performed normally when the long-term control window data accumulation is insufficient at the beginning of the live stream. After performing the above division operation on each emotion dimension, the group emotion change rate corresponding to each of the k emotion dimensions is obtained, k=5. The group emotion change rates corresponding to each dimension together constitute the input data for the dimension-by-dimensional comparison with the interaction trigger thresholds corresponding to each emotion dimension in step S3.

[0057] In one specific embodiment, in step S3, decryption scripts are selected from the private domain script library based on the trigger dimension and the corresponding group emotion change rate. The decryption scripts are then synthesized into speech to drive the digital human to output an action sequence to the private domain live streaming room, including:

[0058] The trigger dimension and the corresponding group emotion change rate are input into the script selection rules in the private domain script library. The script selection rules divide the group emotion change rate into a first value interval, a second value interval, and a third value interval. Based on the trigger dimension and the value interval to which the group emotion change rate belongs, the corresponding dissolving script is matched from the private domain script library.

[0059] The deconstructed speech is input into the speech synthesis module to obtain the deconstructed speech audio; the deconstructed speech audio is input into the digital human skeleton binding model, and the digital human skeleton binding model outputs an action sequence and lip-sync video stream that are aligned with the audio duration of the deconstructed speech. The action sequence and lip-sync video stream are then pushed to the private live streaming room to record the control output time.

[0060] Specifically, the private domain dialogue database refers to a dialogue database that is hierarchically indexed according to emotional dimension categories and the range of group emotional change rates. Matching and retrieval are performed using the category of the trigger dimension and the corresponding range of group emotional change rates as the index key. The dialogue selection rules divide the group emotional change rate into three ranges: the first range is between the interaction trigger threshold and twice the interaction trigger threshold, corresponding to a mild emotional buildup state, matching low-intensity de-escalation dialogue; the second range is between twice and three times the interaction trigger threshold, corresponding to a moderate emotional buildup state, matching medium-intensity de-escalation dialogue; and the third range is three times the interaction trigger threshold and above, corresponding to a severe emotional buildup state, matching high-intensity de-escalation dialogue. Multiple trigger dimensions can exist simultaneously. When multiple emotional dimensions are triggered at the same time, the corresponding de-escalation dialogue is matched sequentially in descending order of group emotional change rate, and multiple de-escalation dialogues are merged into a complete de-escalation dialogue text.

[0061] The decrypted speech audio refers to the audio data obtained after inputting the decrypted speech text into the speech synthesis module. The audio format is a PCM-encoded mono audio stream with a sampling rate of 16000Hz. The digital human skeleton binding model is a generation model that uses the decrypted speech audio as the driving signal and outputs a digital human action sequence and a lip-sync video stream. This model takes the temporal characteristics of the decrypted speech audio as input, calculates the joint angles and lip deformation parameters of the digital human skeleton frame by frame, and outputs an action sequence and a lip-sync video stream with a frame rate of 25 frames per second. The total number of frames in the action sequence is strictly aligned with the duration of the decrypted speech audio. The alignment method is to divide the total audio duration by the single frame duration of 40 milliseconds and round down to obtain the total number of frames, ensuring that the action sequence ends synchronously with the decrypted speech audio. The control output time is recorded as a Unix timestamp with millisecond precision when the action sequence and lip-sync video stream begin to be pushed to the private domain live broadcast room. This timestamp is used to locate the start time of the verification time window in step S4.

[0062] In one specific embodiment, in step S4, within the verification time window starting from the control output moment, the interaction effect coefficient is obtained based on the change in the normalized emotion control quantity. The group emotion change rate and the interaction effect coefficient are written into the private domain interaction control file. The private domain interaction control file updates the interaction trigger threshold for the next session and the starting value of the long-term control window, including:

[0063] The interaction effect coefficient is obtained by subtracting the short-term mean of the normalized emotional control quantity at the end of the verification time window from the short-term mean of the control output time, and dividing the difference by the short-term mean of the control output time.

[0064] Write the group emotion change rate and interaction effect coefficient corresponding to the trigger dimension into the private domain interaction control file in the order of the sessions.

[0065] Based on the mean of the group emotion change rate corresponding to the trigger dimension in the historical sessions in the private domain interaction control archive, update the interaction trigger threshold of the trigger dimension corresponding to the next session; based on the mean of the normalized emotion control quantity of private domain users in the historical sessions in the private domain interaction control archive on each emotion dimension, initialize the starting value of the long-term control window for the next session.

[0066] Specifically, the verification time window refers to a fixed time window that starts at the control output time and extends for 30 seconds. The time length is set to 30 seconds because it is consistent with the short-term control window, ensuring that the verification and triggering phases use the same time scale to statistically analyze the normalized emotional control quantity, thus guaranteeing the comparability of the interaction effect coefficient calculation. The interaction effect coefficient is a real value obtained by subtracting the short-term mean of the normalized emotional control quantity at the end of the verification time window from the short-term mean at the control output time, and then dividing by the short-term mean at the control output time. This coefficient reflects the relative change in the normalized emotional control quantity of the triggering dimension after the output of this dissipation speech. A negative value indicates that the short-term mean of the triggering dimension decreases within the verification time window, while a positive value indicates that the short-term mean of the triggering dimension increases within the verification time window. The more negative the value, the greater the dissipation effect of this dissipation speech on the emotional backlog in the triggering dimension.

[0067] The private domain interaction control archive uses the emotion dimension as the primary index and the session number as the secondary index. It records the group emotion change rate and interaction effect coefficient corresponding to each trigger dimension in each session, and appends them in the order of each session after the session ends. The interaction trigger threshold for the corresponding trigger dimension of the next session is determined by the arithmetic mean of the group emotion change rate of all historical sessions of that trigger dimension in the private domain interaction control archive. When the number of historical sessions is less than 3, a preset initial value is used. The starting value of the long-term control window for the next session is determined by the arithmetic mean of the normalized emotion control quantity of all private domain users' historical sessions in the private domain interaction control archive on each emotion dimension. Each emotion dimension is calculated independently to obtain a k-dimensional initial vector as the initial accumulated value of the long-term control window, so that the long-term control window already contains historical emotion baseline information when the next live broadcast starts, rather than accumulating from zero.

[0068] The above describes the digital human intelligent interaction control method in a private domain live streaming scenario in the embodiments of this application. The following describes the digital human intelligent interaction control system in a private domain live streaming scenario in the embodiments of this application. Please refer to [link / reference]. Figure 2 One embodiment of the digital human intelligent interaction control system in a private domain live streaming scenario in this application includes:

[0069] The recognition module is used to perform sentiment recognition processing on each bullet screen text in the private domain live broadcast room bullet screen text stream to obtain a multi-dimensional sentiment intensity vector corresponding to each bullet screen text.

[0070] The analysis module is used to divide the multidimensional emotion intensity vector by the base-2 logarithm of the sum of the number of online users at the corresponding collection time and 2 to obtain the normalized emotion control quantity; the normalized emotion control quantity is accumulated and written into the short-term control window and the long-term control window according to the collection time; the mean of the normalized emotion control quantity in each emotion dimension in the short-term control window is divided by the sum of the mean of the normalized emotion control quantity in each emotion dimension in the long-term control window and the preset smoothing term to obtain the group emotion change rate corresponding to each emotion dimension;

[0071] The comparison module is used to compare the group emotion change rate corresponding to each emotion dimension with the interaction trigger threshold corresponding to each emotion dimension. When the group emotion change rate corresponding to any emotion dimension is not lower than the interaction trigger threshold corresponding to that emotion dimension, the module selects a decryption script from the private domain script library according to the trigger dimension and the corresponding group emotion change rate. The decryption script is then synthesized into speech to drive the digital human to output an action sequence to the private domain live broadcast room, and the control output time is recorded.

[0072] The update module is used to obtain the interaction effect coefficient based on the change in the normalized emotional control quantity within the verification time window starting from the control output time, write the group emotion change rate and the interaction effect coefficient into the private domain interaction control file, and update the interaction trigger threshold and the starting value of the long-term control window for the next session by the private domain interaction control file.

[0073] above Figure 2 The digital human intelligent interaction control system in the private domain live streaming scenario of this invention will be described in detail from the perspective of modular functional entities. The digital human intelligent interaction control device in the private domain live streaming scenario of this invention will be described in detail from the perspective of hardware processing.

[0074] Reference Figure 3 This invention also provides a digital human intelligent interaction control device for private domain live streaming scenarios. This device can be a server, and its internal structure can be as follows: Figure 3As shown, the digital human intelligent interactive control device in this private domain live streaming scenario includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor in this computer design provides computing and control capabilities. The memory of the digital human intelligent interactive control device in this private domain live streaming scenario includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the digital human intelligent interactive control device in this private domain live streaming scenario is used to store the data corresponding to this embodiment. The network interface of the digital human intelligent interactive control device in this private domain live streaming scenario is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.

[0075] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the digital human intelligent interactive control device in the private domain live streaming scenario to which the present invention is applied.

[0076] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the digital human intelligent interaction control method in the private domain live streaming scenario.

[0077] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0078] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a digital human intelligent interactive control device (which may be a personal computer, server, or network device, etc.) in a private domain live streaming scenario to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0079] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent interactive control of digital humans in a private domain live streaming scenario, characterized in that, The method includes: Step S1: Perform sentiment recognition processing on each bullet screen text in the private domain live broadcast room bullet screen text stream to obtain the multi-dimensional sentiment intensity vector corresponding to each bullet screen text. Step S2: Divide the multidimensional emotion intensity vector by the base-2 logarithm of the sum of the number of online users at the corresponding collection time and 2 to obtain the normalized emotion control quantity. The normalized sentiment control quantities are accumulated and written into a short-term control window and a long-term control window according to the collection time, respectively. This includes: writing the normalized sentiment control quantities into a short-term control window in order of collection time, wherein the duration of the short-term control window is 30 seconds; calculating an exponentially weighted moving average of all the normalized sentiment control quantities in the j-th dimension of the short-term control window according to the collection time to obtain the short-term dimension mean; and writing the normalized sentiment control quantities into a long-term control window in order of collection time, wherein the duration of the long-term control window is 300 seconds; calculating an exponentially weighted moving average of all the normalized sentiment control quantities in the j-th dimension of the long-term control window according to the collection time to obtain the long-term dimension mean. Divide the mean of the normalized emotional control quantity in each emotional dimension within the short-term control window by the sum of the mean of the normalized emotional control quantity in each emotional dimension within the long-term control window and the preset smoothing term to obtain the group emotional change rate corresponding to each emotional dimension. Step S3: Compare the group emotion change rate corresponding to each emotion dimension with the interaction trigger threshold corresponding to each emotion dimension. When the group emotion change rate corresponding to any emotion dimension is not lower than the interaction trigger threshold corresponding to that emotion dimension, select a decryption script from the private domain script library according to the trigger dimension and the corresponding group emotion change rate. After the decryption script is synthesized by speech, drive the digital human to output an action sequence to the private domain live broadcast room and record the control output time. Step S4: Within the verification time window starting from the control output time, obtain the interaction effect coefficient based on the change in the normalized emotional control quantity, write the group emotion change rate and the interaction effect coefficient into the private domain interaction control file, and update the interaction trigger threshold and the starting value of the long-term control window for the next session by the private domain interaction control file.

2. The digital human intelligent interaction control method in the private domain live broadcast scene according to claim 1, characterized in that, Step S1 includes: Each bullet screen text in the private live stream is input into a sentiment classification model with the Sigmoid function as the output layer activation function. Based on the output layer of the sentiment classification model, the activation intensity of each bullet screen text in each sentiment dimension is calculated to obtain the activation intensity value of each bullet screen text in each sentiment dimension. The activation intensity values ​​of each emotion dimension corresponding to the same bullet screen text are arranged in order of emotion dimension to obtain the multi-dimensional emotion intensity vector; The multidimensional emotional intensity vector corresponding to each bullet screen text is associated with the sending time of the corresponding bullet screen text and the private domain user identifier to obtain a multidimensional emotional intensity vector sequence with time annotation. The sequence of multidimensional emotion intensity vectors with time stamps is written into the bullet screen emotion buffer queue in the order of sending time. The bullet screen emotion buffer queue outputs the multidimensional emotion intensity vectors corresponding to each collection time and the corresponding private domain user identifiers to step S2.

3. The digital human intelligent interaction control method in the private domain live broadcast scene according to claim 1, characterized in that, In step S2, the multidimensional emotion intensity vector is divided by the base-2 logarithm of the sum of the number of online users at the corresponding collection time and 2 to obtain the normalized emotion control quantity, including: The j-th dimension of the multidimensional emotional intensity vector is denoted as the emotional component, and the number of online users at the corresponding collection time is denoted as the number of live online users. The logarithm of the sum of the number of live online users and 2 is taken to the base 2 to obtain the logarithm of the online scale. Divide the emotional component by the logarithmic value of the online scale to obtain the j-th dimension normalized emotional component; arrange the normalized emotional components corresponding to each emotional dimension in order of emotional dimension to obtain the normalized emotional control quantity.

4. The digital human intelligent interaction control method in the private domain live broadcast scene according to claim 3, characterized in that, In step S2, the mean of the normalized emotional control quantity in each emotional dimension within the short-term control window is divided by the sum of the mean of the normalized emotional control quantity in each emotional dimension within the long-term control window and a preset smoothing term to obtain the group emotional change rate corresponding to each emotional dimension, including: Divide the short-term mean by the sum of the long-term mean and the preset smoothing term to obtain the j-th dimension group sentiment change rate; wherein, the preset smoothing term is 0.001, which is used to prevent division by zero error when the long-term mean is zero; Perform the corresponding division operation on each emotional dimension to obtain the group emotional change rate corresponding to each emotional dimension.

5. The digital human intelligent interaction control method in the private domain live broadcast scene according to claim 1, characterized in that, In step S3, decryption scripts are selected from the private domain script library based on the trigger dimension and the corresponding group emotion change rate. These decryption scripts are then synthesized into speech to drive a digital human to output an action sequence to the private domain live streaming room, including: The trigger dimension and the corresponding group emotion change rate are input into the script selection rules in the private domain script library. The script selection rules divide the group emotion change rate into a first value interval, a second value interval, and a third value interval. Based on the trigger dimension and the value interval to which the group emotion change rate belongs, the corresponding dissolving script is matched from the private domain script library. The deconstructed speech is input into the speech synthesis module to obtain the deconstructed speech audio; the deconstructed speech audio is input into the digital human skeleton binding model, and the digital human skeleton binding model outputs an action sequence and lip-sync video stream that are aligned with the audio duration of the deconstructed speech; the action sequence and lip-sync video stream are pushed to the private live broadcast room, and the control output time is recorded.

6. The digital human intelligent interaction control method in the private domain live broadcast scene according to claim 1, characterized in that, In step S4, within the verification time window starting from the control output moment, the interaction effect coefficient is obtained based on the change in the normalized emotional control quantity. The group emotion change rate and the interaction effect coefficient are written into the private domain interaction control file. The private domain interaction control file updates the interaction trigger threshold for the next session and the starting value of the long-term control window, including: The interaction effect coefficient is obtained by subtracting the short-term mean of the normalized emotional control quantity in the trigger dimension at the end of the verification time window from the short-term mean of the control output time, and dividing the difference by the short-term mean of the control output time. The group emotion change rate and the interaction effect coefficient corresponding to the trigger dimension are written into the private domain interaction control file in the order of the sessions. Based on the mean of the group emotion change rate corresponding to the trigger dimension in the historical sessions of the private domain interaction control archive, update the interaction trigger threshold of the trigger dimension corresponding to the next session; based on the mean of the normalized emotion control quantity of the private domain users in the historical sessions of the private domain interaction control archive on each emotion dimension, initialize the starting value of the long-term control window for the next session.

7. A digital human intelligent interaction control system in a private domain live broadcast scene, characterized in that, The method for implementing the intelligent interactive control of digital humans in a private domain live streaming scenario as described in any one of claims 1-6, wherein the intelligent interactive control system of digital humans in the private domain live streaming scenario comprises: The recognition module is used to perform sentiment recognition processing on each bullet screen text in the private domain live broadcast room bullet screen text stream to obtain a multi-dimensional sentiment intensity vector corresponding to each bullet screen text. The analysis module is used to divide the multidimensional emotion intensity vector by the base-2 logarithm of the sum of the number of online users at the corresponding collection time and 2, to obtain the normalized emotion control quantity. The normalized sentiment control quantities are accumulated and written into a short-term control window and a long-term control window according to the collection time, respectively. This includes: writing the normalized sentiment control quantities into a short-term control window in order of collection time, wherein the duration of the short-term control window is 30 seconds; calculating an exponentially weighted moving average of all the normalized sentiment control quantities in the j-th dimension of the short-term control window according to the collection time to obtain the short-term dimension mean; and writing the normalized sentiment control quantities into a long-term control window in order of collection time, wherein the duration of the long-term control window is 300 seconds; calculating an exponentially weighted moving average of all the normalized sentiment control quantities in the j-th dimension of the long-term control window according to the collection time to obtain the long-term dimension mean. Divide the mean of the normalized emotional control quantity in each emotional dimension within the short-term control window by the sum of the mean of the normalized emotional control quantity in each emotional dimension within the long-term control window and the preset smoothing term to obtain the group emotional change rate corresponding to each emotional dimension. The comparison module is used to compare the group emotion change rate corresponding to each emotion dimension with the interaction trigger threshold corresponding to each emotion dimension. When the group emotion change rate corresponding to any emotion dimension is not lower than the interaction trigger threshold corresponding to that emotion dimension, the module selects a decryption script from the private domain script library according to the trigger dimension and the corresponding group emotion change rate. The decryption script is then synthesized into speech to drive the digital human to output an action sequence to the private domain live broadcast room, and the control output time is recorded. The update module is used to obtain the interaction effect coefficient based on the change in the normalized emotional control quantity within the verification time window starting from the control output time, write the group emotion change rate and the interaction effect coefficient into the private domain interaction control file, and update the interaction trigger threshold and the starting value of the long-term control window for the next session by the private domain interaction control file.

8. A digital human intelligent interaction control device in a private domain live broadcast scene, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the digital human intelligent interaction control method in the private domain live streaming scenario as described in any one of claims 1 to 6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is run by the processor, it causes the processor to execute the digital human intelligent interaction control method in the private domain live streaming scenario as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice data emotion detecting method, device and system

    CN106782615A

  • Live broadcast verbal skill generation method and system based on AI digital human

    CN120542427A