Deep learning digital human live broadcast method and system

By generating a set of audience interaction timing, interaction intensity, emotional interaction fusion, and contextual changes, the problem of lagging facial expressions and movements in virtual digital human live streaming was solved, achieving real-time consistency between virtual digital human and audience interaction and improving the instant response and continuity of live streaming.

CN120915973APending Publication Date: 2025-11-07WEIHAI MENGHU NETWORK TECHNOLOGY CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511199676.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies lack the ability to perform timeline correlation analysis of audience behavior, emotional changes, and contextual data in virtual digital human live streaming, resulting in delayed updates of facial expressions, voice, and actions, which reduces the live stream's real-time responsiveness and continuity.

Method used

By acquiring the frame rate of virtual digital human face capture, keywords in audience comments and bullet screens, and heat values ​​of live screen areas, a set of audience interaction time sequences is generated. This is combined with likes and shares to generate a set of interaction intensity. The emotional interaction fusion set is calculated using the number and weight of comment emotions. Finally, this is fused with the rate of change in online users and the rate of change in topic popularity to form a set of contextual changes, thus achieving the correlation operation and feature fusion of multi-dimensional data on the same time line.

Benefits of technology

It achieves real-time consistency between the virtual digital human's live performance and the audience's emotional and situational changes, improving the live broadcast's instant responsiveness and continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915973A_ABST
    Figure CN120915973A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a deep learning digital human live broadcast method and system, and the method comprises the steps: obtaining a face capture frame rate, comment bullet screen keywords and a picture thermodynamic value, carrying out the matching to generate an audience interaction time sequence set, carrying out the statistics of like and sharing to generate a real-time interaction intensity set, and carrying out the real-time interaction; calculating and generating an emotional interaction fusion set by combining the emotional word quantity, generating a live broadcast situation change set by combining the population change rate and the topic popularity, and forming a live broadcast performance execution set by combining the interest hotspots. In virtual digital human live broadcast, a face capture frame rate is compared with a picture thermodynamic value and matched with a comment keyword to generate an interaction time sequence set, an interaction strength set is generated by combining like giving and sharing, an emotion interaction fusion set is generated by utilizing an emotion quantity and a weight, and a situation change set is generated by fusing a people number change rate and topic popularity. And outputting an expression voice action control execution set in combination with the interest hotspots.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a digital person live broadcast method and system based on deep learning. BACKGROUND

[0002] The field of artificial intelligence technology includes the intersection of computer science and artificial intelligence, involving how to use machines to simulate human intelligent activities, aiming to realize intelligent decision-making and autonomous learning capabilities. The core content of artificial intelligence technology includes deep learning, natural language processing, computer vision, speech recognition, etc., mainly focusing on simulating and enhancing human intelligence through data analysis, pattern recognition, and autonomous decision-making systems. In this field, the widespread application of deep learning technologies, especially convolutional neural networks (CNN), recurrent neural networks (RNN), and other models, enables artificial intelligence to handle more complex and diverse data, and has made significant progress in image recognition, speech recognition, and natural language understanding. The development of artificial intelligence technology not only promotes the automation and intelligence of traditional industries, but also provides strong technical support for new industrial innovation.

[0003] Among them, the digital person live broadcast method and system based on deep learning refers to a technical method and system that generates virtual digital persons using deep learning technology and conducts real-time live broadcast. For the generation of virtual characters, real-time interaction, emotional expression, and other technical matters in digital person live broadcast, a solution is proposed to automatically generate digital person images with realistic and interactive features through deep learning technology. Specifically, deep learning algorithms are used to analyze and synthesize the appearance, movements, and language expressions of virtual digital persons, enabling them to interact with the audience in real time. In addition, it also involves the implementation of emotional recognition and feedback mechanisms in digital person live broadcast, allowing virtual digital persons to adjust their behavior based on audience feedback through deep learning technology, thereby enhancing the audience's interactive experience.

[0004] The existing technology often separates the processing of audience behavior, emotional changes, and situational data in virtual digital person live broadcast, lacks correlation analysis of data within the same timeline, and makes it difficult for emotions and interactive data to form a unified performance drive. When the audience's activity level or emotional state fluctuates rapidly, the virtual digital person's expressions, voices, and movements update with a lag, causing the performance content to be disconnected from the audience's current experience, reducing the live broadcast's immediate response capability and continuity. SUMMARY

[0005] To solve the technical problems existing in the prior art, the embodiments of the present application provide a digital person live broadcast method and system based on deep learning. The technical solution is as follows: A digital person live broadcast method based on deep learning, comprising the following steps: S1: Obtain a virtual digital human face capture frame rate, audience comment barrage keywords and live picture region heat value, correspond and compare the virtual digital human face capture frame rate and the live picture region heat value according to a time line, match and combine the comparison result and the audience comment barrage keywords according to the same time point, and generate an audience interaction time sequence set; S2: Statistically add the number of likes and the number of shares in the audience interaction time sequence set, and calculate the ratio value with the total number of audiences in the same time period, vector combine the calculation result and the number of comments, and generate a real-time interaction intensity set; S3: According to the real-time interaction intensity set, extract the number of positive and negative emotion words in the comment barrage keywords, multiply the number of two categories by the corresponding emotion weight respectively and sum, calculate the difference value of the two sum values and perform ratio operation with the total number of comments, combine the operation result with the real-time interaction intensity set, and generate a sentiment interaction fusion set; S4: Vector splicing operation is performed on the sentiment interaction fusion set and the online audience change rate in the same time line, the change amplitude of the splicing vector is calculated, and the numerical value is superimposed with the topic heat change rate, to generate a live situation change set.

[0006] As a further scheme of the application, the audience interaction time sequence set includes a picture attention track distribution, a comment semantic category and a heat concentration block, the real-time interaction intensity set includes an interaction frequency vector, an active audience proportion and a comment density interval, the sentiment interaction fusion set includes an emotion change trend value, an emotion proportion structure and an emotion interaction correlation feature, and the live situation change set includes a situation change amplitude sequence, a hot topic switching record and an audience gathering and scattering fluctuation mode.

[0007] As a further scheme of the application, the audience interaction time sequence set includes a picture attention track distribution, a comment semantic category and a heat concentration block, the real-time interaction intensity set includes an interaction frequency vector, an active audience proportion and a comment density interval, the sentiment interaction fusion set includes an emotion change trend value, an emotion proportion structure and an emotion interaction correlation feature, and the live situation change set includes a situation change amplitude sequence, a hot topic switching record and an audience gathering and scattering fluctuation mode. S101: Obtain a frame rate data sequence output by a virtual digital human face capture device and a heat value sequence collected by a live picture region heat sensor, align the two sequences according to data time stamps, ensure one-to-one correspondence at each time point, calculate the absolute difference value of the face capture frame rate value and the region heat value at each time point, integrate the difference values of all time points into a sequence, and generate a face capture heat difference sequence; S102: Obtain the audience comment barrage data stream, the data containing a time stamp and text content, analyze the barrage text for each time point, extract the keywords in the preset keyword list, count the number of keywords appearing at each time point, take the counting result as a sequence element, and generate a barrage keyword heat sequence; S103: Align the face-catching heat force difference sequence and the barrage keyword heat sequence according to the same timestamp, combine the face-catching heat force difference sequence value and the barrage keyword heat sequence value in each matching time point into a single data unit, and establish a viewer interaction time sequence set.

[0008] As a further scheme of the present application, the acquisition step of the real-time interaction intensity set is: S201: Acquire the like number and share number data in the viewer interaction time sequence set, statistically summarize the like number and the share number, perform an accumulation and merging operation on the like number statistical result and the share number statistical result, integrate the two types of interactive behavior data, and generate interactive behavior summary data; S202: Call the interactive behavior summary data and the total number of viewers in the same time period, perform a ratio calculation process on the interactive behavior summary data and the total number of viewers, perform a correlation analysis operation on the two data, establish a proportion relationship between the interactive behavior and the audience size, and obtain interactive participation ratio data; S203: Based on the interactive participation ratio data and the comment quantity data, integrate the interactive participation ratio data and the comment quantity according to a vector combination rule, perform a vectorization assembly process on the two data, form a multi-dimensional interactive data structure, and establish a real-time interaction intensity set.

[0009] As a further scheme of the present application, the acquisition step of the emotional interaction fusion set is: S301: Extract positive emotional words and negative emotional words from the comment barrage keywords according to the real-time interaction intensity set, perform polarity recognition and classification processing on the emotional words, classify the emotional words into different categories according to the positive and negative attributes, count the occurrence frequency of the emotional words in each category, and establish an emotional word polarity distribution table; S302: Call the emotional word polarity distribution table and the corresponding emotional weight data, respectively perform a multiplication operation on the positive emotional word frequency and the negative emotional word frequency and the corresponding weight, perform a summation process on the positive weighted result and the negative weighted result, calculate the difference value of the two summation results, perform a ratio operation on the difference value and the total number of comments, and obtain an emotional tendency index; S303: Based on the emotional tendency index and the real-time interaction intensity set, perform a data combination operation on the emotional tendency index and the real-time interaction intensity set, perform a fusion and integration process on the two types of data, construct a composite data structure containing interaction intensity and emotional tendency, and establish an emotional interaction fusion set.

[0010] As a further scheme of the present application, the acquisition step of the live broadcast situation change set is: S401: The emotional interaction fusion set is vector spliced with the online number change rate in the same time line, the two types of data are matched and corresponded according to time nodes, a vector splicing operation is performed to construct a composite vector structure, the emotional interaction data and the number change data are integrated, and a composite vector data is established; S402: The composite vector data is called to perform change amplitude measurement processing, topic heat change rate data are obtained, the vector change amplitude and the topic heat change rate are superimposed and integrated, the superimposed result is reordered in time sequence, a time sequence arranged change data structure is constructed, and time sequence change data table is obtained; S403: Based on the time sequence change data table, change amplitude screening processing is performed, the data change amplitude is compared with the set limit threshold value, data points exceeding the limit threshold value are screened out, data points meeting the screening condition are summarized and arranged, a data set containing abnormal change characteristics is constructed, and a live situation change set is generated.

[0011] As a further scheme of the application, the method further comprises: S5: Based on the live situation change set, group interest hot spot weights are extracted, the hot spot weights are multiplied by the numerical values in the corresponding positions of the live situation change set and then summed, the obtained numerical value is proportionally combined with the interest hot spot appearance frequency, and a live performance execution set directly controlling the expression change, voice rhythm and body action amplitude of the virtual digital person in the live broadcast is formed; The live performance execution set comprises expression change control signals, voice rhythm adjustment signals and body action amplitude control signals.

[0012] As a further scheme of the application, the live performance execution set comprises expression change control signals, voice rhythm adjustment signals and body action amplitude control signals. S501: Group interest hot spot weights in the live situation change set are extracted, interest hot spot identification processing is performed on the data in the live situation change set, the hot spot content is weighted according to the audience attention and interaction intensity, various interest hot spots are classified and arranged according to the influence degree, a weight data structure of different hot spot categories is constructed, and group interest hot spot weights are obtained. S502: The group interest hot spot weights are called to perform corresponding position fusion processing with the live situation change set, the hot spot weight data and the situation change data are integrated according to the position relationship, the fusion data is comprehensively processed, a composite data structure reflecting the hot spot influence intensity is constructed, and hot spot influence intensity data are established. S503: Based on the hot spot influence intensity data, interest hot spot appearance frequency information is obtained, the hot spot influence intensity data and the interest hot spot appearance frequency are proportionally combined, the two types of data are integrated, an execution instruction set containing expression change, voice rhythm and body action amplitude control parameters is constructed, and a live performance execution set is generated.

[0013] As a further scheme of the present application, the control signal generation process of the live performance execution set contains a feedback closed-loop mechanism: After inputting the expression change control signal into the virtual digital human face driving model, real-time collection of audience bullet screen emotional feedback data is performed; Through the emotional interaction fusion set, the emotional tendency deviation state is detected, and when the deviation degree exceeds the baseline range, the emotional weight adaptive adjustment is triggered; Based on the picture area heat value change trend after the execution of the body movement amplitude control signal, the heat block weight parameter is dynamically updated; The emotional weight and the heat block parameter after the fusion adjustment are combined to reconstruct the signal dynamic generation model of the live performance execution set.

[0014] A digital human live broadcast system based on deep learning, the system comprises: A face and picture capturing module acquires a virtual digital human face capture frame rate, a live picture area heat value, and audience comment bullet screen keywords, compares the face capture frame rate and the live picture area heat value on the same time line, matches and combines the comparison results and the audience comment bullet screen keywords at the same time point, and generates an audience interaction time sequence set; An interaction intensity calculation module calculates the number of likes and the number of shares based on the audience interaction time sequence set, adds the two data and calculates the ratio with the total number of audiences in the corresponding time period, vector combines the calculation result and the number of comments, and generates a real-time interaction intensity set; An emotional fusion generation module extracts the number of positive emotional words and the number of negative emotional words in the comment bullet screen keywords according to the real-time interaction intensity set, multiplies the two types of numbers by the corresponding emotional weight and sums them up, calculates the difference between the two types of sum values and performs ratio operation with the total number of comments, multiplies the obtained group emotional deviation with the real-time interaction intensity set in multiple dimensions, and generates an emotional interaction fusion set; A situation change recognition module calls the emotional interaction fusion set and the online number change rate, performs vector splicing operation on the two data on the same time line, calculates the change amplitude of the splicing vector and performs numerical superposition with the topic heat change rate, and generates a live broadcast situation change set; A live performance output module extracts the group interest hotspot weight based on the live broadcast situation change set, multiplies the group interest hotspot weight and the live broadcast situation change set at the corresponding positions and sums them up, and combines the interest hotspot appearance frequency in proportion to generate a live performance execution set.

[0015] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects: In the virtual digital person live broadcast, the face capture frame rate is compared with the picture area heat value in the time line, and the interactive time sequence set is formed by matching the comment barrage keywords, the interaction intensity set is generated by combining the like and share data, the emotion interaction fusion set is calculated by using the comment emotion quantity and weight, the situation change set is formed by fusing the online number change rate and the topic heat change rate, finally, the expression, voice and action control execution set is output by combining the situation change and group interest hot spot, realizing the correlation operation and feature fusion of multi-dimensional data in the same time line, so that the live performance is consistent with the audience emotion and situation change in real time. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 The method flowchart of the present application is shown in the figure; Figure 2 The acquisition flowchart of the audience interactive time sequence set of the present application is shown in the figure; Figure 3 The acquisition flowchart of the real-time interaction intensity set of the present application is shown in the figure; Figure 4 The acquisition flowchart of the emotion interaction fusion set of the present application is shown in the figure; Figure 5 The acquisition flowchart of the live broadcast situation change set of the present application is shown in the figure; Figure 6 The acquisition flowchart of the live broadcast performance execution set of the present application is shown in the figure. DETAILED DESCRIPTION

[0017] The technical solutions in the present application will be described below with reference to the drawings.

[0018] In the embodiments of the present application, the words such as "example", "for example" are used to represent as an example, illustration or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. In fact, the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be one of the two.

[0019] In the embodiments of the present application, "image" and "picture" can be used interchangeably at times. It should be pointed out that when the distinction is not emphasized, the meanings expressed are consistent. "Of", "corresponding" and "corresponding" can be used interchangeably at times. It should be pointed out that when the distinction is not emphasized, the meanings expressed are consistent.

[0020] In the embodiments of the present application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1. When the distinction is not emphasized, the meanings expressed are consistent.

[0021] In order to make the technical problems, technical solutions and advantages of the present application clearer, specific embodiments will be described in detail below with reference to the drawings.

[0022] Please refer to Figure 1 The present application provides a technical solution: a digital person live broadcast method based on deep learning, comprising the following steps: S1: Obtain the virtual digital person face capture frame rate, audience comment barrage keywords and live broadcast picture area heat value, and compare the virtual digital person face capture frame rate with the live broadcast picture area heat value according to the time line, match and combine the comparison results with the audience comment barrage keywords according to the same time point, and generate an audience interaction time sequence set; S2: Count and add the number of likes and the number of shares in the audience interaction time sequence set, and calculate the ratio with the total number of audiences in the same time period, and combine the calculation results with the number of comments to generate a real-time interaction intensity set; S3: According to the real-time interaction intensity set, extract the number of positive and negative emotional words in the comment barrage keywords, multiply the number of two categories by the corresponding emotional weight respectively and sum, calculate the difference value of the two sum values and perform ratio operation with the total number of comments, combine the operation result with the real-time interaction intensity set to generate a sentiment interaction fusion set; S4: Perform vector splicing operation on the sentiment interaction fusion set and the online audience change rate in the same time line, calculate the change amplitude of the spliced vector and perform numerical superposition with the topic heat change rate, reorder the superposition results according to time sequence and filter out the data points with a change amplitude exceeding the set limit to generate a live broadcast situation change set; S5: Based on the live broadcast situation change set, extract the group interest hot spot weight, multiply the hot spot weight with the live broadcast situation change set at the corresponding position and sum, merge the obtained numerical value with the interest hot spot appearance frequency according to the proportion to form a live broadcast performance execution set which can directly control the virtual digital person's expression change, voice rhythm and body movement amplitude in live broadcast.

[0023] The audience interaction time sequence set includes picture attention track distribution, comment semantic category and heat concentration block, the real-time interaction intensity set includes interaction frequency vector, active audience proportion and comment density interval, the sentiment interaction fusion set includes emotion change trend value, emotion proportion structure and emotion interaction correlation feature, the live broadcast situation change set includes situation change amplitude sequence, hot topic switching record and audience gathering and scattering fluctuation mode, and the live broadcast performance execution set includes expression change control signal, voice rhythm adjustment signal and body movement amplitude control signal.

[0024] Please refer to Figure 2 The acquisition step of the audience interaction time sequence set is: S101: Obtain the frame rate data sequence output by the virtual digital human face capture device and the heat value sequence collected by the live picture area heat sensor, align the two sequences according to the data timestamp, ensure one-to-one correspondence at each time point, calculate the absolute difference value of the face capture frame rate value and the area heat value at each time point, integrate the difference values of all time points into a sequence to generate a face capture heat difference sequence; Obtain the frame rate data sequence output by the virtual digital human face capture device and the heat value sequence collected by the live picture area heat sensor, at timestamp T001, the face capture device, specifically model OptiTrack-S250e, outputs frame rate data of frames per second (FPS), at the same time, the heat sensor array covering the live picture, specifically model FLIR-AX8, collects the average heat value of the key area (the upper body area of the virtual digital human) of heat units, which are obtained by normalizing the original readings of the sensor, the timestamp alignment process matches the frame rate data and the heat value data with millisecond-level precision, for example, at T002, the frame rate is obtained as FPS, and the heat value is heat units, at T003, the frame rate is obtained as FPS, and the heat value is heat units, calculate the absolute difference value of the face capture frame rate value and the area heat value at each time point, the operation process is , wherein is the face capture heat difference value at t, at T001, the difference value is , at T002, the difference value is , at T003, the difference value is , integrate the difference values of all time points into a sequence, specifically, arrange and combine the calculation results of 60 data points collected every second within 10 seconds into a data vector in chronological order, and the vector constitutes the face capture heat difference sequence.

[0025] S102: Obtain the audience comment barrage data stream, the data contains timestamp and text content, parse the barrage text at each time point, extract the keywords in the preset keyword list, count the number of occurrences of each keyword at each time point, take the counting result as a sequence element, and generate a barrage keyword heat sequence; A viewer comment barrage data stream is acquired, which is received in real time from a live broadcast platform server through a WebSocket interface, and each piece of data is in JSON format and includes a timestamp, a user ID, and text content. For example, at T001+500 ms, the comment {“timestamp”: T001500, “uid”: “user001”, “text”: “This digital person is amazing, 666”} is received. The comment text is parsed for each time point, where the “time point” is defined as a 1-second time window, for example, [T001, T002). Keywords in a preset keyword list are extracted, the keyword list is constructed based on historical live broadcast data and general network language, and is stored in a Redis database. The keywords are classified into interactive keywords (such as “666”, “amazing”, and “follow”), product keywords (such as “link”, “price”, and “purchase”), and question keywords (such as “who” and “how to do”). In the [T001, T002) time window, a total of 150 comments are received, including 35 comments containing “666”, 18 comments containing “amazing”, and 25 comments containing “link”. The number of occurrences of each keyword at each time point is counted, and the total number of occurrences of the keywords in the time window is . The counting result is taken as a sequence element, and the element is the comment keyword heat value of the time window [T001, T002), denoted as . The keyword heat values of each subsequent 1-second time window are calculated in the same way, for example, the keyword heat value of the [T002, T003) window is . The heat values are arranged in chronological order to generate a comment keyword heat sequence.

[0026] S103: Align the face capture heat force difference sequence and the comment keyword heat sequence according to the same timestamp. For the face capture heat force difference sequence value and the comment keyword heat sequence value in each matching time point, a single data unit is combined to establish a viewer interaction time sequence set. Align the face capture heat force difference sequence and the comment keyword heat sequence according to the same timestamp. Since the sampling frequency of the face capture heat force difference sequence is much higher than that of the comment keyword heat sequence (the former is millisecond-level, and the latter is second-level), down-sampling aggregation is used for alignment. Specifically, the arithmetic mean of all face capture heat force difference values in a 1-second window (for example, [T001, T002)) is calculated to match the comment keyword heat value of the window. It is assumed that in the [T001, T002) window, 600 difference values are collected, and the arithmetic mean is . The comment keyword heat value of the window is . For the face capture heat force difference sequence value and the comment keyword heat sequence value in each matching time point, a single data unit is combined, which is a two-dimensional vector in the form of At the time point of [T001, T002), the data unit is Similarly, at the time point of [T002, T003), the average difference value is The keyword heat value is The data unit is Arrange the data units at all time points in sequence to form a two-dimensional array to establish the audience interaction time sequence set.

[0027] Please refer to Figure 3 The acquisition steps of the real-time interaction intensity set are as follows: S201: Acquire the like number and share number data in the audience interaction time sequence set, statistically summarize the like number and statistically summarize the share number, perform an accumulation and merging operation on the like number statistical result and the share number statistical result, integrate the two types of interaction behavior data, and generate interaction behavior summary data; Acquire the like number and share number data in the audience interaction time sequence set. The two data are pulled at the end of each time window (for example, 1 second) through the API interface provided by the live broadcast platform. In the time window [T001, T002), the like number returned by the API is , and the share number is Statistically summarize the like number and the share number. The statistical summary here means that the data in a single time window is directly taken as the statistical result of the period, the like number statistical result is accumulated and merged with the share number statistical result, and the operation process is , wherein is the interaction behavior summary data of the t time window. In the [T001, T002) window, the value is In the next time window [T002, T003), the like number acquired is , and the share number is The interaction behavior summary data is Integrate the two types of interaction behavior data to generate the interaction behavior summary data.

[0028] S202: Call the interaction behavior summary data and the total number of audiences in the same time period, perform a ratio calculation on the interaction behavior summary data and the total number of audiences, perform an association analysis operation on the two data, establish the proportional relationship between the interaction behavior and the audience size, and obtain the interaction participation ratio data; Call the interaction behavior summary data and the total number of audiences in the same time period. The total number of audiences data is also acquired at the end of each time window through the live broadcast platform API. In the time window [T001, T002), the total number of audiences is People, the interactive behavior summary data and the total number of audience to perform ratio calculation processing, the operation process is , wherein is the interactive participation ratio data of t time window, in the [T001, T002) window, the ratio is , the correlation analysis operation is performed on the two data, and the correlation analysis here is the ratio calculation described above, which reflects the proportion of the audience who actively likes or shares under the current audience size, in the [T002, T003) window, the total number of audience increases to People, the interactive behavior summary data is , then the interactive participation ratio is , the proportion relationship between the interactive behavior and the audience size is established, and the interactive participation ratio data is obtained.

[0029] S203: On the interactive participation ratio data and the comment quantity data, the interactive participation ratio data and the comment quantity are integrated according to the vector combination rule, the vectorization assembly processing is performed on the two data, the multi-dimensional interactive data structure is formed, and the real-time interactive intensity set is established; Based on the interactive participation ratio data and the comment quantity data, the comment quantity data is obtained from the original barrage stream in S102 step, that is, the total number of received barrage in each time window, in the [T001, T002) window, the total number of barrage is , the interactive participation ratio data and the comment quantity are integrated according to the vector combination rule, the combination rule is to form a two-dimensional vector by juxtaposing two scalar data, the first dimension of the vector is the interactive participation ratio, and the second dimension is the normalized comment quantity, the normalization processing of the comment quantity is based on the maximum comment quantity in the last 10 minutes, assuming that the maximum single second comment quantity in the 10 minutes is 500, the normalization operation is , in the [T001, T002) window, the normalized comment quantity is , the vectorization assembly processing is performed on the two data, and the generated vector is , after substituting the numerical value, it is , similarly, in the [T002, T003) window, the total number of barrage is , the normalized comment quantity is , the interactive participation ratio is , then the vector is , the multi-dimensional interactive data structure is formed, and the real-time interactive intensity set is established.

[0030] Please refer to Figure 4 , the acquisition steps of the emotional interaction fusion set are: S301: Extract positive and negative emotion words from the comment barrage keywords according to the real-time interaction intensity set, perform polarity recognition and classification processing on the emotion words, classify the emotion words into different categories according to the positive and negative attributes, count the occurrence frequency of emotion words in each category, and establish an emotion word polarity distribution table; According to the real-time interaction intensity set, the positive and negative emotion words in the comment barrage keywords are extracted. This process calls the parsed barrage text content and uses a predefined emotion dictionary for matching. The dictionary contains positive words (such as "like", "good-looking", "support") and negative words (such as "boring", "trash", "stuttering"). Perform polarity recognition and classification processing on the emotion words. For example, for the barrage "too good-looking, always support", the system identifies the positive words "good-looking" and "support". For the barrage "the picture is stuttering, not interesting", the negative words "stuttering" and "not interesting" are identified. According to the positive and negative attributes, the emotion words are classified into different categories. The occurrence frequency of emotion words in each category is counted. In the 150 barrages in the time window [T001, T002), a total of 80 positive emotion words and 15 negative emotion words are counted. An emotion word polarity distribution table is established.

[0031] Table 1: Emotion word polarity distribution table

[0032] As shown in Table 1, the table records the statistical number of positive and negative emotion words in different time windows. In the [T002, T003) window, 110 positive emotion words and 12 negative emotion words are counted in 190 barrages.

[0033] S302: Call the emotion word polarity distribution table and corresponding emotion weight data, perform multiplication operation on the positive and negative emotion word frequencies respectively with the corresponding weight, perform summation processing on the positive and negative weighted results, calculate the difference value of the two summation results, perform ratio operation on the difference value and the total number of comments, and obtain the sentiment tendency index; Call the emotion word polarity distribution table and corresponding emotion weight data. The setting of emotion weight is based on a data set containing 5000 labeled sentiment polarity (-1 to +1) barrages for regression analysis. In the experiment, word frequency is used as the independent variable and user post-score is used as the dependent variable. The weight coefficient is obtained by least squares fitting. The experimental results show that the contribution of positive emotion words to user satisfaction is about 1.5 times that of negative emotion words. Therefore, the positive emotion weight is set to , and the negative emotion weight is Perform multiplication operation on the positive and negative emotion word frequencies respectively with the corresponding weight. In the [T001, T002) window, the positive weighted result is , and the negative weighted result is Summing the positive and negative weighted results, the description here should be calculating the difference, calculating the difference between the two types of summation results, obtaining the weighted sentiment difference The difference is divided by the total number of comments, which is The operation process is The sentiment tendency index is The sentiment tendency index is obtained.

[0034] S303: Based on the sentiment tendency index and the real-time interaction intensity set, the sentiment tendency index and the real-time interaction intensity set are combined to perform data combination operation, and the two types of data are integrated and processed to construct a composite data structure containing interaction intensity and sentiment tendency, and a sentiment interaction fusion set is established. Based on the sentiment tendency index and the real-time interaction intensity set, the sentiment tendency index is The corresponding vector in the real-time interaction intensity set is The sentiment tendency index and the real-time interaction intensity set are combined to perform data combination operation, which is vector splicing. The sentiment tendency index is added to the end of the real-time interaction intensity vector as a new dimension. The two types of data are integrated and processed, and the processed vector form is After substituting the values of the [T001, T002) window, we get In the [T002, T003) window, the sentiment tendency index is The real-time interaction intensity vector is The fused vector is A composite data structure containing interaction intensity and sentiment tendency is constructed, and a sentiment interaction fusion set is established.

[0035] Please refer to Figure 5 The acquisition steps of the live situation change set are: S401: The sentiment interaction fusion set and the online number change rate are combined to perform vector splicing processing under the same time line, and the two types of data are matched and corresponding according to the time node. The vector splicing operation constructs a composite vector structure, integrates the sentiment interaction data and the number change data, and establishes a composite vector data; The sentiment interaction fusion set and the online number change rate are combined to perform vector splicing processing under the same time line. The online number change rate is The difference between the online numbers of adjacent time windows is obtained by dividing the number of the previous window, that is Using the data of [T001, T002) and [T002, T003), the online number changes from 2500 to 2650, and the change rate of [T002, T003) window is , the emotional interaction fusion vector of the window [T002, T003) is matched and corresponded with the online number change rate of the window The vector splicing operation is performed to construct a composite vector structure, and the spliced vector is After substituting the numerical value, it is The emotional interaction data and the number change data are integrated to establish a composite vector data.

[0036] S402: Call the composite vector data to perform change amplitude measurement processing, obtain topic heat change rate data, perform superposition integration operation on the vector change amplitude and the topic heat change rate, perform reordering processing on the superposition result in time sequence, construct a time sequence arranged change data structure, and obtain a time sequence change data table; The measurement of the change amplitude is realized by calculating the Euclidean distance between the composite vectors of the continuous time windows, and the operation process is Where n is the vector dimension, which is 4 here, and the data of needs to be calculated. Assuming that people, then The composite vector is The change amplitude of the window T002-T003 is Obtain the topic heat change rate data, which is obtained by monitoring the topic related to the live content on the social media platform and calculating the discussion quantity change rate by an independent crawler module. Assuming that the topic heat change rate in the window [T002, T003) is Perform superposition integration operation on the vector change amplitude and the topic heat change rate to obtain the total change amplitude Then Perform reordering processing on the superposition result in time sequence to construct a time sequence arranged change data structure, and obtain a time sequence change data table.

[0037] Table 2: Time sequence change data table

[0038] As shown in Table 2, the table lists the key indicators in the consecutive three time windows and the total change amplitude calculated finally, which is used for subsequent change amplitude screening processing.

[0039] S403: Perform change amplitude screening processing based on the time sequence change data table, compare the data change amplitude with the set limit threshold value, select the data points exceeding the limit threshold value, and induce and arrange the data points meeting the screening condition to construct a data set containing abnormal change characteristics, and generate a live situation change set; ​The change range screening processing is performed based on the time sequence change data table, and the data change range is compared with a set limit threshold value . The setting of the limit threshold value is obtained by statistical analysis of the total change range of 10 seconds before and after the "burst point event" (such as a sudden increase in the number of people and a gift burst) in the past 100 live broadcasts . The value of the 85% quantile point is taken as the threshold value. Through statistics, the value is . The total change range calculated in the [T002, T003) window is , which is greater than the threshold value . Therefore, the data point at this time point is screened out. The data point contains a complex vector corresponding to the time and a total change range value . The data points meeting the screening conditions are summarized and arranged, and all the screened data points and their associated information (time stamp, original data vector, and change range value) are stored in a new set to construct a data set containing abnormal change characteristics and generate a live situation change set.

[0040] Please refer to Figure 6 . The acquisition steps of the live performance execution set are as follows: S501: Extract the group interest hot spot weight in the live situation change set, perform interest hot spot recognition processing on the data in the live situation change set, distribute the weights of the hot spot content according to the audience attention and interaction intensity, classify and arrange various interest hot spots according to the influence degree, construct a weight data structure of different hot spot categories, and obtain the group interest hot spot weight; Extract the group interest hot spot weight in the live situation change set, perform interest hot spot recognition processing on the data in the live situation change set, this processing process is to analyze the barrage keywords in the time window corresponding to the screened data point, the main keywords identified in the S102 step at the screened time point [T002, T003) are "link", "price", and "purchase", which together point to the "commodity delivery" interest hot spot. The weights of the hot spot content are distributed according to the audience attention and interaction intensity. The weight calculation formula is , wherein is the total number of keywords related to the hot spot, is the total change range that triggers screening, and the total number of keywords related to "commodity delivery" is (hypothetical value) in [T002, T003). The total change range is . Therefore, the weight of the hot spot is . Various interest hot spots are classified and arranged according to the influence degree, a weight data structure of different hot spot categories is constructed, and the group interest hot spot weight is obtained.

[0041] Table 3: Group interest hotspot weight table

[0042] As shown in Table 3, the table lists the interest hotspots identified at different context change points and their calculated weight values.

[0043] S502: Call the group interest hotspot weight and the live broadcast context change set for corresponding position fusion processing, integrate the hotspot weight data and the context change data according to the position relationship, perform comprehensive processing on the fusion data, construct a composite data structure reflecting the hotspot influence intensity, and establish the hotspot influence intensity data; Call the group interest hotspot weight and the live broadcast context change set for corresponding position fusion processing, which means that the hotspot weight calculated in the last step is weighted with the original composite vector that triggered the hotspot, and the hotspot weight data and the context change data are integrated according to the position relationship. Specifically, the weight is multiplied by the interaction-related dimensions in the composite vector , namely the interaction participation ratio and the normalized comment quantity . In the [T002, T003) window, the weight is , the composite vector is , the weighted interaction dimension becomes , and the sentiment and number change dimensions remain unchanged. The comprehensive processing here refers to calculating the norm of the weighted vector as a quantitative indicator of the hotspot influence intensity, which is calculated as , and a composite data structure reflecting the hotspot influence intensity is constructed, and the hotspot influence intensity data is established.

[0044] S503: Obtain the interest hotspot appearance frequency information based on the hotspot influence intensity data, perform proportional merging processing on the hotspot influence intensity data and the interest hotspot appearance frequency, integrate the two types of data, construct an execution instruction set containing expression change, speech rhythm, and body movement amplitude control parameters, and generate a live performance execution set; The control signal generation process of the live performance execution set includes a feedback closed-loop mechanism: After the expression change control signal is input into the virtual digital human face driving model, real-time collection of audience bullet screen emotional feedback data is performed; Detect the emotional tendency deviation state through the emotional interaction fusion set, and trigger emotional weight adaptive adjustment when the deviation degree exceeds the baseline range; Based on the picture area heat value change trend after execution of the body movement amplitude control signal, dynamically update the heat block weight parameter; By integrating and adjusting the emotion weights and heatmap parameters, the signal dynamic generation model of the live performance execution set is reconstructed. Information on the frequency of interest-related trending topics is obtained based on the intensity of their influence. For example, the trending topic "product promotion" appeared 3 times in the past hour. Subsequently, the intensity of the hotspot's impact will be determined. The frequency is summed using weighted averages. The merging process is performed, where the weighting coefficients are... and The final control signal magnitude was determined based on regression analysis of historical live broadcast data and then calculated using numerical values. This magnitude is then converted into facial expression changes through a preset linear mapping rule. Speech rhythm and range of motion Specific control parameters are used to generate a live performance execution set. This execution set initiates a feedback closed-loop mechanism the instant the application is applied. For example, during the execution of... After the mapped "hearty laughter" emoji, the system immediately collects the comments from the following 5 seconds, calculates the degree of deviation in the sentiment index, and if the deviation value is as follows... It exceeded the preset stability threshold. This triggers an adaptive adjustment of the emotion weight, and based on... The calculation will shift the positive sentiment weight from Updated to Simultaneously, the system detected that... After the driver performs a large-scale limb movement, the heat value of the key area of ​​the image changes from Rise to This positive feedback will trigger an update to the thermal block weight parameters, changing the calculated weight of the corresponding region from... Upgraded to Ultimately, the adjusted sentiment weights were incorporated. With thermal block parameters The dynamic generation model will be directly applied to the calculation of the execution set generation at the next live broadcast scenario change point.

[0045] A deep learning-based digital human live streaming system, the system comprising: The face and image capture module acquires the virtual digital human face capture frame rate, the heat value of the live screen area, and the keywords of the audience comments and bullet screens. It compares the face capture frame rate and the heat value of the live screen area on the same timeline, and matches and combines the comparison results with the keywords of the audience comments and bullet screens at the same time point to generate a set of audience interaction time sequences. An interaction intensity calculation module, based on the audience interaction timing set statistics like number and share number, adds the two data and calculates the ratio with the total number of audience in the corresponding time period, combines the calculation result with the comment number to generate a real-time interaction intensity set; An emotional fusion generation module, according to the real-time interaction intensity set, extracts the number of positive emotional words and the number of negative emotional words in the comment barrage keywords, multiplies the two types of numbers by the corresponding emotional weights and sums them up, calculates the difference between the two types of summed values and performs ratio operation with the total number of comments, multiplies the obtained group emotional offset with the real-time interaction intensity set to generate an emotional interaction fusion set; A situation change recognition module, calls the emotional interaction fusion set and the online number change rate, performs vector splicing operation on the two data in the same time line, calculates the change amplitude of the splicing vector and performs numerical superposition with the topic heat change rate to generate a live broadcast situation change set; A live broadcast performance output module, based on the live broadcast situation change set, extracts the group interest hotspot weight, multiplies the group interest hotspot weight with the live broadcast situation change set at the corresponding position and sums them up, and merges with the interest hotspot appearance frequency in proportion to generate a live broadcast performance execution set.

[0046] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A deep learning digital human live streaming method, characterized in that, The method comprises the following steps: S1: obtaining a virtual digital human face capture frame rate, audience comment barrage keywords, and live picture region heat value, comparing the virtual digital human face capture frame rate with the live picture region heat value according to a time line, matching the comparison result with the audience comment barrage keywords according to the same time point, and generating an audience interaction time sequence set; S2: statistically adding the number of likes and the number of shares in the audience interaction time sequence set, performing ratio calculation on the total number of audiences in the same time period, performing vector combination on the calculation result and the number of comments, and generating a real-time interaction intensity set; S3: extracting the number of positive and negative emotional words in the comment barrage keywords according to the real-time interaction intensity set, multiplying the number of two types by the corresponding emotional weight respectively and summing, calculating the difference value of the two types of sum values and performing ratio operation on the total number of comments, combining the operation result with the real-time interaction intensity set, and generating a sentiment interaction fusion set; S4: performing vector splicing operation on the sentiment interaction fusion set and the online audience change rate in the same time line, calculating the change amplitude of the spliced vector and performing numerical superposition on the topic heat change rate, and generating a live situation change set.

2. The deep learning digital human live streaming method of claim 1, wherein: The audience interaction time sequence set comprises a picture attention track distribution, a comment semantic category, and a heat concentration block, the real-time interaction intensity set comprises an interaction frequency vector, an active audience proportion, and a comment density interval, the sentiment interaction fusion set comprises an emotional change trend value, an emotional proportion structure, and an emotional interaction correlation feature, and the live situation change set comprises a situation change amplitude sequence, a hot topic switching record, and an audience gathering and scattering fluctuation mode.

3. The deep learning digital human live streaming method of claim 1, wherein, The obtaining step of the audience interaction time sequence set is: S101: obtaining a frame rate data sequence output by a virtual digital human face capture device and a heat value sequence collected by a live picture region heat sensor, aligning the two sequences according to data time stamps, ensuring one-to-one correspondence at each time point, calculating the absolute difference value of the face capture frame rate value and the region heat value at each time point, integrating the difference values of all time points into a sequence to generate a face capture heat difference sequence; S102: obtaining the audience comment barrage data stream, the data containing a time stamp and text content, analyzing the barrage text at each time point, extracting keywords in a preset keyword list, counting the number of keywords appearing at each time point, taking the counting result as a sequence element, and generating a barrage keyword heat sequence; S103: aligning the face capture heat difference sequence and the barrage keyword heat sequence according to the same time stamp, combining the face capture heat difference sequence value and the barrage keyword heat sequence value in each matched time point into a single data unit, and establishing an audience interaction time sequence set.

4. The deep learning digital human live streaming method of claim 1, wherein, The obtaining step of the real-time interaction intensity set is: S201: obtaining the number of likes and the number of shares in the audience interaction time sequence set, statistically counting the number of likes and the number of shares, performing cumulative merging operation on the number of likes and the number of shares, integrating the two types of interaction behavior data, and generating interaction behavior summary data; S202: Call the interactive behavior summary data and the total number of audiences in the same time period, perform ratio calculation processing on the interactive behavior summary data and the total number of audiences, perform relevance analysis operation on the two data, establish the proportional relationship between the interactive behavior and the audience size, obtain the interactive participation ratio data; S203: Based on the interactive participation ratio data and the comment quantity data, integrate the interactive participation ratio data and the comment quantity according to the vector combination rule, perform vectorization assembly processing on the two data, form a multi-dimensional interactive data structure, and establish a real-time interactive intensity set.

5. The deep learning digital human live streaming method of claim 1, wherein, The acquisition step of the emotional interaction fusion set is: S301: According to the real-time interactive intensity set, extract positive and negative emotional words in the comment bullet screen keywords, perform polarity recognition and classification processing on the emotional words, classify the emotional words into different categories according to the positive and negative attributes, count the occurrence frequency of the emotional words in each category, and establish an emotional word polarity distribution table; S302: Call the emotional word polarity distribution table and the corresponding emotional weight data, respectively perform multiplication operation on the positive and negative emotional word frequencies and the corresponding weights, sum the positive and negative weighted results, calculate the difference value of the two summed results, and perform ratio operation on the difference value and the total number of comments to obtain the emotional tendency index; S303: Based on the emotional tendency index and the real-time interactive intensity set, perform data combination operation on the emotional tendency index and the real-time interactive intensity set, perform fusion integration processing on the two types of data, construct a composite data structure containing interactive intensity and emotional tendency, and establish an emotional interaction fusion set.

6. The deep learning digital human live streaming method of claim 1, wherein, The acquisition step of the live broadcast situation change set is: S401: Perform vector splicing processing on the emotional interaction fusion set and the online number change rate in the same time line, match the two types of data according to the time nodes, perform vector splicing operation to construct a composite vector structure, integrate the emotional interaction data and the number change data, and establish a composite vector data; S402: Call the composite vector data to perform change amplitude measurement processing, obtain topic heat change rate data, perform superposition integration operation on the vector change amplitude and the topic heat change rate, perform reordering processing on the superposition result according to time sequence, construct a time sequence arranged change data structure, and obtain a time sequence change data table; S403: Based on the time sequence change data table, perform change amplitude screening processing, compare the data change amplitude with the set limit threshold value, select the data points exceeding the limit threshold value, and arrange the data points meeting the screening condition to construct a data set containing abnormal change characteristics, and generate a live broadcast situation change set.

7. The deep learning digital human live streaming method of claim 1, wherein, The method further comprises: S5: Based on the live broadcast situation change set, extract the group interest hot spot weight, multiply the hot spot weight and the live broadcast situation change set in the corresponding position, and then sum, merge the obtained value with the interest hot spot appearance frequency according to the proportion, form a live broadcast performance execution set which can directly control the expression change, voice rhythm and body movement amplitude of the virtual digital person in the live broadcast; The live performance execution set includes expression change control signals, speech rhythm adjustment signals, and body movement amplitude control signals.

8. The deep learning digital human live streaming method of claim 1, wherein, The obtaining step of the live performance execution set is: S501: Extract the group interest hotspot weight in the live context change set, perform interest hotspot identification processing on the data in the live context change set, distribute weights to the hotspot content according to the audience attention and interaction intensity, classify and arrange the various interest hotspots according to the influence degree, construct the weight data structure of different hotspot categories, and obtain the group interest hotspot weight; S502: Call the group interest hotspot weight and the live context change set for corresponding position fusion processing, integrate the hotspot weight data and the context change data according to the position relationship, perform comprehensive processing on the fusion data, construct a composite data structure reflecting the hotspot influence intensity, and establish the hotspot influence intensity data; S503: Based on the hotspot influence intensity data, obtain the interest hotspot frequency information, perform proportional merging processing on the hotspot influence intensity data and the interest hotspot frequency, integrate the two types of data, construct an execution instruction set containing expression change, speech rhythm, and body movement amplitude control parameters, and generate a live performance execution set.

9. The deep learning digital human live streaming method of claim 1, wherein, The control signal generation process of the live performance execution set includes a feedback closed loop mechanism: After the expression change control signal is input into the virtual digital human face driving model, the emotional feedback data of the audience's bullet screen is collected in real time; Through the emotional interaction fusion set, the emotional tendency deviation state is detected, and when the deviation degree exceeds the baseline range, the emotional weight self-adaptive adjustment is triggered; Based on the trend of the picture area heat value change after the execution of the body movement amplitude control signal, the heat block weight parameter is dynamically updated; The adjusted emotional weight and the heat block parameter are fused to reconstruct the signal dynamic generation model of the live performance execution set.

10. A deep learning digital human live streaming system, characterized in that, The system is used for the deep learning digital human live method in any one of claims 1-9, and the system comprises: A face and picture capturing module acquires a virtual digital human face capture frame rate, a live picture area heat value, and audience comment bullet screen keywords, compares the face capture frame rate and the live picture area heat value in the same time line, matches and combines the comparison result and the audience comment bullet screen keywords at the same time point, and generates an audience interaction time sequence set. An interaction intensity calculation module calculates the number of likes and the number of shares based on the audience interaction time sequence set, adds the two data, calculates the ratio of the total number of audiences in the corresponding time period, vector combines the calculation result and the number of comments, and generates a real-time interaction intensity set. An emotional fusion generation module extracts the number of positive emotional words and the number of negative emotional words in the comment bullet screen keywords according to the real-time interaction intensity set, multiplies the two types of numbers by the corresponding emotional weight and sums them up, calculates the difference between the two types of summed values and the ratio of the total number of comments, multiplies the obtained group emotional deviation by the real-time interaction intensity set in multiple dimensions, and generates an emotional interaction fusion set. The context change recognition module calls the emotional interaction fusion set and the online number change rate, performs vector splicing operation on the two data in the same time line, calculates the change amplitude of the spliced vector, and performs numerical superposition on the topic heat change rate to generate a live broadcast context change set; The live broadcast performance output module extracts a group interest hot spot weight based on the live broadcast context change set, multiplies the group interest hot spot weight and the live broadcast context change set in a corresponding position, sums up, and merges with the interest hot spot appearance frequency in proportion to generate a live broadcast performance execution set.

Citation Information

Patent Citations

  • Interactive live broadcast system based on total-sensing digital human and use method of interactive live broadcast system

    CN116233536A

  • Digital human live broadcast interaction method and system based on artificial intelligence

    CN117880566A

  • AI-based digital human generation method and digital human live broadcast system

    CN118608663A

  • Live broadcast interaction method and system based on AI digital human

    CN119071521A

  • Digital human intelligent interaction and posture expression synthesis method based on multi-modal synchronization

    CN120068923A