A digital human broadcasting platform and method based on computer vision
By constructing a computer vision-based digital human broadcasting platform, analyzing training video footage and generating emotion transition strategies, the problem of abrupt emotional expression in digital human broadcasting was solved, achieving smooth transitions in emotional expression and adaptability to marketing scenarios, thus improving user experience.
Patent Information
- Application Number
- CN202510568033.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Existing digital human broadcasting technology suffers from excessive emotional expression during marketing content delivery, resulting in abrupt and unnatural emotional expressions that negatively impact content dissemination and user experience, thus hindering the promotion and popularization of digital human technology.
By analyzing training video footage using computer vision methods, a phased three-dimensional dataset of emotions and a dataset of emotion transition features are constructed. Combined with the marketing content information to be broadcast, a digital human emotion transition strategy is generated to ensure that emotion expression conforms to human cognitive habits and achieves a smooth transition.
It significantly improved the adaptability of digital human emotional expression to marketing scenarios, enhanced the experience of the target audience, ensured that the process of digital human emotional changes conformed to human cognitive habits, and enhanced the effect of digital human broadcasting.
Smart Images

Figure CN120499443B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of AI digital human, in particular to a digital human broadcasting platform and method based on computer vision. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, digital human technology has gradually become an important form of expression in the fields of film and television, live broadcast, e-commerce marketing, etc. Digital human can provide real-time and continuous content services, greatly improving the efficiency of content dissemination. Digital human is becoming an important way of content output.
[0003] However, the existing digital human broadcasting technology has the problem of excessive emotional expression in the process of marketing content broadcasting, which makes the emotional expression of digital human abrupt and unnatural, thereby negatively affecting the content dissemination effect and user experience, and is not conducive to the promotion and popularization of digital human technology. SUMMARY
[0004] The present application provides a digital human broadcasting platform and method based on computer vision to solve the above technical problems.
[0005] In a first aspect, the present application provides a digital human broadcasting method based on computer vision, which comprises: obtaining a training video material set, performing three-dimensional emotion analysis on different marketing broadcasting stages according to the training video material set, and determining a three-dimensional stage emotion data set; analyzing the three-dimensional stage emotion data set to determine an emotion transition feature data set; obtaining to-be-broadcast marketing content information, determining a digital human emotion transition strategy based on the three-dimensional stage emotion data set and the emotion transition feature data set according to the to-be-broadcast marketing content information; and generating and outputting digital human broadcasting content according to the to-be-broadcast marketing content information and the digital human emotion transition strategy.
[0006] Through the present solution, the training video material set is analyzed from the perspective of multi-modal data fusion, the emotional characteristics of real people in different marketing stages during broadcasting are obtained, a three-dimensional stage emotion data set is constructed, and on this basis, an emotion transition feature data set representing the change of emotional characteristics in the emotion transition process of real people is obtained. Combined with the to-be-broadcast marketing content information, a digital human emotion transition strategy is generated to generate corresponding digital human broadcasting content, which is provided to the target audience, significantly improving the adaptability of digital human emotion expression and marketing scenarios, realizing smooth transition of digital human emotion, making the digital human emotion change process highly consistent with human cognitive habits, and improving the experience of the target audience.
[0007] Optionally, the marketing broadcasting stage includes a warm-up stage, a product introduction stage, a product demonstration stage, a promotion stage, and a summary stage; and the three-dimensional emotion includes excitement, affinity, and professionalism.
[0008] By the scheme, the marketing broadcast stage is divided into a warm-up stage, a product introduction stage, a product demonstration stage, a promotion stage and a summary stage, and the core emotional targets required to be expressed in different stages are standardized by an emotional three-dimensional, which includes excitement, affinity and professionalism, so that the subsequent analysis process of the digital human emotion conversion strategy has significant stage pertinence, and the emotional expression of the digital human has higher accuracy and consistency, and the emotional expression of the digital human is more humanized.
[0009] Optionally, the emotional three-dimensional analysis of different marketing broadcast stages according to the training video material set to determine the stage emotional three-dimensional data set comprises:
[0010] The training video material set is analyzed, and visual information, voice information and text information corresponding to each training video material are extracted; each training video material is divided into a plurality of video segments according to a preset time window length; the background switching frequency and the gesture action density of each video segment are extracted according to the visual information; the real-time speech speed and the real-time volume of each video segment are extracted according to the voice information; the stage keyword set of each video segment is extracted based on a preset stage keyword dictionary according to the text information; the stage mapping index corresponding to each video segment is determined based on the stage keyword set, the background switching frequency, the gesture action density, the real-time speech speed and the real-time volume; each training video material is divided into marketing stages according to the stage mapping index corresponding to each video segment, and a stage video material set is determined; the stage video material set is analyzed to determine the emotional three-dimensional index information corresponding to each marketing stage; and the stage video material set and the emotional three-dimensional index information are integrated to construct the stage emotional three-dimensional data set.
[0011] By the scheme, multi-modal feature fusion is used to extract feature parameters that can map the marketing stage of a video from the dimensions of vision, voice and text, and then the stage mapping index that can map the marketing stage of a video segment is quantified to accurately determine the marketing stage of each video segment, and then the stage video material set is integrated, on the basis of which the emotional three-dimensional features of each marketing stage are analyzed to determine the emotional three-dimensional index information corresponding to different marketing stages. After the stage video material set and the emotional three-dimensional index information are integrated, the stage emotional three-dimensional data set is constructed, which contains video materials corresponding to different marketing stages and real human emotional three-dimensional features in different marketing stages, and provides an accurate data basis for the subsequent analysis process of the digital human emotion conversion strategy.
[0012] Optionally, based on the set of stage keywords, according to the background switching frequency, gesture action density, real-time speech speed and real-time volume, a stage mapping index corresponding to each video segment is determined, specifically as follows:
[0013]
[0014] wherein S(t) is the stage mapping index corresponding to the video segment in the current time window t, s is the current evaluation stage, S a is the set of marketing stages, a is the text influence weight, w i is the stage influence weight of the i-th stage keyword, k i is the i-th stage keyword, F(k i ) is the word frequency of the i-th stage keyword, β is the audio influence weight, A(t) is the audio influence factor in the current time window t, γ is the visual influence weight, V(t) is the visual influence factor in the current time window t, λ1 is the speech speed influence weight, x(t) is the real-time speech speed of the video segment in the current time window t, μ x is the average speech speed of the current video, σ x is the speech speed standard deviation of the current video, λ2 is the volume influence weight, y(t) is the real-time volume of the video segment in the current time window t, μ y is the average volume of the current video, σ y is the volume standard deviation of the current video, η1 is the background influence weight, b(t) is the background switching frequency of the video segment in the current time window t, η2 is the gesture influence weight, and d(t) is the gesture action density of the video segment in the current time window t.
[0015] By the scheme, the comprehensive scores of the visual features, speech features and text features in the current video segment in each evaluation stage in the set of marketing stages are evaluated by mathematical analysis means, the maximum value among them is taken as the stage mapping index, and the evaluation stage corresponding to the stage mapping index is the marketing stage corresponding to the current video segment, so that the marketing stage to which the video segment belongs is accurately positioned, and the accuracy of the marketing stage division of the video material is improved.
[0016] Optionally, the analysis of the stage video material set determines the emotional three-dimensional index information corresponding to each marketing stage, comprising: according to the image action analysis model, analyzing the stage video material set to determine the gesture average acceleration and posture stability index of each marketing stage; according to the deep learning face recognition model, analyzing the stage video material set to determine the expression affinity index of each marketing stage; according to the speech processing model, analyzing the stage video material set to determine the speech fundamental frequency standard deviation, speech affinity index and professional term density of each marketing stage; according to the gesture average acceleration, posture stability index, expression affinity index, speech fundamental frequency standard deviation and speech affinity index, respectively determine the excitement index, professionalism index and affinity index of each marketing stage, and construct the emotional three-dimensional index information, specifically as follows:
[0017]
[0018] wherein, E x is the excitement index, κ1 is the fundamental frequency influence weight, κ2 is the acceleration influence weight, f p is the speech fundamental frequency standard deviation, v g is the gesture average acceleration, E p is the professionalism index, θ1 is the posture influence weight, S p is the posture stability index, θ2 is the term influence weight, n t is the professional term density, E a is the affinity index, τ1 is the expression influence weight, U s is the expression affinity index, τ2 is the speech influence weight, E w is the speech affinity index.
[0019] Through the scheme, the excitement index, professionalism index and affinity index of each marketing stage are quantified respectively according to the gesture average acceleration, posture stability index, expression affinity index, speech fundamental frequency standard deviation and speech affinity index by using mathematical analysis means, and the emotional three-dimensional index information is constructed, so that the emotional three-dimensional index information can scientifically and accurately reflect the emotional characteristic targets required to be reached in different influencing stages, and the adaptation degree of subsequent digital human emotion conversion strategy and marketing stage is improved.
[0020] Optionally, the analysis of the phase emotion three-dimensional data set to determine the emotion transition feature data set comprises: determining a plurality of emotion transition time nodes according to the phase mapping index corresponding to each video segment in the phase emotion three-dimensional data set; and analyzing the phase video material set according to the plurality of emotion transition time nodes, extracting the visual information, the voice information and the text information of the video material corresponding to each emotion transition time node within a preset emotion transition focus range, and constructing the emotion transition feature data set.
[0021] According to the scheme, the emotion transition time nodes are accurately analyzed and recognized, and on this basis, the visual information, the voice information and the text information of the video material within the preset emotion transition focus range with the emotion transition time node as the midpoint are extracted, the emotion transition feature data set is constructed, the emotion change transition features of real humans in the marketing phase transition process are fully reflected, and the emotion transition strategy of the digital person is scientifically based on data.
[0022] Optionally, the determination of the plurality of emotion transition time nodes according to the phase mapping index corresponding to each video segment in the phase emotion three-dimensional data set is specifically as follows:
[0023]
[0024] wherein, Δt p is the emotion transition time node, S(t) is the phase mapping index corresponding to the video segment in the current time window t, θ is a preset mutation threshold, and M(t) is the middle time node in the current time window t.
[0025] According to the scheme, the mutation in the phase mapping index is analyzed according to the phase mapping index corresponding to each video segment in the phase emotion three-dimensional data set by using mathematical analysis means, the middle time node in the time window corresponding to the phase mapping index when the mutation occurs is taken as the emotion transition time node, and the emotion transition time node is accurately positioned.
[0026] Optionally, determining the digital human emotion conversion strategy based on the staged emotion 3D dataset and the emotion transition feature dataset, according to the marketing content information to be broadcast, includes: analyzing the marketing content information to be broadcast, determining the content to be broadcast corresponding to different marketing stages, and determining the transition time range between different marketing stages based on the content to be broadcast; analyzing the emotion transition feature dataset, determining the amplitude range of facial expression transition, the amplitude range of action transition, the amplitude range of tone transition, and the amplitude range of text emotional polarity transition; based on the transition time range, and constrained by the amplitude range of facial expression transition, the amplitude range of action transition, the amplitude range of tone transition, and the amplitude range of text emotional polarity transition, determining the target emotion 3D feature values corresponding to different time points within the transition time range; and constructing the digital human emotion conversion strategy based on the target emotion 3D feature values corresponding to different time points.
[0027] This solution utilizes a dynamic mechanism to determine the transition time range, dynamically analyzing the emotional transition time between different marketing stages based on the content to be broadcast. This ensures that the temporal characteristics of the emotional transition process align with the actual emotional transitions of real people. Furthermore, by incorporating multimodal amplitude constraints, the solution quantifies the three-dimensional characteristic values of the target emotion at different time points within the transition time range. Based on this, a digital human emotional transition strategy is constructed. This strategy clearly defines the required levels of excitement, affinity, and professionalism for the digital human at different time points during the emotional transition process, thereby improving the smoothness of the digital human's emotional transition and ensuring that the digital human's emotional characteristics match its current marketing state.
[0028] Optionally, the determination of the target emotion three-dimensional feature values corresponding to different time points within the transition time range is specifically based on the following formula:
[0029]
[0030] Among them, E out (t) represents the three-dimensional feature value of the target emotion corresponding to the current time point t, P i Let t be the current three-dimensional feature value of the i-th emotion category, ψ be the preset time influence coefficient corresponding to the three-dimensional feature value of the i-th emotion category, and t be the current three-dimensional feature value of the i-th emotion category. start t is the starting time point of the transition time range. end E represents the end time point of the transition time range. i This represents the i-th type of three-dimensional sentiment index corresponding to the current marketing stage.
[0031] This solution utilizes mathematical analysis, employing a nonlinear time decay model and a differentiated time influence coefficient, to precisely quantify the three-dimensional feature values of the target emotion required at different time points. This ensures that the emotional transition process of the digital human highly conforms to the inertial characteristics of human emotional transitions, thereby reducing the unnaturalness of emotional changes in the digital human and improving the smoothness of emotional transitions.
[0032] Secondly, this application provides a computer vision-based digital human broadcasting platform, the platform comprising: an emotion feature analysis module, used to acquire a training video material set, and based on the training video material set, perform three-dimensional emotion analysis on different marketing broadcasting stages to determine a stage-specific three-dimensional emotion dataset; an emotion transition analysis module, used to analyze the stage-specific three-dimensional emotion dataset to determine an emotion transition feature dataset; a strategy analysis module, used to acquire marketing content information to be broadcast, and based on the stage-specific three-dimensional emotion dataset and the emotion transition feature dataset, determine a digital human emotion transition strategy according to the marketing content information to be broadcast; and an output module, used to generate and output digital human broadcasting content according to the marketing content information to be broadcast and the digital human emotion transition strategy. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 This is a schematic diagram illustrating an application scenario provided in one embodiment of this application;
[0035] Figure 2 A flowchart illustrating a computer vision-based digital human broadcasting method provided in one embodiment of this application;
[0036] Figure 3 This is a schematic diagram of the structure of a computer vision-based digital human broadcasting platform provided in an embodiment of this application. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0038] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0039] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0040] Existing digital human broadcasting technology suffers from excessive emotional expression during marketing content delivery, resulting in abrupt and unnatural emotional expressions from the digital human. This negatively impacts the effectiveness of content dissemination and user experience, hindering the promotion and popularization of digital human technology.
[0041] Based on this, this application provides a computer vision-based digital human broadcasting platform and method. From the perspective of multimodal data fusion, the training video material set is analyzed to obtain the emotional characteristics exhibited by the real person during different marketing stages of the broadcasting process. A three-dimensional emotional dataset for each stage is constructed. Based on this, an emotion transition feature dataset representing the changes in emotional characteristics during the emotional transition process of the real person is analyzed. Combined with the marketing content information to be broadcast, a digital human emotion transition strategy is generated, thereby generating corresponding digital human broadcasting content. This content is then provided to the target audience, significantly improving the adaptability of digital human emotional expression to the marketing scenario, achieving a smooth transition of digital human emotions, and ensuring that the process of digital human emotional change highly conforms to human cognitive habits, thus improving the experience of the target audience and ultimately enhancing the effectiveness of digital human broadcasting.
[0042] Figure 1 This is a schematic diagram illustrating an application scenario provided by this application. When using digital humans to deliver marketing content, the method provided in this application significantly improves the adaptability of digital human emotional expression to the marketing scenario, achieving a smooth transition of digital human emotions.
[0043] Specifically, the method of this application is applied to any server, which obtains a set of training video materials provided by the user. From the perspective of multimodal data fusion, the training video material set is analyzed to obtain the emotional characteristics exhibited by the real person during different marketing stages of the broadcast. A three-dimensional emotional dataset for each stage is constructed. Based on this, an emotional transition feature dataset representing the changes in emotional characteristics during the emotional transition process of the real person is obtained. Combined with the marketing content information to be broadcast provided by the user, a digital human emotional transition strategy is generated, thereby generating corresponding digital human broadcast content. This corresponding digital human broadcast content is then provided to the target audience, significantly improving the adaptability of digital human emotional expression to the marketing scenario, achieving a smooth transition of digital human emotions, and making the process of digital human emotional change highly consistent with human cognitive habits, thus improving the experience of the target audience. Specific implementation methods can be found in the following embodiments.
[0044] Figure 2 This is a flowchart illustrating a computer vision-based digital human broadcasting method according to an embodiment of this application. The method of this embodiment can be applied to the server in the above scenario. Figure 2 As shown, the method includes:
[0045] S201. Obtain the training video material set. Based on the training video material set, perform three-dimensional emotion analysis on different marketing broadcast stages to determine the three-dimensional emotion dataset for each stage.
[0046] The training video material set can be a collection of complete broadcast videos by live anchors in different marketing scenarios, along with their associated metadata. This training video material set can be provided by users. Marketing broadcast stages can be different content broadcasting stages within a marketing scenario, such as product introduction stages or promotional stages. The emotional dimension can be three dimensions used to characterize the emotional features of the broadcaster in a marketing scenario, including excitement, professionalism, and approachability. The staged emotional dimension dataset can be a collection of data containing the emotional dimension features of broadcasters in different marketing stages.
[0047] Specifically, existing digital human broadcasting technologies typically use fixed emotion templates, failing to dynamically and smoothly adjust emotional expressions according to the stage characteristics of marketing content. This results in abrupt emotional shifts during marketing stage transitions, and the emotional characteristics presented vary to different stages of different products. For example, during the product introduction stage of electronic products, the required level of professionalism is higher (to reduce consumer doubts), while during the promotion stage, the required level of excitement is higher (to stimulate consumer emotions). Using fixed emotion templates cannot guarantee the effectiveness of marketing broadcasts. This solution uses computer vision, speech extraction, and natural language processing technologies to extract visual, speech, and text information from real-person anchor video materials. It comprehensively analyzes the three-dimensional emotional characteristics exhibited by real-person anchors at different marketing stages within marketing videos, constructing a staged three-dimensional emotion dataset. This provides a scientific data foundation for subsequently building digital human emotion transition strategies that conform to human behavioral habits.
[0048] S202. Analyze the three-dimensional dataset of phased emotions to determine the dataset of emotion transition characteristics.
[0049] An emotion transition feature dataset can be a collection of data containing various feature changes related to emotional expression during the transition of a live streamer in the marketing phase, such as changes in facial expressions and movements.
[0050] Specifically, there are two main reasons for the unnatural emotional changes in digital humans. First, the changes in emotional characteristics during emotional transitions lack a trend and are often sudden, such as a sudden increase in tone of voice. Second, there is a lack of coordination between different types of emotional characteristics during emotional transitions; for example, there are obvious changes in facial expressions, but the movements and tone of voice do not match the changes in amplitude. By analyzing the three-dimensional dataset of staged emotions, we can extract the emotional change nodes of real anchors during the switching process of marketing stages. Using the emotional change nodes as a benchmark, we can further analyze the amplitude of a series of emotional characteristic changes of real anchors within the corresponding time range before and after the node, such as the amplitude of facial expression changes and the amplitude of movement changes. In this way, we can construct an emotional transition feature dataset, which serves as a scientific data basis for achieving smooth emotional transitions in subsequent digital human emotional transition processes.
[0051] S203. Obtain information on the marketing content to be broadcast, and determine the digital human's emotion conversion strategy based on the phased three-dimensional emotion dataset and the emotion conversion feature dataset.
[0052] The marketing content to be broadcast can be product marketing content that needs to be broadcast by the digital human, such as marketing scripts, product parameters, and promotional plans. This marketing content can be provided by the user. The digital human emotion conversion strategy can be a strategy to control the three-dimensional emotional characteristics of the digital human at different times during different marketing stages.
[0053] Specifically, to avoid mechanical broadcasting by digital humans, the semantics of marketing content need to be dynamically linked with the emotional expression of digital humans. By integrating the analysis of the broadcast content structure and the characteristics of emotional transition, it is ensured that the emotional changes of digital humans not only conform to the marketing logic but also have the natural fluency of human broadcasting. By analyzing the stage division and transition time range of the content to be broadcast, and using the transition amplitude in the emotional transition feature dataset as the constraint boundary, mathematical analysis methods are used to quantify the three-dimensional values of the target emotion required at different times, generating a digital human emotion transition strategy that takes into account both the content focus and the audience experience.
[0054] S204. Generate and output digital human broadcast content based on the marketing content information to be broadcast and the digital human emotion conversion strategy.
[0055] Digital human broadcasting content can be digital human broadcasting videos that include elements such as digital human movements, speech synthesis, background effects, and subtitles.
[0056] Specifically, based on the digital human emotion conversion strategy, the three-dimensional target emotion values at different time points are used as a basis to map the three-dimensional target emotion values to digital human control parameters such as facial muscle control parameters, limb movement trajectories, and speech synthesis parameters. Combined with background requirements and special effects insertion rules, the video is rendered to generate digital human broadcast content that conforms to the video platform's encoding specifications. The digital human broadcast content is then provided to the target audience through the video platform.
[0057] This solution analyzes the training video footage from a multimodal data fusion perspective to identify the emotional characteristics of real people during different marketing stages of the broadcast. A three-dimensional emotional dataset representing each stage is constructed. Based on this, an emotional transition feature dataset characterizing the changes in emotional features during the emotional transition process is derived. Combined with the marketing content to be broadcast, a digital human emotional transition strategy is generated, leading to the creation of corresponding digital human broadcast content. This content is then provided to the target audience, significantly improving the adaptability of digital human emotional expression to the marketing scenario, achieving a smooth transition of digital human emotions, and ensuring that the process of digital human emotional changes highly aligns with human cognitive habits, thereby enhancing the experience for the target audience.
[0058] In some embodiments, the marketing broadcast phase includes a warm-up phase, a product introduction phase, a product demonstration phase, a promotion phase, and a summary phase; the emotional dimension includes excitement, affinity, and professionalism.
[0059] The warm-up phase can be a stage in the marketing process where lighthearted topics are used to attract the audience's attention and establish initial trust and interaction. The product introduction phase can be a stage where the product's functions, technical parameters, and core advantages are systematically explained. The product demonstration phase can be a stage where the actual usage effects of the product are demonstrated through physical displays or virtual simulations. The promotion phase can be a stage where limited-time offers and free gifts are announced to stimulate consumer decisions. The summary phase can be a stage where the marketing process reviews the product's value proposition and guides subsequent purchasing behavior.
[0060] Excitement level can be quantified by factors such as changes in tone of voice, range of body movements, and speaking speed to represent the persuasiveness of the broadcast. Affinity can be quantified by factors such as the naturalness of facial expressions, frequency of eye contact, and interactive nature of language to represent the friendliness perceived in the broadcast. Professionalism can be quantified by factors such as the accuracy of terminology use, stability of posture, and logical coherence to represent the professional credibility of the broadcast content.
[0061] Specifically, the core emotional goals differ significantly across different marketing phases. The warm-up phase needs to reduce user resistance, while the promotion phase needs to generate a sense of urgency. Without phase differentiation, the digital human might adopt a single emotional mode, leading to a mismatch between content and user expectations. For example, maintaining high excitement during the summary phase might interfere with users' reception of core information, while a lack of professionalism during the product demonstration phase would weaken the persuasiveness of the technology. The core emotional goals of different marketing phases are mainly reflected through a comprehensive approach using three emotional dimensions: excitement, professionalism, and affinity. A single emotional dimension is insufficient to fully represent complex emotional expressions, such as "excitement" or "affinity" itself. There is a risk of being too simplistic or one-sided. However, by comprehensively analyzing emotions through three dimensions—excitement, professionalism, and affinity—digital humans can more realistically convey the tone, attitude, and intention in information. Consumers are often very sensitive to emotional expression, and inaccurate or biased emotional output may lead to ineffective communication or loss of trust. The introduction of multiple dimensions makes the emotional expression of digital humans more accurate and consistent, avoiding expression errors. Moreover, the diversity of human emotional expression is usually driven by multiple dimensions of emotions simultaneously. Simulating this mechanism is a key way for AI digital humans to improve the realism and naturalness of interaction, which helps to make digital humans more human-like.
[0062] This solution divides the marketing broadcast phase into a warm-up phase, a product introduction phase, a product demonstration phase, a promotion phase, and a summary phase. The core emotional goals to be expressed in each phase are standardized using three dimensions of emotion: excitement, affinity, and professionalism. This ensures that the subsequent analysis of digital human emotion conversion strategies is highly phase-specific, while also making the emotional expression of digital humans more accurate and consistent, and more human-like.
[0063] In some embodiments, the training video material set is analyzed to extract visual, audio, and textual information corresponding to each training video material; each training video material is divided into time windows according to a preset time window length to determine several video segments; based on the visual information, the background switching frequency and gesture density of each video segment are extracted; based on the audio information, the real-time speech rate and real-time volume of each video segment are extracted; based on a preset stage keyword dictionary, a set of stage keywords for each video segment is extracted according to the textual information; based on the set of stage keywords, the stage mapping index corresponding to each video segment is determined according to the background switching frequency, gesture density, real-time speech rate, and real-time volume; based on the stage mapping index corresponding to each video segment, each training video material is divided into marketing stages to determine a set of stage video materials; the set of stage video materials is analyzed to determine the three-dimensional emotion index information corresponding to each marketing stage; based on the set of stage video materials and the three-dimensional emotion index information, a stage-specific three-dimensional emotion dataset is constructed.
[0064] Visual information can be features extracted from video footage through image analysis, including background visuals, gestures, body movements, and facial expressions. Audio information can be quantifiable parameters in the audio stream of the video footage, such as speech rate, volume, fundamental frequency, and emotional tone. Text information can be corresponding subtitles or speech-to-text content in the video footage, including keywords and semantic information. The preset time window length can be a fixed time unit (e.g., 5 seconds / segment) for segmenting the video, used for refined analysis of stage features. Background switching frequency can be the number of significant changes in the video background within a unit of time (e.g., scene changes or dynamic element updates), reflecting visual complexity. Gesture density can be the frequency of hand movements within a unit of time (e.g., waving, pointing), used to measure body language activity. Real-time speech rate can be the average number of words spoken per second (words / second) in the current video segment, reflecting the pace of information delivery. Real-time volume can be the average speech energy (decibels) in the current video segment, used to determine emotional intensity. The preset stage keyword dictionary can be a collection of keywords that can map different marketing stages of video content. The stage keyword set can be a collection of keywords within the current video segment that can map to the current marketing stage. The stage mapping index can be a quantitative indicator used to characterize the marketing stage of the current video segment. The stage video material set can be a collection containing all video segments corresponding to each marketing stage. The emotional 3D index information can be quantitative information used to characterize the emotional goals of the marketing stage, including a three-dimensional vector of excitement, affinity, and professionalism.
[0065] Specifically, the distribution of marketing stages varies across different video materials, and these stages are typically not consecutive; for example, product introduction and product demonstration stages are often interspersed. Analyzing the features in video materials that can map the marketing stage of a video segment from visual, audio, and textual perspectives can significantly improve the accuracy of segmenting video segments into marketing stages. This can be achieved through OpenCV (Open Source Computer Vision). The library (an open-source computer vision library) detects background transitions (a transition is defined as a difference of >30% between adjacent frames). During promotional phases, background transitions are typically infrequent to avoid distracting users. The MediaPipe gesture recognition model is used to statistically analyze hand movement distances and determine gesture density (number of actions per second). Product demonstrations require high-density pointing gestures to enhance expressive flexibility. The Librosa library is used to extract speech rate (number of syllables / duration) and volume (RMS energy value). During the warm-up phase, excessively fast speech and high volume can cause user anxiety, so the speech rate is usually slower and the volume lower. In contrast, the speech rate and volume are typically higher during promotional phases to create a sense of urgency and promotional intensity. The BERT model is used to perform semantic analysis on subtitles, matching them to a pre-defined keyword dictionary for each phase and statistically analyzing keyword frequencies. These phase keywords are a direct basis for dividing marketing phases (for example, the word "demonstration" appears significantly more frequently during product demonstrations than in other phases).
[0066] Based on the above parameters, mathematical analysis is used to quantify the stage mapping index corresponding to each video segment, thereby mapping the marketing stage of each video segment and realizing the marketing stage division of each training video material. The training video materials with completed marketing stage division are integrated into a stage video material set. Through mathematical analysis, the three-dimensional emotion index information corresponding to different marketing stages is comprehensively analyzed and quantified from aspects such as facial expressions, actions, and voice. After integrating the stage video material set and the three-dimensional emotion index information, a staged emotion three-dimensional dataset is constructed. This staged emotion three-dimensional dataset contains both video materials corresponding to different marketing stages and the three-dimensional emotion features of real people under different marketing stages, providing an accurate data foundation for the subsequent analysis of digital human emotion conversion strategies.
[0067] This solution utilizes multimodal feature fusion to extract feature parameters that map the marketing stage of a video from three dimensions: visual, speech, and text. These feature parameters are then combined to quantify the stage mapping index, accurately determining the marketing stage of each video segment. This results in a stage-specific video material set. Based on this, the three-dimensional emotional features of each marketing stage are analyzed to determine the corresponding three-dimensional emotional index information. By integrating the stage-specific video material set and the emotional three-dimensional index information, a stage-specific emotional three-dimensional dataset is constructed. This dataset includes both video materials corresponding to different marketing stages and the three-dimensional emotional features of real people at different marketing stages, providing an accurate data foundation for subsequent analysis of digital human emotion conversion strategies.
[0068] In some embodiments, based on the set of stage keywords, the stage mapping index corresponding to each video segment is determined according to the background switching frequency, gesture action density, real-time speech rate and real-time volume, specifically as follows (1):
[0069]
[0070] Where S(t) is the stage mapping index corresponding to the video segment in the current time window t, and s is the current evaluation stage. a For the marketing stage set, α is the text influence weight, w i Let k be the stage influence weight of the keyword in the i-th stage. i For the i-th stage keyword, F(k) i Let ) represent the word frequency of the keyword in the i-th stage, β represent the audio influence weight, A(t) represent the audio influence factor in the current time window t, γ represent the visual influence weight, V(t) represent the visual influence factor in the current time window t, λ1 represent the speech rate influence weight, x(t) represent the real-time speech rate of the video segment in the current time window t, and μ represent the speech rate of the video segment in the current time window t. x σ represents the average speaking speed of the current video. x Let λ be the standard deviation of the speech rate in the current video, λ2 be the weight of the volume effect, y(t) be the real-time volume of the video segment in the current time window t, and μ be the standard deviation of the speech rate in the current video. y σ represents the average volume of the current video. y Let η be the standard deviation of the current video volume, η1 be the background influence weight, b(t) be the background switching frequency of the video segment under the current time window t, η2 be the gesture influence weight, and d(t) be the gesture action density of the video segment under the current time window t.
[0071] The current pre-assessment phase can be the target marketing phase that is currently pending judgment.
[0072] The marketing phase set can be any set of all possible marketing phases, including warm-up, product introduction, demonstration, promotion, and closing phases.
[0073] The text influence weight can be the influence weight of text features on the stage mapping index (e.g., α = 0.4 indicates that text accounts for 40%), and the text influence weight can be obtained by fitting historical data.
[0074] The stage influence weight can be the importance weight of a keyword (such as "discount") in a certain stage to the current evaluation stage. The stage influence weight can be quantified based on the frequency and distinctiveness of the keyword in the target stage.
[0075] The audio influence weight can be the influence weight of speech features on the stage mapping index in the current evaluation stage, and the audio influence weight can be obtained by fitting historical data.
[0076] The audio impact factor can be a speech feature score calculated by combining real-time speech rate and volume.
[0077] The visual impact weight can be the weight of the influence of visual features on the stage mapping index in the current evaluation stage. The visual impact weight can be obtained by fitting historical data.
[0078] The visual impact factor can be a visual feature score obtained by quantifying the combined background switching frequency and gesture density.
[0079] The speech rate influence weight can be the weight of the speech rate's influence on the audio influence factor.
[0080] The volume influence weight can be the weight of the influence of volume on the audio influence factor.
[0081] Background influence weight can be the weight of the influence of background switching frequency on visual influence factors.
[0082] The gesture influence weight can be the influence weight of gesture density on the visual influence factor.
[0083] The weights of speech rate, volume, background, and gesture were all determined through experimental optimization.
[0084] Specifically, through formula (1) The visual, audio, and text features of the current video segment are weighted and quantified for each evaluation stage within the marketing stage set. The maximum value is taken as the stage mapping index, and the evaluation stage corresponding to this index is the marketing stage for the current video segment. Z-score standardization was applied to real-time speech rate and volume, and then the relationship between real-time speech rate and volume and audio impact factor was described by weighted quantization to obtain the audio impact factor. The relationship between background switching frequency and gesture density and visual impact factor was described by weighted quantization using (t) = η1·b(t) + η2·d(t) to obtain the visual impact factor.
[0085] This solution utilizes mathematical analysis to evaluate the comprehensive scores of visual, audio, and textual features in the current video segment under each evaluation stage within the marketing stage set. The maximum value is taken as the stage mapping index, and the evaluation stage corresponding to this stage mapping index is the marketing stage corresponding to the current video segment. This achieves precise positioning of the marketing stage to which the video segment belongs and improves the accuracy of dividing the marketing stages of video materials.
[0086] In some embodiments, based on the image motion analysis model, the video material set of each marketing stage is analyzed to determine the average acceleration of gestures and the posture stability index of each marketing stage; based on the deep learning face recognition model, the video material set of each marketing stage is analyzed to determine the facial affinity index of each marketing stage; based on the speech processing model, the video material set of each marketing stage is analyzed to determine the speech fundamental frequency standard deviation, speech affinity index, and professional terminology density of each marketing stage; based on the average acceleration of gestures, the posture stability index, the facial affinity index, the speech fundamental frequency standard deviation, and the speech affinity index, the excitement index, professionalism index, and affinity index of each marketing stage are determined respectively, and the three-dimensional emotion index information is constructed, specifically as follows (2):
[0087]
[0088] Among them, E x The excitation index is κ1, where κ1 is the weight of the fundamental frequency influence, κ2 is the weight of the acceleration influence, and f is the excitation index. p v represents the standard deviation of the fundamental frequency of speech. g E represents the average acceleration of the gesture. p The professionalism index is represented by θ1, which is the weight of the attitude influence. p Here, θ² is the attitude stability index, θ₂ is the terminology influence weight, and n... t For the density of technical terms, E a U is the affinity index, τ1 is the weight of facial expression influence, and U s τ2 represents the facial affinity index, E represents the influence weight of speech, and E represents the facial expression affinity index. w This refers to the voice affinity index.
[0089] The aforementioned influence weights can all be quantitative values used to characterize the degree of influence of the current parameter on the quantification result, and the aforementioned influence weights are all obtained through regression analysis of experimental data.
[0090] The average acceleration of hand gestures can be the average rate of change of hand movement speed during the marketing phase, reflecting the activity level of body language.
[0091] The posture stability index can be the reciprocal of the displacement variance of key trunk points (such as the shoulder and hip). The larger the value, the more stable the posture.
[0092] The facial affinity index can be a weighted average of the probability that a person's facial expression tends to be positive and optimistic (such as smiling). The higher the value, the stronger the affinity.
[0093] The standard deviation of speech fundamental frequency can be the degree of fluctuation of speech fundamental frequency (pitch). The larger the standard deviation, the more obvious the fluctuation of pitch.
[0094] The voice affinity index can be calculated as the ratio of low-frequency energy in the voice (the ratio between the energy in the 200-500Hz frequency band and the total energy), reflecting the softness of the voice.
[0095] Terminology density can be defined as the ratio of the number of times a term appears to the total number of words per unit of time.
[0096] Specifically, the gesture recognition and posture estimation modules in MediaPipe are used to analyze the range of changes in hand gestures and postures in video materials corresponding to different marketing stages, determining the average acceleration of gestures and the posture stability index for each marketing stage. OpenFace is used to evaluate the facial motion unit scores of people in video materials corresponding to different marketing stages, determining the facial affinity index for each marketing stage. PyWorld Vocoder is used to analyze the audio data of people in video materials corresponding to different marketing stages, determining the standard deviation of the fundamental frequency of speech for each marketing stage. The Librosa library is used to analyze the audio data of people, quantifying the proportion of low-frequency energy in speech to derive the speech affinity index. Based on professional domain dictionary matching technology, the proportion of industry-specific terms mentioned in the audio data of people is analyzed to determine the density of professional terms.
[0097] When a person is excited, their speech will show more fluctuations, and their body movements will increase in amplitude, as shown in formula (2). The effects of the standard deviation of speech fundamental frequency and gesture acceleration on excitability are described. An activation function is used to ensure the non-negativity of the results, and a normalization factor is introduced. Limiting the range of results ensures that the numerical value of arousal directly reflects the current level of emotional excitement; professionalism is primarily reflected in posture and the use of professional terminology. More stable posture indicates a more professional performance, and frequent use of terminology indicates proficiency in professional knowledge. This is achieved through θ1·tanh(10·S) p )+θ2·log(1+n tThis paper describes the impact of posture stability and terminology density on professionalism. Due to significant differences in posture characteristics among individuals, a nonlinear transformation function is used to compress posture stability to fit a wider range. Once terminology density reaches a certain level, its impact on professionalism slows down (all reaching a high level of professionalism). The influence of terminology density is then mitigated through a convergent change using a logarithmic function. Finally, the paper uses min(1,τ1·U) as an example. s +τ2·E w This describes the combined effect of facial affinity index and linguistic affinity index on affinity. To ensure the rationality and comparability of affinity, the min(1,x) function is used to limit it to the range of [0,1].
[0098] This solution utilizes mathematical analysis to quantify the excitement, professionalism, and affinity indices for each marketing stage based on average gesture acceleration, posture stability index, facial expression affinity index, speech fundamental frequency standard deviation, and speech affinity index. This constructs a three-dimensional emotional index, enabling the three-dimensional emotional index to scientifically and accurately reflect the emotional characteristic goals to be achieved at different impact stages, thereby improving the adaptability of subsequent digital human emotion conversion strategies to the marketing stages.
[0099] In some embodiments, several emotion transition time nodes are determined based on the stage mapping index corresponding to each video segment in the staged emotion 3D dataset; based on the several emotion transition time nodes, the stage video material set is analyzed, and visual, audio and text information of the video material corresponding to each emotion transition time node within the preset emotion transition focus range is extracted to construct an emotion transition feature dataset.
[0100] The emotional shift point can be the point in time when emotional characteristics change significantly between marketing stages.
[0101] The preset focus range of emotion transition can be a preset time window centered on the time node of emotion transition, used to extract transition features. The preset focus range of emotion transition can be set according to the statistical results of historical data.
[0102] Specifically, real human emotional transitions involve a gradual process (e.g., the speech rate gradually slows down 3 seconds before the end of a promotional phase). This gradual process is reflected in a buffer time range in terms of temporal characteristics. Within this buffer time range, the emotional transition from one marketing phase to another is completed. Through mathematical analysis, the abrupt changes in the mapping index of different phases within the three-dimensional dataset of phased emotions are analyzed. The time point corresponding to the video segment where the abrupt change occurs is taken as the emotional transition time node. Then, using the preset emotional transition focus range as the buffer time range benchmark, visual, audio, and textual information corresponding to the video material within the preset emotional transition focus range with the emotional transition time node as the midpoint is extracted to construct an emotional transition feature dataset. This dataset comprehensively reflects the transitional characteristics of real human emotional changes during the marketing phase transition process, serving as a scientific data basis for digital human emotional transition strategies.
[0103] This solution enables precise analysis and identification of emotional transition time points. Based on this, using a preset emotional transition focus range as a buffer time range benchmark, it extracts visual, audio, and textual information corresponding to video materials within the preset emotional transition focus range with the emotional transition time point as the midpoint. This constructs an emotional transition feature dataset to comprehensively reflect the emotional change transition characteristics of real humans during the marketing transition process, serving as a scientific data basis for digital human emotional transition strategies.
[0104] In some embodiments, several emotion transition time points are determined based on the stage mapping index corresponding to each video segment in the staged emotion three-dimensional dataset, specifically as shown in the following formula (3):
[0105]
[0106] Where, Δt p S(t) represents the time node for emotion transition, S(t) represents the stage mapping index corresponding to the video segment under the current time window t, θ represents the preset mutation threshold, and M(t) represents the intermediate time node of the current time window t.
[0107] The preset mutation threshold can be the critical rate of change of the stage mapping index used to characterize the occurrence of emotional shifts.
[0108] Intermediate time nodes can be time nodes located at the midpoint of the corresponding time axis.
[0109] Specifically, using the second derivative in formula (3) The trend of the rate of change of the phase mapping index is described. If the second derivative has a value greater than the preset mutation threshold, it indicates that the change of the phase mapping index has fluctuated significantly, which can be regarded as a significant change in emotion. At this time, the intermediate time node within the time window corresponding to the phase mapping index that caused the above-mentioned situation of exceeding the preset mutation threshold is taken as the emotion conversion time node, so as to achieve accurate positioning of the emotion conversion time node.
[0110] This solution utilizes mathematical analysis to analyze abrupt changes in the stage mapping index corresponding to each video segment within the stage-based emotional 3D dataset. The intermediate time node within the time window of the corresponding stage mapping index when the abrupt change occurs is taken as the emotion transition time node, thus achieving accurate positioning of the emotion transition time node.
[0111] In some embodiments, the system analyzes the marketing content information to be broadcast, determines the content to be broadcast corresponding to different marketing stages, and determines the transition time range between different marketing stages based on the content to be broadcast; it analyzes the emotion transition feature dataset to determine the amplitude range of facial expression transition, the amplitude range of action transition, the amplitude range of tone transition, and the amplitude range of text emotional polarity transition; based on the transition time range, and constrained by the amplitude ranges of facial expression transition, action transition, tone transition, and text emotional polarity transition, it determines the three-dimensional feature values of the target emotion corresponding to different time points within the transition time range; and it constructs a digital human emotion transition strategy based on the three-dimensional feature values of the target emotion corresponding to different time points.
[0112] The transition time range can be a dynamic transition time interval between adjacent marketing stages, determined by the length of the transition broadcast content between adjacent stages.
[0113] The range of facial expression transition amplitudes can be the maximum range of change allowed for digital human facial expression parameters (such as the curvature of the corners of the mouth and the height of the eyebrows) during the transition period.
[0114] The range of motion transition amplitude can be the maximum range of change that digital human gestures are allowed to undergo during the transition period.
[0115] The pitch transition amplitude range can be the range of rise and fall allowed for the fundamental frequency of the speech during the transition period.
[0116] The amplitude range of the text sentiment polarity transition can be the allowed gradient of change in the text sentiment value (positive / negative) during the transition phase.
[0117] The three-dimensional characteristic value of target emotion can be a combination of quantitative indicators of excitement, professionalism, and affinity that the digital human needs to achieve at a specific point in time.
[0118] Specifically, the transitions between different marketing stages vary in content complexity (e.g., a gradual transition from warm-up to product introduction vs. a rapid conclusion from promotion to summary). Mechanically setting a fixed transition duration can lead to a disrupted rhythm. By analyzing the semantic density (density of technical terms) and logical structure (frequency of transition words) of the content to be broadcast, the transition time range can be dynamically calculated. For example, when multiple price discounts are detected in the promotion stage, the transition time is automatically extended to reserve sufficient space for emotional rendering, avoiding cognitive breaks caused by information overload. Within the transition time range, the digital human's expressions, movements, tone of voice, and broadcast text must all match the emotional transition process to prevent unnatural emotional transitions (e.g., excessive facial expression changes can lead to unnatural facial muscle movements in the digital human, even violating physiological limits). This is achieved through image analysis and audio analysis techniques. Using technology and natural language processing, video footage from different marketing stages within the emotion transition feature dataset is analyzed in a unified manner to determine the amplitude ranges of facial expressions, actions, tone of voice, and text sentiment polarity during the emotional transition of a real person at different marketing stages. This constrains the changes in emotional characteristics during the emotional transition of the digital human. Based on this constraint, mathematical analysis is used to quantify the three-dimensional feature values of the target emotion at different time points within the transition time range, and a digital human emotion transition strategy is constructed accordingly. This strategy clarifies the required level of excitement, affinity, and professionalism for the digital human at different time points during the emotion transition, thereby improving the smoothness of the digital human's emotion transition and ensuring that the digital human's emotional characteristics are consistent with its current marketing state.
[0119] This solution utilizes a dynamic mechanism to determine the transition time range, dynamically analyzing the emotional transition time between different marketing stages based on the content to be broadcast. This ensures that the temporal characteristics of the emotional transition process align with the actual emotional transitions of real people. Furthermore, by incorporating multimodal amplitude constraints, the solution quantifies the three-dimensional characteristic values of the target emotion at different time points within the transition time range. Based on this, a digital human emotional transition strategy is constructed. This strategy clearly defines the required levels of excitement, affinity, and professionalism for the digital human at different time points during the emotional transition process, thereby improving the smoothness of the digital human's emotional transition and ensuring that the digital human's emotional characteristics match its current marketing state.
[0120] In some embodiments, the target emotion three-dimensional feature values corresponding to different time points within the transition time range are determined, specifically by the following formula (4):
[0121]
[0122] Among them, E out (t) represents the three-dimensional feature value of the target emotion at the current time point t, P iLet t be the current three-dimensional feature value of the i-th emotion category, ψ be the preset time influence coefficient corresponding to the three-dimensional feature value of the i-th emotion category, and t be the current three-dimensional feature value of the i-th emotion category. start t is the starting time point of the transition time range. end E is the end time point of the transition time range. i This represents the i-th type of three-dimensional sentiment index corresponding to the current marketing stage.
[0123] The preset time impact coefficient can be an adjustment parameter that controls the rate of change of various emotional dimensions during the transition period. The preset time impact coefficient can be obtained based on the analysis of the emotional transition rhythm in historical excellent marketing cases.
[0124] Specifically, human emotion transition is not a linear process, but usually presents the characteristics of "rapid start-smooth transition-precise convergence". If the linear interpolation algorithm is directly used to quantify the three-dimensional feature value of the target emotion, it is easy to cause the digital human to change slowly in the early stage of the emotion transition and suddenly change in the late stage of the emotion transition, which is an unnatural phenomenon. By constructing a nonlinear time decay model through formula (4), the position change of the current time point in the time axis corresponding to the transition time range, the difference between the current three-dimensional emotion index and the three-dimensional emotion index required by the corresponding marketing stage are used as the variable benchmark. The preset time influence coefficient of the corresponding emotion feature is introduced to quantify the target emotion three-dimensional feature value corresponding to the current time point.
[0125] This solution utilizes mathematical analysis, employing a nonlinear time decay model and a differentiated time influence coefficient, to precisely quantify the three-dimensional feature values of the target emotion required at different time points. This ensures that the emotional transition process of the digital human highly conforms to the inertial characteristics of human emotional transitions, thereby reducing the unnaturalness of emotional changes in the digital human and improving the smoothness of emotional transitions.
[0126] Figure 3 A schematic diagram of the structure of a computer vision-based digital human broadcasting platform provided in one embodiment of this application is shown below. Figure 3 As shown, a computer vision-based digital human broadcasting platform 300 in this embodiment includes: an emotion feature analysis module 301, an emotion conversion analysis module 302, a strategy analysis module 303, and an output module 304.
[0127] The emotion feature analysis module 301 is used to acquire a training video material set, and to perform three-dimensional emotion analysis on different marketing broadcast stages based on the training video material set to determine the stage-specific three-dimensional emotion dataset.
[0128] The emotion transition analysis module 302 is used to analyze the staged emotion three-dimensional dataset and determine the emotion transition feature dataset.
[0129] The strategy analysis module 303 is used to acquire information on marketing content to be broadcast, and based on the phased emotion three-dimensional dataset and the emotion conversion feature dataset, determine the digital human emotion conversion strategy according to the information on marketing content to be broadcast.
[0130] The output module 304 is used to generate and output digital human broadcast content based on the marketing content information to be broadcast and the digital human emotion conversion strategy.
[0131] Optionally, in the emotion feature analysis module 301, the marketing broadcast stage includes a warm-up stage, a product introduction stage, a product demonstration stage, a promotion stage, and a summary stage; the three-dimensional emotion includes excitement, affinity, and professionalism.
[0132] Optionally, the emotion feature analysis module 301 is specifically used for:
[0133] Analyze the training video material set and extract the visual information, audio information and text information corresponding to each training video material;
[0134] Based on the preset time window length, each training video material is divided into time windows to determine several video segments;
[0135] Based on the visual information, extract the background switching frequency and gesture density of each video segment;
[0136] Based on the audio information, extract the real-time speech rate and real-time volume of each video segment;
[0137] Based on a preset stage keyword dictionary, the set of stage keywords for each video segment is extracted according to the text information;
[0138] Based on the set of stage keywords, the stage mapping index corresponding to each video segment is determined according to the background switching frequency, gesture density, real-time speech rate and real-time volume.
[0139] Based on the stage mapping index corresponding to each video segment, each training video material is divided into marketing stages to determine the stage video material set.
[0140] Analyze the video material collection for each stage to determine the three-dimensional emotional index information corresponding to each marketing stage;
[0141] Based on the set of video materials for each stage and the three-dimensional emotion index information, the stage-specific three-dimensional emotion dataset is constructed.
[0142] Optionally, when the emotion feature analysis module 301 determines the stage mapping index corresponding to each video segment based on the stage keyword set and according to the background switching frequency, gesture density, real-time speech rate, and real-time volume, the specific formula is as follows:
[0143]
[0144] Where S(t) is the stage mapping index corresponding to the video segment in the current time window t, and s is the current evaluation stage. a For the marketing stage set, α is the text influence weight, w i Let k be the stage influence weight of the keyword in the i-th stage. i For the i-th stage keyword, F(k) i Let ) represent the word frequency of the keyword in the i-th stage, β represent the audio influence weight, A(t) represent the audio influence factor in the current time window t, γ represent the visual influence weight, V(t) represent the visual influence factor in the current time window t, λ1 represent the speech rate influence weight, x(t) represent the real-time speech rate of the video segment in the current time window t, and μ represent the speech rate of the video segment in the current time window t. x σ represents the average speaking speed of the current video. x Let λ be the standard deviation of the speech rate in the current video, λ2 be the weight of the volume effect, y(t) be the real-time volume of the video segment in the current time window t, and μ be the standard deviation of the speech rate in the current video. y σ represents the average volume of the current video. y η1 is the standard deviation of the current video volume, b(t) is the background influence weight, b(t) is the background switching frequency of the video segment under the current time window t, η2 is the gesture influence weight, and d(t) is the gesture action density of the video segment under the current time window t.
[0145] Optionally, when analyzing the set of video materials for each marketing stage and determining the three-dimensional emotional index information corresponding to each stage, the emotion feature analysis module 301 is specifically used for:
[0146] Based on the image motion analysis model, the video material set of the aforementioned stages is analyzed to determine the average acceleration of gestures and the posture stability index for each marketing stage.
[0147] Based on a deep learning facial recognition model, the video material set of the aforementioned stages is analyzed to determine the facial affinity index for each marketing stage;
[0148] Based on the speech processing model, the video material set of the aforementioned stages is analyzed to determine the speech fundamental frequency standard deviation, speech affinity index, and professional terminology density for each marketing stage;
[0149] Based on the average acceleration of the gesture, the posture stability index, the facial expression affinity index, the standard deviation of the voice fundamental frequency, and the voice affinity index, the excitement index, professionalism index, and affinity index for each marketing stage are determined respectively, and the three-dimensional emotion index information is constructed, specifically by the following formula:
[0150]
[0151] Among them, E x The excitation index is defined as follows: κ1 is the weight of the fundamental frequency influence, κ2 is the weight of the acceleration influence, and f... p v is the standard deviation of the fundamental frequency of the speech. g E is the average acceleration of the gesture. p Let θ1 be the professionalism index, and S be the attitude influence weight. p Let θ2 be the attitude stability index, θ2 be the term influence weight, and n be the nth index. t E represents the density of the terminology. a Let U be the affinity index, τ1 be the weight of facial expression influence, and U be the weight of facial expression influence. s Let E be the facial expression affinity index, τ2 be the voice influence weight, and E be the facial expression affinity index. w The voice affinity index is defined as follows.
[0152] Optionally, the emotion conversion analysis module 302 is specifically used for:
[0153] Based on the stage mapping index corresponding to each video segment in the staged emotion 3D dataset, several emotion transition time nodes are determined.
[0154] Based on several emotional transition time points, the video material set of the stage is analyzed, and the visual information, voice information and text information of the video material corresponding to each emotional transition time point within the preset emotional transition focus range are extracted to construct the emotional transition feature dataset.
[0155] Optionally, when the emotion transition analysis module 302 determines several emotion transition time nodes based on the stage mapping index corresponding to each video segment in the staged emotion three-dimensional dataset, the specific formula is as follows:
[0156]
[0157] Where, Δt p Let S(t) be the time node for the emotion transition, S(t) be the stage mapping index corresponding to the video segment under the current time window t, θ be the preset mutation threshold, and M(t) be the intermediate time node of the current time window t.
[0158] Optionally, the strategy analysis module 303 is specifically used for:
[0159] Analyze the marketing content information to be broadcast, determine the content to be broadcast corresponding to different marketing stages, and determine the transition time range between different marketing stages based on the content to be broadcast.
[0160] Analyze the emotion transition feature dataset to determine the amplitude ranges for facial expression transitions, action transitions, tone transitions, and text sentiment polarity transitions.
[0161] Based on the transition time range, and constrained by the facial expression transition amplitude range, the action transition amplitude range, the tone transition amplitude range, and the text emotion polarity transition amplitude range, the target emotion three-dimensional feature values corresponding to different time points within the transition time range are determined.
[0162] The digital human emotion conversion strategy is constructed based on the three-dimensional feature values of the target emotion at different time points.
[0163] Optionally, when determining the target emotion three-dimensional feature values corresponding to different time points within the transition time range, the strategy analysis module 303 uses the following formula:
[0164]
[0165] Among them, E out (t) represents the three-dimensional feature value of the target emotion corresponding to the current time point t, P i Let t be the current three-dimensional feature value of the i-th emotion category, ψ be the preset time influence coefficient corresponding to the three-dimensional feature value of the i-th emotion category, and t be the current three-dimensional feature value of the i-th emotion category. start t is the starting time point of the transition time range. end E represents the end time point of the transition time range. i This represents the i-th type of three-dimensional sentiment index corresponding to the current marketing stage.
[0166] The platform in this embodiment can be used to execute the methods of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
Claims
1. A digital human broadcasting method based on computer vision, characterized in that, include: Obtain a training video material set, and based on the training video material set, perform three-dimensional emotion analysis on different marketing broadcast stages to determine the stage-specific three-dimensional emotion dataset; Analyze the aforementioned three-dimensional dataset of staged emotions to determine the dataset of emotion transition characteristics; Obtain information on marketing content to be broadcast, and based on the phased emotion 3D dataset and the emotion conversion feature dataset, determine the digital human emotion conversion strategy according to the information on marketing content to be broadcast; Based on the marketing content information to be broadcast and the digital human emotion conversion strategy, generate and output digital human broadcast content; The marketing broadcast phase includes a warm-up phase, a product introduction phase, a product demonstration phase, a promotion phase, and a summary phase. The three dimensions of emotion include excitement, affinity, and professionalism.
2. The method according to claim 1, characterized in that, The step involves performing a three-dimensional emotional analysis on different marketing broadcast stages based on the training video material set to determine a stage-specific three-dimensional emotional dataset, including: Analyze the training video material set and extract the visual information, audio information and text information corresponding to each training video material; Based on the preset time window length, each training video material is divided into time windows to determine several video segments; Based on the visual information, extract the background switching frequency and gesture density of each video segment; Based on the audio information, extract the real-time speech rate and real-time volume of each video segment; Based on a preset stage keyword dictionary, the set of stage keywords for each video segment is extracted according to the text information; Based on the set of stage keywords, the stage mapping index corresponding to each video segment is determined according to the background switching frequency, gesture density, real-time speech rate and real-time volume. Based on the stage mapping index corresponding to each video segment, each training video material is divided into marketing stages to determine the stage video material set. Analyze the video material collection for each stage to determine the three-dimensional emotional index information corresponding to each marketing stage; Based on the set of video materials for each stage and the three-dimensional emotion index information, the stage-specific three-dimensional emotion dataset is constructed.
3. The method according to claim 2, characterized in that, Based on the set of stage keywords, and according to the background switching frequency, gesture density, real-time speech rate, and real-time volume, the stage mapping index corresponding to each video segment is determined, specifically by the following formula: ; in, For the current time window The stage mapping index corresponding to the next video segment. This is the current evaluation phase. For marketing phases, Weighting the text. For the first The stage-specific influence weight of keywords. For the first Key words for each stage For the first Keyword frequency at each stage The weighting of audio influences For the current time window The audio impact factor below, Weighting based on visual impact. Current time window The visual impact factor below Speech rate affects weighting. For the current time window The real-time speech rate of the next video segment The average speaking speed of the current video. The standard deviation of the current video's speech rate. The weighting is affected by the volume. For the current time window Real-time volume of the next video segment The average volume of the current video. The standard deviation of the current video volume. The background influences the weight. Current time window The background switching frequency mentioned in the following video segment, Gestures affect weight. For the current time window The density of the gestures in the following video segment.
4. The method according to claim 2, characterized in that, The analysis of the video material set for each stage determines the three-dimensional sentiment index information corresponding to each marketing stage, including: Based on the image motion analysis model, the video material set of the aforementioned stages is analyzed to determine the average acceleration of gestures and the posture stability index for each marketing stage. Based on a deep learning facial recognition model, the video material set of the aforementioned stages is analyzed to determine the facial affinity index for each marketing stage; Based on the speech processing model, the video material set of the aforementioned stages is analyzed to determine the standard deviation of the speech fundamental frequency, the speech affinity index, and the density of professional terms for each marketing stage. Based on the average acceleration of the gesture, the posture stability index, the facial expression affinity index, the standard deviation of the voice fundamental frequency, and the voice affinity index, the excitement index, professionalism index, and affinity index for each marketing stage are determined respectively, and the three-dimensional emotion index information is constructed, specifically by the following formula: ; in, The excitement index is... The fundamental frequency influences the weight. To influence the weights of acceleration, The standard deviation of the fundamental frequency of the speech is... The average acceleration of the gesture. The so-called professionalism index, The attitude affects the weights. Let be the attitude stability index. The terminology affects the weight. For the density of the aforementioned technical terms, The affinity index is... The weighting is affected by facial expressions. The facial expression affinity index is... The weighting is affected by voice. The voice affinity index is defined as follows.
5. The method according to claim 4, characterized in that, The analysis of the stage-based three-dimensional emotion dataset determines the emotion transition feature dataset, including: Based on the stage mapping index corresponding to each video segment in the staged emotion 3D dataset, several emotion transition time nodes are determined. Based on several emotional transition time points, the video material set of the stage is analyzed, and the visual information, voice information and text information of the video material corresponding to each emotional transition time point within the preset emotional transition focus range are extracted to construct the emotional transition feature dataset.
6. The method according to claim 5, characterized in that, The determination of several emotion transition time nodes based on the stage mapping index corresponding to each video segment in the staged emotion 3D dataset is specifically formulated as follows: ; in, This refers to the time point of the emotional transition. For the current time window The stage mapping index corresponding to the next video segment. To preset the mutation threshold, For the current time window The intermediate time point.
7. The method according to claim 4, characterized in that, The process of determining a digital human emotion conversion strategy based on the staged emotion 3D dataset and the emotion conversion feature dataset, and according to the marketing content information to be broadcast, includes: Analyze the marketing content information to be broadcast, determine the content to be broadcast corresponding to different marketing stages, and determine the transition time range between different marketing stages based on the content to be broadcast. Analyze the emotion transition feature dataset to determine the amplitude ranges for facial expression transitions, action transitions, tone transitions, and text sentiment polarity transitions. Based on the transition time range, and constrained by the range of facial expression transition amplitude, the range of action transition amplitude, the range of tone transition amplitude, and the range of text emotional polarity transition amplitude, the three-dimensional feature values of the target emotion corresponding to different time points within the transition time range are determined. The digital human emotion conversion strategy is constructed based on the three-dimensional feature values of the target emotion at different time points.
8. The method according to claim 7, characterized in that, The determination of the target emotion three-dimensional feature value corresponding to different time points within the transition time range is specifically made by the following formula: ; in, The current time point The corresponding three-dimensional feature values of the target emotion, For the current number Three-dimensional feature values of emotion For the first The preset time influence coefficient corresponding to the three-dimensional feature values of emotion. This refers to the starting time point of the transition time range. This refers to the end time point of the transition time range. The first corresponding to the current marketing stage Three-dimensional sentiment index.
9. A computer vision-based digital human broadcasting platform, characterized in that, The method applied to any one of claims 1-8 includes: The emotion feature analysis module is used to acquire a training video material set, and to perform three-dimensional emotion analysis on different marketing broadcast stages based on the training video material set to determine the stage-specific three-dimensional emotion dataset. The emotion transition analysis module is used to analyze the staged emotion three-dimensional dataset and determine the emotion transition feature dataset; The strategy analysis module is used to acquire information on marketing content to be broadcast, and based on the phased emotion 3D dataset and the emotion conversion feature dataset, determine the digital human emotion conversion strategy according to the information on marketing content to be broadcast. The output module is used to generate and output digital human broadcast content based on the marketing content information to be broadcast and the digital human emotion conversion strategy.
Citation Information
Patent Citations
News intelligent broadcasting method, device and apparatus and storage medium
CN112541078A
Emotion expression method, system and equipment for voice-driven 3D virtual human
CN119418722A