Digital human interaction strategy optimization method based on big data identification
By collecting multimodal interaction data and constructing a user emotion change trajectory map using a pre-trained emotion recognition model, combined with historical interaction big data evaluation and ranking, the problem of the difficulty in adaptively adjusting interaction strategies in digital human interaction technology is solved, thereby improving users' emotional response capabilities and interaction effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING XINZHI ART TESTING TECH CO LTD
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-08
AI Technical Summary
Existing digital human interaction technologies struggle to adapt their interaction strategies to changes in user emotional states in high-frequency, multi-turn, and emotionally sensitive interaction scenarios, resulting in rigid interaction styles, unbalanced response rhythms, and a decline in user experience.
By collecting multimodal interaction data, constructing user emotion feature vectors using a pre-trained emotion recognition model, generating emotion change trajectory maps, and combining historical interaction big data for evaluation and ranking, adaptive optimization of interaction strategies is achieved, and strategy evaluation indicators are continuously corrected through user feedback mechanisms.
It enhances the emotional responsiveness and interaction strategy rationality of digital humans in complex interaction scenarios, realizes dynamic optimization and continuous iteration of interaction strategies, and improves user satisfaction.
Smart Images

Figure CN121996341A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent interaction and digital human technology, and in particular to a method for optimizing digital human interaction strategies based on big data recognition. Background Technology
[0002] With the rapid development of artificial intelligence, big data analytics, and multimodal perception technologies, digital humans, as intelligent interactive carriers integrating natural language processing, speech recognition, computer vision, and affective computing, have been widely applied in scenarios such as intelligent customer service, virtual assistants, digital government, digital education, and immersive interaction. Current digital human interaction technology has evolved from early rule-driven and template-based single-turn dialogue modes to context-aware interaction modes that integrate text, voice, and facial expressions, achieving some progress in natural interaction and complete information expression. However, most existing digital human systems focus on improving the accuracy of "single-interaction content generation," paying more attention to semantic understanding, multimodal feature fusion, and the matching degree of output content, while neglecting the long-term statistical and dynamic modeling of the user's emotional evolution, emotional stability, and strategy response effects during continuous interaction. Especially in high-frequency, multi-turn, and emotionally sensitive interaction scenarios, digital humans often struggle to adaptively adjust their interaction strategies according to changes in the user's emotional state, leading to rigid interaction styles, unbalanced response rhythms, and even problems such as emotional intensification or a decline in user experience. Therefore, how to identify and model user emotional changes based on multimodal interaction data using big data analytics, and thereby dynamically optimize interaction strategies, has become a key issue that urgently needs to be addressed in the field of digital human intelligent interaction technology.
[0003] CN119299805B discloses a "Big Data-Driven Intelligent Interaction Method and System for Digital Humans." This technical solution extracts interactive audio and video frame sets from historical user interaction videos, performs image smoothing and equalization processing on the video frames, and combines this with speech recognition results to extract the temporal information of enhanced video frames, interactive text, and interactive audio. Based on this, multimodal features are interactively fused to generate target interactive text, interactive voice, and interactive emoticons, ultimately constructing an interactive video of the digital human to achieve intelligent interaction with the user. This solution, from the perspective of multimodal perception and feature fusion, improves the digital human's ability to understand user input information, and to a certain extent enhances the accuracy and consistency of interactive content and expression, providing strong support for the implementation of digital human interaction technology in real-world application scenarios.
[0004] However, the technical solution disclosed in CN119299805B focuses primarily on the synchronous extraction and fusion modeling of multimodal features, and the optimization of single-turn or partial interactive content generation effects. It fails to explicitly model and analyze changes in the user's emotional state throughout consecutive dialogue rounds. Specifically, while the solution can identify multimodal information such as text, voice, and facial expressions, it does not quantify key emotional parameters such as user emotional polarity, intensity, and stability. Furthermore, it does not construct a trajectory model reflecting the evolution of user emotions throughout the interaction process, making it difficult to identify patterns of emotional transition from stable to fluctuating, or from positive to negative. At the interaction strategy level, this existing technology primarily improves expression matching by generating target interactive content and facial expressions. It lacks a strategy effectiveness evaluation mechanism based on historical interaction big data, cannot horizontally compare and rank the actual performance of different interaction strategies in scenarios with similar emotional changes, and does not address the issue of continuously updating strategy evaluation indicators through user feedback. This makes it difficult for digital humans to achieve true strategy self-optimization when facing complex emotion-driven interaction scenarios, and their interaction behavior remains somewhat static and passive.
[0005] In summary, existing digital human intelligent interaction technologies generally suffer from insufficient granularity in recognizing user emotional changes, a lack of emotional evolution modeling mechanisms, and difficulty in dynamically optimizing interaction strategies based on big data. This invention addresses these shortcomings by providing a digital human interaction strategy optimization method based on big data recognition. It collects and integrates multimodal interaction data, including text, voice, facial expressions, and duration, and uses a pre-trained emotion recognition model to construct user emotion feature vectors. Furthermore, it generates a user emotion change trajectory map to identify emotional state transition patterns in continuous interactions. Based on this, it evaluates and ranks the actual effects of different interaction strategies using historical interaction big data, achieving adaptive optimization of interaction strategies. A user feedback mechanism continuously refines the strategy evaluation indicators, thereby significantly improving the emotional responsiveness and rationality of digital humans in complex interaction scenarios. This invention belongs to the field of intelligent interaction technology, specifically involving digital human interaction, big data analysis, and emotion computing, providing an engineering-practical technical path for achieving highly adaptable and highly continuous digital human intelligent interaction. Summary of the Invention
[0006] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract and title of the invention. Such simplifications or omissions shall not be used to limit the scope of the present invention.
[0007] Given that existing digital human interaction technologies generally focus on single-turn interaction content generation and multimodal feature fusion, lack systematic modeling of the evolution of emotional states during continuous user interaction, and that interaction strategies are difficult to evaluate and dynamically optimize based on historical big data, resulting in insufficient stability and adaptability of the interaction experience in complex, multi-turn, and emotionally sensitive interaction scenarios, this invention is proposed.
[0008] Therefore, the problem to be solved by this invention is how to accurately identify and model the user's emotional state and its changing patterns based on multimodal interaction data, and combine historical interaction big data to realize the evaluation, ranking and adaptive optimization of digital human interaction strategies, thereby improving the rationality of the emotional response of digital humans in continuous interaction scenarios and the overall interaction effect.
[0009] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, embodiments of the present invention provide a method for optimizing digital human interaction strategies based on big data recognition, comprising, Collect multimodal interaction data generated by users during digital human interaction and store it in the interaction database; The multimodal interaction data is labeled with emotional state based on a pre-trained emotion recognition model to extract the emotion feature vector of the user's current dialogue round. Construct a trajectory map of user emotional changes, and identify the transition patterns of user emotional states by calculating the difference in emotional feature vectors between consecutive dialogue rounds; According to the conversion mode, a set of candidate interaction strategies is retrieved from the strategy library, the strategy evaluation index of the candidate interaction strategies is calculated, and the set of candidate interaction strategies is sorted according to the strategy evaluation index. The candidate interaction strategy with the best ranking is selected as the optimized interaction strategy. The interaction strategy parameters of the optimized interaction strategy are transmitted to the digital human execution terminal, user feedback data on the optimized interaction strategy is collected, and the feedback data is written back to the interaction database to update the strategy evaluation indicators.
[0010] Secondly, embodiments of the present invention provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any step of the above-described digital human interaction strategy optimization method based on big data recognition.
[0011] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the above-described digital human interaction strategy optimization method based on big data recognition.
[0012] Compared with existing technologies, the advantages of this invention are as follows: By collecting multimodal interaction data and storing it in an interaction database, a complete user interaction behavior profile is established, achieving comprehensive recording and traceability of the interaction process and improving the reliability of data-driven decision-making; by using a pre-trained emotion recognition model to label the multimodal data with emotional states and extract emotional feature vectors, unstructured interaction information is transformed into quantifiable emotion indicators, enabling real-time and objective identification of user emotional states and providing accurate input for dynamic emotion analysis; by constructing a user emotion change trajectory map and calculating the difference in emotional feature vectors between consecutive dialogue rounds, the user's emotional state can be dynamically captured. The evolution patterns and transformation modes enable quantitative tracking of emotional trends, helping to identify key points of emotional transition and providing a basis for personalized strategy matching. Based on emotional transformation modes, candidate interaction strategies are retrieved from the strategy library, and historical interaction records with the same patterns are used to calculate strategy evaluation indicators. This achieves dynamic optimization of strategy recommendations and fusion of historical experience. A quantitative ranking mechanism ensures the selection of the optimal strategy, improving the adaptability and effectiveness of interaction strategies. By transmitting optimized strategy parameters to the digital human execution end and collecting user feedback data to write back to the database, continuous iteration and adaptive adjustment of interaction strategies are achieved, ultimately improving the emotional adaptability and user satisfaction of digital human interaction. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart of a method for optimizing digital human interaction strategies based on big data recognition. Detailed Implementation
[0014] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0015] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort should fall within the scope of protection of this invention.
[0016] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0017] As mentioned in the background section, existing digital human intelligent interaction technologies generally suffer from problems such as insufficient granularity in recognizing user emotional changes, lack of emotional evolution modeling mechanisms, and difficulty in dynamically optimizing interaction strategies based on big data. To address these issues, this invention provides a digital human interaction strategy optimization method based on big data recognition.
[0018] Reference Figure 1 , Figure 1 This is a flowchart illustrating a digital human interaction strategy optimization method based on big data recognition according to an embodiment of the present invention. Figure 1 As shown, a digital human interaction strategy optimization method based on big data recognition includes: S1: Collect multimodal interaction data generated by users during digital human interaction and store it in the interaction database.
[0019] It should be noted that multimodal interaction data includes text input content, speech frequency feature parameters, facial expression recognition coordinates, and interaction duration parameters. A dialogue turn is defined as the process from "the user completing one input (including text, speech, etc.) to the digital human providing a corresponding response" within a single interaction session. A globally unique dialogue turn number is generated for each dialogue turn, and this number increments chronologically within the session. This definition serves as the basic time unit for subsequent data sharding, sentiment analysis, and strategy optimization.
[0020] S1.1: When a user initiates an interactive session with the digital human, the interactive system generates a unique session identifier, which includes the user's identity code, the session start timestamp, and the interactive terminal type marker.
[0021] S1.2: Capture the text input content entered by the user through the text input interface of the interactive terminal, and segment and mark the text input content according to the input time sequence. Each segment of text input content is attached with a corresponding input start and end timestamp, forming a text data packet with time sequence labeling. The text input content includes the original character sequence typed by the user, input method candidate word switching records, and text editing modification traces.
[0022] S1.3: Acquire the user's voice signal stream through the microphone array of the interactive terminal, perform preprocessing operations on the voice signal stream, and extract voice frequency feature parameters from the preprocessed voice signal stream.
[0023] Preferably, the preprocessing operations include removing environmental background noise, endpoint detection to segment effective speech segments, and signal normalization; the speech frequency characteristic parameters include the fundamental frequency mean, fundamental frequency standard deviation, formant frequency distribution, speech rate value, and speech energy envelope curve.
[0024] S1.4: The user's facial image sequence is acquired at a fixed frame rate through the camera module of the interactive terminal. Face detection and key point localization are performed on each frame of the facial image sequence, facial expression recognition coordinates are extracted, and the facial expression recognition coordinates corresponding to each frame of the image are organized into a temporal coordinate matrix. S1.5: Record the interaction duration parameters of the user in a single interaction session. The timing module of the interactive terminal collects the above-mentioned durations with millisecond-level precision and binds the interaction duration parameters with the corresponding dialogue round number. It should be noted that the facial expression recognition coordinates include the inner and outer endpoints of the eyebrows, the vertex of the eye contour, the nose tip positioning point, the left and right endpoints of the corners of the mouth, and the control point of the jaw contour; the interaction duration parameters include the total duration of the conversation, the waiting time for a single-turn dialogue response, the duration for the user to read the digital human's reply, and the interval duration for continuous user input.
[0025] S1.6: Perform timeline alignment on the text input content, speech frequency feature parameters, facial expression recognition coordinates, and interaction duration parameters. Using the session identifier as the primary key, encapsulate the aligned multimodal interaction data into a structured data record. The structured data record includes a data type field, a collection time field, a data content field, and a quality assessment field. S1.7: Perform data integrity verification on structured data records, detect whether there are missing segments or abnormal values in multimodal interaction data, fill the detected missing segments with adjacent valid data through interpolation, mark the detected abnormal values and record the abnormality type label, and generate the verified multimodal interaction data record. Specifically, the interpolation and imputation operation needs to be performed differently depending on the data type: for continuous time-series numerical data (such as interaction duration parameters, speech energy envelope curves), linear or spline interpolation is used; for discrete or categorical data (such as text fragments, sudden loss of key point coordinates), it is marked as missing and forward imputation or assigned a system default value is used, while the imputation type is recorded in the quality assessment field.
[0026] S1.8: Write the verified multimodal interaction data records into the interaction database; Preferably, the interactive database establishes a partitioned storage structure according to data modality type. Text input content is stored in the text data partition, speech frequency feature parameters are stored in the speech data partition, facial expression recognition coordinates are stored in the visual data partition, and interaction duration parameters are stored in the behavior data partition. Each partition establishes an associated index through a session identifier.
[0027] S2: Based on the pre-trained emotion recognition model, perform emotion state annotation processing on the multimodal interaction data and extract the emotion feature vector of the user's current dialogue round; It should be noted that the emotional feature vector includes emotional polarity values, emotional intensity coefficients, and emotional stability indicators.
[0028] S2.1: Read the multimodal interaction data corresponding to the current dialogue round from the interaction database, retrieve the associated records according to the session identifier, and load them into the memory buffer of the sentiment analysis and processing module; S2.2: Perform feature extraction on the multimodal interaction data to extract multimodal sentiment representation vectors, which include text sentiment feature sub-vectors, speech sentiment feature sub-vectors, visual sentiment feature sub-vectors, and behavioral sentiment feature sub-vectors; S2.2.1: Perform text sentiment feature extraction on the text input content. The text input content is converted into a text semantic embedding vector through a pre-trained text encoder. The text encoder adopts a multi-layer bidirectional attention network structure to weight and encode sentiment words, negation words and degree adverbs in the text input content, and outputs a text sentiment feature sub-vector with fixed dimension. S2.2.2: Perform speech emotion feature extraction operation on speech frequency feature parameters. Input the fundamental frequency mean, fundamental frequency standard deviation, formant frequency distribution, speech rate value and speech energy envelope curve into the pre-trained speech emotion encoder. The speech emotion encoder performs feature mapping on the temporal change pattern of speech frequency feature parameters through a one-dimensional convolutional layer and outputs speech emotion feature sub-vectors. S2.2.3: Perform visual emotion feature extraction operation on facial expression recognition coordinates, calculate the displacement and displacement velocity of each key point between adjacent frames in the temporal coordinate matrix, and use the vertical displacement of the inner and outer ends of the eyebrows, the change in the opening and closing of the eye contour apex, and the horizontal displacement of the left and right ends of the corners of the mouth as facial expression action features input to the pre-trained facial expression emotion encoder, where the facial expression emotion encoder outputs visual emotion feature sub-vectors. S2.2.4: Perform behavioral sentiment feature extraction on the interaction duration parameter, calculate the ratio of single-turn dialogue response waiting time to historical average response time as the response latency index, calculate the variance of the interval duration of continuous user input as the input rhythm fluctuation value, and combine the response latency index and the input rhythm fluctuation value into a behavioral sentiment feature sub-vector.
[0029] S2.2.5: Input the text sentiment feature sub-vectors, speech sentiment feature sub-vectors, visual sentiment feature sub-vectors, and behavioral sentiment feature sub-vectors into the multimodal fusion layer of the pre-trained sentiment recognition model. The multimodal fusion layer uses a cross-modal attention mechanism to dynamically assign weights to each sub-vector, calculates the correlation coefficient matrix between each modality, and performs a weighted summation of each sub-vector based on the correlation coefficient matrix to generate the fused multimodal sentiment representation vector.
[0030] Furthermore, the specific formula for the multimodal sentiment representation vector is as follows: ; in, This is the final generated multimodal sentiment representation vector; i and j are modality indices, representing text 1, speech 2, visual 3, and behavioral 4 modalities, respectively; Let be the sentiment feature sub-vector of the i-th modality; is the correlation coefficient between the i-th modality and the current dialogue context (represented by a context encoding vector Vetex); The correlation temperature coefficient is a trainable or preset scalar parameter greater than 0, used to sharpen or smooth the weight distribution. Let i be the data quality confidence score for the i-th modality in the current round. For example, for the visual modality, It can be calculated based on the bounding box confidence of face detection and the inverse of the variance of key point localization; This is the Sigmoid function, used to map confidence scores to the (0,1) interval and perform adjustments before normalization.
[0031] S2.3: Input the multimodal sentiment representation vector into the pre-trained sentiment recognition model, and output the sentiment feature vector, specifically including: The sentiment classification layer includes a sentiment polarity prediction branch, a sentiment intensity prediction branch, and a sentiment stability prediction branch. The sentiment polarity prediction branch maps multimodal sentiment representation vectors through a fully connected network and outputs a sentiment polarity value. The sentiment polarity value ranges from negative one to positive one, with negative values indicating a negative sentiment tendency and positive values indicating a positive sentiment tendency. The emotion intensity prediction branch calculates the mean absolute value of the activation values of each dimension in the multimodal emotion representation vector as the raw intensity score, and maps the raw intensity score to the interval between zero and one through a normalization function, outputting the emotion intensity coefficient. The closer the emotion intensity coefficient is to one, the stronger the user's emotional expression. The emotion stability prediction branch retrieves historical emotion polarity value records from the interaction database for the three consecutive rounds preceding the current dialogue round. It calculates the standard deviation between the historical emotion polarity value records and the current emotion polarity value as the emotion fluctuation amplitude. After taking the reciprocal of the emotion fluctuation amplitude and normalizing it, it outputs the emotion stability index. The closer the emotion stability index value is to one, the more stable the user's emotional state is. The emotional polarity value, emotional intensity coefficient, and emotional stability index are concatenated in a fixed order to form the emotional feature vector of the user's current dialogue round. The emotional feature vector is a three-dimensional numerical vector. After binding the emotional feature vector with the current dialogue round number and the session identifier, it is written into the emotional feature storage table of the interaction database.
[0032] S3: Construct a trajectory map of user emotional changes, and identify the transition patterns of user emotional states by calculating the difference in emotional feature vectors between consecutive dialogue rounds; S3.1: From the sentiment feature storage table of the interaction database, retrieve all historical dialogue round records of the current interaction session according to the session identifier, sort them in ascending order of dialogue round number, extract the sentiment feature vector corresponding to each round, and form a temporal sequence of sentiment feature vectors. Each element in the temporal sequence of sentiment feature vectors contains the sentiment polarity value, sentiment intensity coefficient and sentiment stability index of the corresponding round.
[0033] S3.2: Perform difference calculation on the sentiment feature vectors of two adjacent dialogue rounds in the temporal sequence of sentiment feature vectors to obtain the sentiment change difference vector, specifically including: The change in emotional polarity is obtained by subtracting the emotional polarity value of round N from the emotional polarity value of round N+1; the change in emotional intensity is obtained by subtracting the emotional intensity coefficient of round N from the emotional intensity coefficient of round N+1; and the change in emotional stability is obtained by subtracting the emotional stability index of round N from the emotional stability index of round N+1. The change in emotional polarity, the change in emotional intensity, and the change in emotional stability are combined into an emotional change difference vector, which represents the direction and magnitude of the user's emotional state transition between two adjacent dialogue rounds. The above difference calculation operation is repeated for all adjacent round pairs in the emotional feature vector time sequence to generate an emotional change difference vector sequence.
[0034] Furthermore, before performing pattern encoding, a preset set of threshold parameters needs to be loaded. These thresholds are determined based on large-scale analysis of historical interaction data. For example, the polarity change threshold can be taken as the 75th percentile of the absolute value of historical emotional polarity changes; the upper and lower thresholds for intensity changes can be taken as the 66th and 33rd percentiles of the absolute value of historical emotional intensity changes, respectively; and the stability change threshold can be set to 0.1 based on experience. All thresholds support dynamic calibration on the system management end based on new data distribution.
[0035] S3.3: Perform pattern encoding on the emotion change difference vector sequence to determine the polarity migration direction label, intensity change amplitude label, and stability trend label, and generate a single-step transformation pattern encoding sequence, specifically including: The polarity migration direction label is determined based on the positive or negative sign of the change in emotional polarity. The polarity migration direction label includes three types: positive migration, negative migration, and stable. If the change in emotional polarity is greater than the preset polarity change threshold, it is marked as positive migration. If the change in emotional polarity is less than the negative value of the polarity change threshold, it is marked as negative migration. In other cases, it is marked as stable. The intensity change range label is determined based on the absolute value of the change in emotional intensity. The intensity change range label includes three types: severe change, moderate change, and slight change. If the absolute value of the change in emotional intensity is greater than the preset upper threshold of intensity change, it is marked as severe change. If the absolute value of the change in emotional intensity is between the upper threshold of intensity change and the preset lower threshold of intensity change, it is marked as moderate change. If the absolute value of the change in emotional intensity is less than the lower threshold of intensity change, it is marked as slight change. The stability trend label is determined based on the positive or negative sign and value of the change in emotional stability. The stability trend label includes three types: tending to be stable, tending to fluctuate, and remaining unchanged. If the change in emotional stability is greater than the preset stability change threshold, it is marked as tending to be stable. If the change in emotional stability is less than the negative value of the stability change threshold, it is marked as tending to fluctuate. In other cases, it is marked as remaining unchanged. The polarity migration direction label, intensity change amplitude label, and stability trend label are combined to form a single-step transition pattern code. The single-step transition pattern code is a triple structure. For all elements in the emotional change difference vector sequence, the corresponding single-step transition pattern code is generated to form a single-step transition pattern code sequence. The single-step transition pattern code sequence reflects the gradual evolution of the user's emotional state during continuous dialogue.
[0036] S3.4: Construct a user emotion change trajectory map based on the single-step conversion mode encoding sequence; Specifically, the user emotion change trajectory map adopts a directed graph structure. The nodes in the graph represent discrete emotion state categories, and the directed edges in the graph represent the transformation relationship between emotion states. Each directed edge is accompanied by a single-step transformation pattern encoding as an edge attribute, and the number of times this transformation occurs in the current session is recorded as the edge weight.
[0037] Furthermore, after the user emotion change trajectory map is constructed, it is mainly used for long-term conversation analysis and visualization insights; for real-time strategy optimization, the main basis is the latest generated single-step transformation pattern encoding sequence in the current conversation, and the continuous pattern recognition operation in step S3.5 is performed on this sequence.
[0038] S3.5: Perform continuous pattern recognition on the user's emotional change trajectory map, and extract continuous single-step conversion pattern encoding segments on the single-step conversion pattern encoding sequence using a sliding window, and splice the single-step conversion pattern encodings in each segment in sequence to form a composite conversion pattern encoding. S3.6: Match the composite conversion pattern encoding with the preset typical emotion conversion pattern library, calculate the pattern similarity score, and select the conversion pattern category corresponding to the sample with the highest pattern similarity score as the conversion pattern; It should be noted that the composite transition pattern encoding characterizes the trajectory of a user's emotional evolution in a multi-turn continuous dialogue; the typical emotional transition pattern library stores historical transition pattern samples labeled with the interaction risk level, and the transition patterns are written into the transition pattern record table of the interaction database after being bound to the session identifier.
[0039] S4: Retrieve a set of candidate interaction strategies from the strategy library according to the conversion mode, calculate the strategy evaluation index of the candidate interaction strategies, sort the set of candidate interaction strategies according to the strategy evaluation index, and select the candidate interaction strategy with the best ranking as the optimized interaction strategy. S4.1: Retrieve historical interaction records with the same conversion pattern from the interaction database, and extract a subset of historical application records that match each candidate interaction strategy, specifically including: The polarity migration direction label, intensity change magnitude label, and stability trend label are combined to form the conversion mode retrieval key, which serves as the main index condition for subsequent strategy library query operations. Access the preset strategy library based on the conversion mode search key, retrieve all interactive strategy entries in the strategy library that match the conversion mode search key and form a candidate interactive strategy set. The strategy library stores interactive strategy entries organized by conversion mode category. Each interactive strategy entry contains a unique strategy identifier, applicable conversion mode range, strategy parameter configuration, and strategy creation timestamp. A preliminary screening operation is performed on the candidate interaction strategy set. The system reads the current user's historical interaction preference records, which include the user's acceptance rating for different language styles, the user's satisfaction rating for different response speeds, and the user's click-through rate data for different content types. Strategy entries in the candidate interaction strategy set that clearly conflict with the historical interaction preference records are removed, generating a filtered candidate interaction strategy set. The rules for determining obvious conflicts need to be predefined. For example, if the user's acceptance rating for a certain "language style" is consistently below 2 points (out of 5), then all strategy entries that force the use of that language style are removed. In addition, for cold start scenarios (such as when the candidate strategy set is empty, or when the number of times all strategies have been applied in the past is below the minimum confidence threshold), the system should enable the default strategy generation module. This module maps to a set of predefined, conservative baseline interaction strategies based on the conversion mode retrieval key to ensure that the interaction can continue, while collecting initial feedback data in this scenario.
[0040] The historical interaction record table in the interaction database is queried according to the conversion mode search key. Historical interaction records with the same conversion mode are retrieved. The historical interaction records include historical session identifiers, historical application policy identifiers, historical interaction result tags, and historical user feedback ratings. All retrieved historical interaction records are grouped and aggregated according to the historical application policy identifiers.
[0041] For each candidate interaction strategy in the filtered candidate interaction strategy set, extract a subset of historical application records that match the unique strategy identifier of that candidate interaction strategy from the grouped and aggregated historical interaction records.
[0042] S4.2: Count the total number of records in the historical application record subset as the number of times the strategy is applied in history, and calculate the percentage of records in the historical application record subset with positive historical interaction result tags as the historical success rate of the strategy. S4.3: Calculate the arithmetic mean of all historical user feedback scores in the historical application record subset as the strategy average feedback score, and at the same time calculate the standard deviation of historical user feedback scores as the strategy feedback stability index. S4.4: Calculate the time decay weight based on the number of days between the creation time of the historical interaction record and the current time. Multiply the historical user feedback rating by the time decay weight, sum the weighted values, and then divide by the sum of all time decay weights to obtain the strategy timeliness weighted score. S4.5: Multiply the historical success rate of the strategy, the average feedback score of the strategy, the stability index of the strategy feedback, and the timeliness-weighted score of the strategy by their respective dimension weight coefficients, and then sum them to obtain the strategy evaluation index. The specific formula is as follows: ; in, As a strategy evaluation indicator; is the multiplication operator; d is the evaluation dimension index, which in turn represents the historical success rate of the strategy, the average feedback score of the strategy, the inverse normalized value of the strategy feedback stability index, and the time-weighted score of the strategy. The evaluation score for the d-th dimension (all normalized to the [0,1] interval); It is a very small positive number (such as 1e-8) to prevent the entire integral from being zero when a certain dimension is zero; Let be the weight coefficient of the d-th dimension; This is the stability penalty strength coefficient; This is the strategy feedback stability index; the larger the value, the more unstable the strategy effect. Scores in four dimensions The arithmetic mean.
[0043] It should be noted that the strategy evaluation indicators The value range is (0,1). The higher the score, the better and more stable and reliable the strategy is in historical applications. The strategy feedback stability index is normalized to the range of zero to one after taking the reciprocal. The closer the value is to one, the more stable the strategy is in different user groups. The larger the interval days, the smaller the time decay weight. The dimension weight coefficient is pre-configured according to the optimization goal of the current application scenario.
[0044] S4.6: Sort all candidate interaction strategies in the filtered candidate interaction strategy set in descending order according to the numerical value of the strategy evaluation index, and generate a strategy ranking list; S4.7: Select the candidate interaction strategy ranked first from the strategy sorting list as the optimized interaction strategy, read the strategy parameter configuration corresponding to the optimized interaction strategy, bind the unique strategy identifier and strategy parameter configuration of the optimized interaction strategy with the current session identifier, and write them into the strategy execution record table of the interaction database. Preferably, the candidate interaction strategy ranked first in the strategy ranking list has the highest strategy evaluation index value, indicating that the strategy has achieved the best overall effect in historical applications for interaction scenarios with the same conversion mode; the strategy parameter configuration includes the initial value of the language style adjustment factor, the initial value of the response speed control parameter, and the initial value of the content recommendation weight coefficient.
[0045] S5: Transmit the interaction strategy parameters of the optimized interaction strategy to the digital human execution terminal, collect user feedback data on the optimized interaction strategy, and write the feedback data back to the interaction database to update the strategy evaluation indicators. It should be noted that the interaction strategy parameters include language style adjustment factors, response speed control parameters, and content recommendation weight coefficients. Regarding the strategy activation and feedback timeline: In this process, there is a one-round interval between strategy optimization and execution. Specifically, based on all interaction data from round N and earlier, the sentiment conversion pattern is analyzed, and an optimized interaction strategy is generated and distributed for the digital human's response in round N+1. The feedback data generated by the user after interacting with the digital human in round N+1 is an evaluation of the strategy used by the digital human in round N+1. This feedback will be used to update the strategy library, influencing strategy selection when encountering similar patterns in the future.
[0046] S5.1: Read the optimized interaction strategy and strategy parameter configuration bound to the current session identifier from the strategy execution record table of the interaction database, parse the initial value of the language style adjustment factor, the initial value of the response speed control parameter, and the initial value of the content recommendation weight coefficient in the strategy parameter configuration, and encapsulate the above three initial values into an interaction strategy parameter data package; S5.2: Perform parameter validation on the interaction strategy parameter data packet, perform boundary truncation correction on parameter values that exceed the valid range, and generate a validated interaction strategy parameter data packet, specifically including: Verify whether the initial value of the language style adjustment factor is within the preset valid range of language style parameters, verify whether the initial value of the response speed control parameter is within the preset valid range of response speed parameters, and verify whether the initial value of the content recommendation weight coefficient is within the preset valid range of recommendation weight parameters.
[0047] S5.3: Establish a data transmission channel with the digital human execution end, and send the verified interactive strategy parameter data packet to the digital human execution end through the data transmission channel. The data transmission channel adopts an encrypted transmission protocol. After receiving the verified interactive strategy parameter data packet, the digital human execution end returns a parameter reception confirmation signal and records the timestamp of the parameter reception confirmation signal as the time point when the strategy takes effect.
[0048] S5.4: The digital human execution end adjusts the digital human's speech generation module according to the initial value of the language style adjustment factor. The initial value of the language style adjustment factor controls the formality, emotional expression intensity, and vocabulary richness of the digital human's output text. When the initial value of the language style adjustment factor approaches the positive extreme value, the speech generation module outputs a more enthusiastic and conversational response text; when the initial value of the language style adjustment factor approaches the negative extreme value, the speech generation module outputs a more formal and professional response text.
[0049] S5.5: The digital human execution end adjusts the digital human's response timing module according to the initial value of the response speed control parameter. The initial value of the response speed control parameter controls the delay time from receiving user input to outputting response content. When the initial value of the response speed control parameter is large, the response timing module adds an appropriate thinking pause interval to simulate the rhythm of natural conversation. When the initial value of the response speed control parameter is small, the response timing module shortens the response delay to provide an instant feedback experience.
[0050] S5.6: The digital human execution end adjusts the digital human's content filtering module according to the initial value of the content recommendation weight coefficient. The initial value of the content recommendation weight coefficient includes three components: information content weight, emotional support content weight, and guidance content weight. The content filtering module sorts the candidate response content according to the above three weight components and prioritizes outputting the response content corresponding to the content category with the higher weight component.
[0051] S5.7: The digital human execution end generates and outputs interactive response content according to the adjusted parameter configuration, and simultaneously starts the user feedback collection module. The user feedback collection module displays the feedback evaluation entry to the user through the interactive interface. The feedback evaluation entry includes a satisfaction rating control, an interactive experience label selection control, and a text feedback input control.
[0052] S5.8: Collect user satisfaction ratings, experience tag sets, and text feedback content, and encapsulate them into feedback data records. Add the current session identifier, the current dialogue round number, the unique identifier of the optimized interaction strategy, and the feedback submission timestamp to the feedback data records to form structured feedback data records. Furthermore, the user submits a satisfaction rating value through a satisfaction rating control, where the satisfaction rating value ranges from one to five integers; selects a set of experience tags through an interactive experience tag selection control, which includes preset tag options such as timely response, relevant content, friendly attitude, and accurate understanding; and inputs text feedback content through a text feedback input control.
[0053] S5.9: Write the structured feedback data records into the user feedback storage table of the interaction database, retrieve the historical evaluation index records corresponding to the strategy in the interaction database based on the unique strategy identifier in the structured feedback data records, incorporate the satisfaction score into the historical user feedback score sequence of the strategy, and recalculate the average feedback score and the strategy feedback stability index. Specifically, this update optimizes the global evaluation metrics of the general policy in the policy library; in the future, all user sessions that match the same conversion pattern will benefit from this feedback update during policy retrieval and evaluation, thereby achieving continuous evolution based on collective experience.
[0054] S5.10: Update the subdivision evaluation data of this strategy according to the tag type selected by the user in the experience tag set, and write the updated subdivision evaluation data into the strategy evaluation index table of the interaction database; when the user selects the timely response tag, increase the positive evaluation count of the response speed of this strategy; when the user selects the content-related tag, increase the positive evaluation count of the content matching of this strategy. S5.11: Perform sentiment analysis on the text feedback content, extract the number of positive evaluation keywords and negative evaluation keywords in the text feedback content, calculate the ratio of the number of positive evaluation keywords to the total number of keywords as the text feedback positive rate, incorporate the text feedback positive rate as an auxiliary evaluation factor into the calculation formula of the strategy evaluation index, update the value of the strategy evaluation index of the optimized interaction strategy in the strategy library, and complete the closed-loop feedback optimization of the strategy effect.
[0055] In summary, this invention establishes a complete user interaction behavior profile by collecting and storing multimodal interaction data in an interaction database, achieving comprehensive recording and traceability of the interaction process and improving the reliability of data-driven decision-making. By utilizing a pre-trained emotion recognition model to label the multimodal data with emotional states and extract emotion feature vectors, unstructured interaction information is transformed into quantifiable emotion indicators, enabling real-time and objective identification of user emotional states and providing accurate input for dynamic emotion analysis. Furthermore, by constructing a user emotion change trajectory map and calculating the difference in emotion feature vectors between consecutive dialogue rounds, the evolution and transformation patterns of user emotional states can be dynamically captured. By switching modes, quantitative tracking of emotional trends is achieved, which helps identify key points of emotional transition and provides a basis for personalized strategy matching. Based on the emotional transition mode, candidate interaction strategies are retrieved from the strategy library, and historical interaction records of the same mode are retrieved to calculate strategy evaluation indicators. This realizes dynamic optimization of strategy recommendation and fusion of historical experience. The quantitative ranking mechanism ensures the selection of the optimal strategy, improving the adaptability and effectiveness of the interaction strategy. By transmitting the optimized strategy parameters to the digital human execution end and collecting user feedback data and writing it back to the database, continuous iteration and adaptive adjustment of the interaction strategy are realized, ultimately improving the emotional adaptability and user satisfaction of the digital human interaction.
[0056] This embodiment also provides a computer device applicable to the digital human interaction strategy optimization method based on big data recognition, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the digital human interaction strategy optimization method based on big data recognition as proposed in the above embodiment.
[0057] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0058] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the digital human interaction strategy optimization method based on big data recognition as proposed in the above embodiments.
[0059] The storage medium proposed in this embodiment and the data storage method proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0060] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for optimizing digital human interaction strategies based on big data recognition, characterized in that: include, Collect multimodal interaction data generated by users during digital human interaction and store it in the interaction database; The multimodal interaction data is labeled with emotional state based on a pre-trained emotion recognition model to extract the emotion feature vector of the user's current dialogue round. Construct a trajectory map of user emotional changes, and identify the transition patterns of user emotional states by calculating the difference in emotional feature vectors between consecutive dialogue rounds; According to the conversion mode, a set of candidate interaction strategies is retrieved from the strategy library, the strategy evaluation index of the candidate interaction strategies is calculated, and the set of candidate interaction strategies is sorted according to the strategy evaluation index. The candidate interaction strategy with the best ranking is selected as the optimized interaction strategy. The interaction strategy parameters of the optimized interaction strategy are transmitted to the digital human execution terminal, user feedback data on the optimized interaction strategy is collected, and the feedback data is written back to the interaction database to update the strategy evaluation indicators.
2. The digital human interaction strategy optimization method based on big data recognition as described in claim 1, characterized in that: The interaction strategy parameters include a language style adjustment factor, a response speed control parameter, and a content recommendation weight coefficient. The digital human execution terminal adjusts the digital human's speech generation module according to the initial value of the language style adjustment factor; the digital human execution terminal adjusts the digital human's response timing module according to the initial value of the response speed control parameter; the digital human execution terminal adjusts the digital human's content filtering module according to the initial value of the content recommendation weight coefficient; the digital human execution terminal generates and outputs interactive response content according to the adjusted parameter configuration, and simultaneously starts the user feedback collection module.
3. The digital human interaction strategy optimization method based on big data recognition as described in claim 1, characterized in that: The calculation method for the strategy evaluation index is as follows: Retrieve historical interaction records with the same conversion pattern from the interaction database, and extract a subset of historical application records that match each candidate interaction strategy; The total number of records in the historical application record subset is taken as the number of times the strategy has been applied in history, and the percentage of records with positive historical interaction result tags in the historical application record subset is taken as the historical success rate of the strategy. The arithmetic mean of all historical user feedback scores in the historical application record subset is calculated as the strategy average feedback score, and the standard deviation of the historical user feedback scores is calculated as the strategy feedback stability index. The time decay weight is calculated based on the number of days between the creation time of the historical interaction record and the current time. The historical user feedback scores are multiplied by the time decay weight, weighted and summed, and then divided by the sum of all time decay weights to obtain the strategy timeliness weighted score. The strategy evaluation index is obtained by multiplying the historical success rate of the strategy, the average feedback score of the strategy, the feedback stability index of the strategy, and the timeliness weighted score of the strategy by their respective dimension weight coefficients and then summing them.
4. The digital human interaction strategy optimization method based on big data recognition as described in claim 3, characterized in that: The method for obtaining the conversion mode is as follows: From the sentiment feature storage table of the interaction database, retrieve all historical dialogue round records of the current interaction session based on the session identifier, sort them in ascending order of dialogue round number, extract the sentiment feature vector corresponding to each round, and form a temporal sequence of sentiment feature vectors; Perform difference calculation on the emotional feature vectors of two adjacent dialogue rounds in the emotional feature vector time sequence to obtain an emotional change difference vector sequence; Perform pattern encoding on the emotional change difference vector sequence to determine polarity migration direction label, intensity change magnitude label and stability trend label, and generate a single-step conversion pattern encoding sequence; Based on the single-step conversion mode encoding sequence, a user emotion change trajectory map is constructed; Perform continuous pattern recognition on the user emotion change trajectory map, and extract continuous single-step conversion pattern encoding segments on the single-step conversion pattern encoding sequence using a sliding window. Then, concatenate the single-step conversion pattern encodings in each segment in sequence to form a composite conversion pattern encoding. The composite conversion pattern encoding is matched and compared with a preset typical emotion conversion pattern library, the pattern similarity score is calculated, and the conversion pattern category corresponding to the sample with the highest pattern similarity score is selected as the conversion pattern.
5. The digital human interaction strategy optimization method based on big data recognition as described in claim 4, characterized in that: The user emotion change trajectory map adopts a directed graph structure. The nodes in the graph represent discrete emotion state categories, and the directed edges in the graph represent the transformation relationship between emotion states. Each directed edge is accompanied by a single-step transformation pattern encoding as an edge attribute, and the number of times this transformation occurs in the current session is recorded as the edge weight.
6. The digital human interaction strategy optimization method based on big data recognition as described in claim 4, characterized in that: The method for extracting the emotional feature vector is as follows: Feature extraction is performed on the multimodal interaction data to extract multimodal sentiment representation vectors, wherein the multimodal sentiment representation vectors include text sentiment feature sub-vectors, speech sentiment feature sub-vectors, visual sentiment feature sub-vectors, and behavioral sentiment feature sub-vectors. The multimodal emotion representation vector is input into a pre-trained emotion recognition model, which outputs an emotion feature vector.
7. The digital human interaction strategy optimization method based on big data recognition as described in claim 6, characterized in that: The multimodal interaction data includes text input content, speech frequency feature parameters, facial expression recognition coordinates, and interaction duration parameters; the emotion recognition model includes an emotion classification layer; the emotion classification layer includes an emotion polarity prediction branch, an emotion intensity prediction branch, and an emotion stability prediction branch.
8. The digital human interaction strategy optimization method based on big data recognition as described in claim 7, characterized in that: The emotion polarity prediction branch maps the multimodal emotion representation vector through a fully connected network and outputs an emotion polarity value. The emotion intensity prediction branch calculates the mean absolute value of the activation values of each dimension in the multimodal emotion representation vector as the original intensity score, and maps the original intensity score to the zero-to-one interval through a normalization function, outputting an emotion intensity coefficient. The emotion stability prediction branch retrieves historical emotion polarity value records from the interaction database for the three consecutive rounds before the current dialogue round, calculates the standard deviation between the historical emotion polarity value records and the current emotion polarity value as the emotion fluctuation amplitude, takes the reciprocal of the emotion fluctuation amplitude and normalizes it before outputting an emotion stability index.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the digital human interaction strategy optimization method based on big data recognition as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the digital human interaction strategy optimization method based on big data recognition as described in any one of claims 1 to 8.
Citation Information
Patent Citations
A big data driven digital human intelligent interaction method and system
CN119299805B