Communication session interaction method and system based on AI digital person

By extracting the intent-emotion coupling features from multi-turn dialogue interaction units and using a dialogue development trajectory prediction model to generate the optimal response strategy, the problem of insufficient dynamic capture of user intent and emotion changes in existing technologies is solved, thereby improving the coherence and naturalness of AI digital human interaction.

CN120848739BActive Publication Date: 2025-11-25CHENGDU IKE IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511358034.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-11-25
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

Existing AI digital human communication and conversation interaction technologies lack dynamic capture of user intentions and emotional changes, resulting in insufficient interaction continuity and affecting user experience.

Method used

By extracting the intent-emotion coupling features from multi-turn dialogue interaction units, a set of expected interaction trajectories is generated using a dialogue development trajectory prediction model, and the optimal response strategy is selected based on path credibility to construct a structured response sequence.

Benefits of technology

It improves the coherence, accuracy, and naturalness of AI digital human-user communication conversations, enhancing the user's interactive experience during the dialogue process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848739B_ABST
    Figure CN120848739B_ABST
Patent Text Reader

Abstract

The application provides a communication session interaction method and system based on an AI digital person, which extracts multiple rounds of dialogue interaction units from a current communication session interaction flow, performs intention and emotion coupling feature extraction, obtains a user intention feature sequence and an emotion polarity feature sequence corresponding to each dialogue interaction unit, inputs the two sequences into a preset dialogue development track prediction model, generates an expected interaction track set of a current dialogue stage of the user, performs matching processing based on the expected interaction track set and a digital person response strategy library, filters out a target interaction track conforming to a current dialogue scene according to a path credibility index, extracts an optimal response strategy parameter set corresponding to the target interaction track, constructs a structured response sequence of the digital person according to the optimal response strategy parameter set, and converts the structured response sequence into digital person output text. The application can improve the coherence, accuracy and naturalness of the communication session interaction between the AI digital person and the user, and enhance the interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more specifically, to a communication session interaction method and system based on AI digital humans. Background Technology

[0002] With the development of artificial intelligence technology, AI digital human communication and conversational interaction technology has been proposed. This technology constructs digital human models with natural language understanding and generation capabilities to achieve intelligent dialogue and interaction with users, and is widely used in scenarios such as intelligent customer service and virtual assistants. Currently, AI digital human communication and conversational interaction technology typically acquires the user's current input statement, extracts the user's intent through an intent recognition model, and then matches the corresponding response content from a pre-set response template library based on the recognized intent to generate the digital human's response statement. However, this interaction method lacks dynamic capture of the evolution of the user's intent during continuous dialogue, and fails to fully consider the impact of changes in the user's emotional state on the development of the dialogue. This results in the generated response content often being limited to the intent matching of the current dialogue round, making it difficult to adapt to the complex changes in user intent and emotion during the dialogue. Consequently, the interaction between the digital human and the user lacks coherence, affecting the user's dialogue experience. Summary of the Invention

[0003] In view of this, the present invention provides a communication session interaction method and system based on AI digital human.

[0004] According to one aspect of the present invention, a communication session interaction method based on an AI digital human is provided. The method includes: extracting multi-turn dialogue interaction units from the current communication session interaction stream, wherein each multi-turn dialogue interaction unit includes a sequence of input statements by a user during a continuous dialogue process and a sequence of response statements generated by the AI ​​digital human, as well as temporal interaction relationship information between the two sequences; performing intent-emotion coupling feature extraction on the multi-turn dialogue interaction units to obtain a user intent feature sequence and an emotion polarity feature sequence corresponding to each dialogue interaction unit, wherein the user intent feature sequence is used to characterize the evolution process of the intent in the user's input statements, and the emotion polarity feature sequence is used to characterize the user's emotional state during the dialogue process. The user intent feature sequence and the emotional polarity feature sequence are input into a preset dialogue development trajectory prediction model to generate a set of expected interaction trajectories for the current dialogue stage. The set of expected interaction trajectories includes multiple possible dialogue development paths and their corresponding path credibility indices. Based on the set of expected interaction trajectories, a matching process is performed with a digital human response strategy library. Target interaction trajectories that conform to the current dialogue scenario are selected according to the path credibility indices, and the optimal response strategy parameter set corresponding to the target interaction trajectory is extracted. A structured response sequence for the digital human is constructed based on the optimal response strategy parameter set, and the structured response sequence is converted into digital human output text.

[0005] According to another aspect of the present invention, a computer system is provided, comprising: a processor; and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the method as described above.

[0006] This invention provides a communication conversation interaction method based on AI digital humans. By extracting multi-turn dialogue interaction units containing temporal interaction relationship information from the current communication conversation interaction stream, it achieves complete capture of the user's continuous dialogue process, providing a comprehensive dialogue data foundation for subsequent feature extraction. Intent-emotion coupling feature extraction is performed on the multi-turn dialogue interaction units to obtain user intent feature sequences that can respectively characterize the evolution of user intent and emotion polarity change trajectories. User intent and emotion are treated as interrelated dynamic variables, making the extracted features closer to the intrinsic connection between intent and emotion in real dialogue scenarios. The user intent feature sequences and emotion polarity feature sequences are input into a dialogue development trajectory prediction model to generate a set of expected interaction trajectories containing multiple possible dialogue development paths and corresponding path credibility indices. This achieves effective modeling of the uncertainty of dialogue development, providing multi-dimensional decision-making basis for subsequent strategy matching. Matching processing is performed based on the expected interaction trajectory set and the digital human response strategy library. Target interaction trajectories are selected according to the path credibility indices, and the optimal response strategy parameter set is extracted. This ensures that the strategy selection not only adapts to the current dialogue state but also remains consistent with the predicted dialogue development trajectory, improving the dynamic adaptability of strategy matching. By constructing a structured response sequence based on the optimal response strategy parameter set and converting it into digital human output text, the generated response content is more in line with the user's intention evolution and emotional changes in structure and semantics. This improves the overall coherence, accuracy and naturalness of the AI ​​digital human's communication conversation with the user, and enhances the user's interactive experience during the dialogue process. Attached Figure Description

[0007] Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of the present invention.

[0008] Figure 2 This is a schematic diagram illustrating the implementation process of a communication session interaction method based on AI digital humans, provided in an embodiment of the present invention.

[0009] Figure 3 This is a schematic diagram of the hardware entity of a computer system provided in an embodiment of the present invention. Detailed Implementation

[0010] The communication session interaction method based on AI digital humans provided in this invention can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with computer system 104 via a network. A data storage system can store data that computer system 104 needs to process. The data storage system can be integrated into computer system 104 or located in the cloud or on other network servers. Dialogue interaction data can be stored in the local storage of terminal 102, or in the data storage system or cloud storage associated with computer system 104. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, etc. Computer system 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0011] The communication session interaction method based on AI digital humans provided in this embodiment of the invention is applied to... Figure 1 The computer system 104 in the middle specifically includes, for example, Figure 2 The following steps are included:

[0012] Step S100: Extract multi-turn dialogue interaction units from the current communication session interaction stream. The multi-turn dialogue interaction units include the user's input statement sequence and the response statement sequence generated by the AI ​​digital human during the continuous dialogue process, as well as the temporal interaction relationship information between the two sequences.

[0013] The current communication session interaction flow is a collection of all interactive information generated sequentially in time during a complete communication session. It encompasses various user-inputted statements and corresponding responses from the AI ​​digital human, presenting a dynamic and continuous flow of information. The multi-turn dialogue interaction unit is a relatively independent interaction module obtained after dividing the communication session interaction flow. It focuses on a specific part of the continuous dialogue process, integrating the user's input statement sequence and the AI ​​digital human's generated response statement sequence, while also recording the temporal interaction relationship information between these two sequences. The user input statement sequence is a series of statements entered by the user in chronological order during the continuous dialogue, reflecting the user's intentions and needs at different times. The AI ​​digital human's generated response statement sequence is a series of response statements generated by the AI ​​digital human in response to the user's input statements, demonstrating the AI ​​digital human's understanding and feedback of the user's intentions. The temporal interaction relationship information clarifies the temporal order and interaction logic between the user's input statements and the AI ​​digital human's response statements, such as which input statement corresponds to which response statement, and the interval between dialogue turns.

[0014] Step S200: Extract intent-emotion coupling features from the multi-turn dialogue interaction units to obtain the user intent feature sequence and emotion polarity feature sequence corresponding to each dialogue interaction unit. The user intent feature sequence is used to characterize the evolution process of intent in the user's input statement, and the emotion polarity feature sequence is used to characterize the trajectory of the user's emotion polarity change during the dialogue process.

[0015] Intent-emotion coupling feature extraction fuses user intent and emotion information to extract key features reflecting the relationship between user intent and emotion from multi-turn dialogue interaction units. The user intent feature sequence is a sequence of feature vectors arranged chronologically in the dialogue, where each feature vector corresponds to the semantic representation of an intent unit, recording in detail the development and changes of the user's intent during the dialogue, helping the AI ​​digital human better understand the user's needs and goals. The emotion polarity feature sequence is also a sequence arranged chronologically in the dialogue, where each element corresponds to a description of the emotion state in a dialogue turn, reflecting the trajectory of the user's emotions changing to positive, negative, or neutral polarities during the dialogue, providing the AI ​​digital human with important information about the user's emotional state.

[0016] In one implementation, step S200 may include the following steps S210 to S250:

[0017] Step S210: Perform linguistic feature analysis on the user input sentence sequence in the multi-turn dialogue interaction unit, identify the grammatical structure components and semantic dependencies of each input sentence, and generate a linguistic feature map at the sentence level.

[0018] Linguistic feature analysis utilizes linguistic theories and methods to deeply analyze user-input sentences, revealing their inherent grammatical and semantic structures. Grammatical structural components are the grammatical functions and categories of each part of a sentence, such as subject, predicate, object, attributive, and adverbial, which together constitute the basic framework of the sentence. Semantic dependency relations describe the semantic connections and logical relationships between words in a sentence, such as subject-predicate, verb-object, and modification relations, reflecting the actual meaning of the sentence. A sentence-level linguistic feature map is a graphical representation that visually displays the grammatical structural components and semantic dependency relations of a sentence, facilitating subsequent analysis and processing.

[0019] Linguistic feature parsing of user-input sentence sequences can utilize syntactic and semantic analysis techniques from natural language processing. First, part-of-speech tagging tools are used to tag each word in the sentence, determining its grammatical category. Then, a syntactic analyzer is used to parse the sentence, identifying its grammatical structural components and semantic dependencies. Finally, this information is integrated into a graph to generate a sentence-level linguistic feature map.

[0020] Step S220: Based on linguistic feature maps, delineate the boundaries of intent units, mark continuous sentence fragments with complete semantic expression as independent intent units, and extract the core semantic keyword set of each intent unit.

[0021] Intent unit boundary delineation determines the start and end positions of each independent intent unit based on semantic information reflected in linguistic feature maps. A continuous sentence fragment with complete semantic expression refers to a group of sentences that are semantically independent and complete, collectively expressing a clear intent. An independent intent unit is the basic unit for refining and classifying user intents; each independent intent unit represents a specific intent of the user during the dialogue. The core semantic keyword set is a set of representative keywords extracted from each independent intent unit, accurately summarizing the core semantics of that intent unit.

[0022] In one implementation, step S220 may include the following steps S221 to S225:

[0023] Step S221: Traverse the sentence nodes in the linguistic feature map and identify sentence fragments containing explicit intent markers as the starting boundaries of candidate intent units. Explicit intent markers include interrogative words, imperative verbs, and modal particles.

[0024] Explicit intent markers are words that clearly indicate intent, helping to quickly identify key parts of a sentence that express intent. Interrogative words such as "what," "where," and "how much" are typically used to guide questions, indicating the user's intention to obtain information; imperative verbs such as "please," "let," and "help" express a request or command; modal particles such as "can," "should," and "must" reflect the speaker's attitude and willingness towards the action. The starting boundary of a candidate intent unit refers to the position in a linguistic feature map where it might begin as an independent intent unit.

[0025] Step S222: Calculate the semantic correlation between the starting boundary of the candidate intent unit and the subsequent statement node. When the semantic correlation is lower than the preset threshold, it is determined as the ending boundary of the intent unit, and the boundary range label of the complete intent unit is obtained.

[0026] Semantic relevance is an indicator that measures the semantic similarity and relevance between two statements, reflecting the degree of semantic closeness between them. The preset threshold is a pre-defined value used to determine if the semantic relevance is high enough to identify whether two statements belong to the same intent unit. The intent unit end boundary refers to the ending position of an independent intent unit. When the semantic relevance between the starting boundary of a candidate intent unit and subsequent statement nodes is lower than the preset threshold, it indicates that the semantics of that intent unit has ended, and this position is determined as the intent unit end boundary. The boundary range marking of complete intent units clearly defines the start and end positions of each independent intent unit, providing accurate range information for subsequent analysis. Semantic relevance can be calculated using various methods, such as cosine similarity and edit distance.

[0027] Step S223: Perform dependency parsing on the sentence fragments within the boundary range markers, identify the subject-verb-object core structure, and extract the corresponding content words as a preliminary keyword set.

[0028] Dependency parsing is a method that reveals the grammatical structure of a sentence by analyzing the dependency relationships between words. The subject-verb-object core structure is the most basic grammatical structure of a sentence; the subject represents the doer of the action, the verb represents the action, and the object represents the object of the action. Content words are words with actual meaning, such as nouns, verbs, and adjectives, which can accurately express the core semantics of the sentence. The preliminary keyword set is a group of content words extracted from sentence fragments within marked boundaries; it forms the basis for further selection of core semantic keywords.

[0029] Step S224: Input the preliminary keyword set into the pre-trained domain dictionary matching model, filter out professional terms related to the current dialogue domain, merge with the preliminary keyword set to remove duplicates and generate the core semantic keyword set.

[0030] The pre-trained domain dictionary matching model is a model trained on a large amount of domain-related data, capable of accurately identifying professional terms and vocabulary related to a specific domain. The current dialogue domain refers to the topics or areas involved in the user's conversation, such as tourism, shopping, and healthcare. Professional terms and vocabulary are words with specific meanings and usages within a particular domain, capable of more accurately expressing the core information of that domain. The core semantic keyword set is a set of keywords obtained by filtering, merging, and deduplicating keywords based on the initial keyword set, capable of more accurately summarizing the core semantics of independent intent units.

[0031] Step S225: Assign a weight value calculated based on the TF-IDF algorithm to each keyword in the core semantic keyword set, and sort them in descending order of weight values ​​to obtain the core semantic keyword sequence.

[0032] The TF-IDF algorithm (Term Frequency-Inverse Document Frequency) is a statistical method used to evaluate the importance of a word in a collection of documents. Term frequency (TF) is the frequency with which a word appears in a document, reflecting its relative importance. Inverse document frequency (IDF) measures the prevalence of a word across the entire document collection, reflecting its uniqueness. The weight value is a numerical value calculated using the TF-IDF algorithm, taking into account both term frequency and inverse document frequency, thus more accurately reflecting the word's importance within the core semantic keyword set. The core semantic keyword sequence is obtained by arranging the core semantic keywords in descending order of their weight values, highlighting the order of importance of the core semantic keywords.

[0033] Step S230: Perform temporal association analysis on the core semantic keyword set, calculate the intent evolution coefficient based on the semantic similarity and time interval between adjacent intent units, construct the user intent feature sequence based on the intent evolution coefficient, and each feature vector in the user intent feature sequence corresponds to the semantic representation of an intent unit.

[0034] Temporal correlation analysis is a data analysis method that considers the time factor and is used to study the interrelationships and patterns of change in data over time. Semantic similarity is an indicator that measures the degree of semantic similarity between two intent units, reflecting the semantic proximity of adjacent intent units. Time interval refers to the number of dialogue rounds between adjacent intent units, reflecting the time span of intent evolution. Intent evolution coefficient is an indicator that comprehensively considers semantic similarity and time interval, used to describe the evolution of intent between adjacent intent units, such as whether the intent is continuous or has changed. User intent feature sequence is a sequence of feature vectors arranged in the temporal order of dialogue interaction, where each feature vector corresponds to the semantic representation of an intent unit, clearly showing the evolution of user intent during the dialogue process.

[0035] In one implementation, step S230 may include the following steps S231 to S235:

[0036] Step S231: Convert the core semantic keyword set of two adjacent intent units into a word vector space representation, and calculate the semantic similarity value between the two sets.

[0037] Word vector space representation is a method of converting words or sets of words into vectors. It maps the semantic information of words into a high-dimensional vector space, making semantically similar words closer together in the vector space. Semantic similarity score is a numerical value that measures the degree of semantic similarity between two sets of core semantic keywords, reflecting the degree of semantic association between adjacent intent units.

[0038] Converting the core semantic keyword set of two adjacent intent units into a word vector space representation can be achieved using pre-trained word vector models such as Word2Vec and GloVe. First, each keyword is converted into a corresponding word vector. Then, all word vectors in the set are summed or averaged to obtain the vector representation of the set. Next, methods such as cosine similarity and Euclidean distance are used to calculate the similarity score between the two vectors, which serves as the semantic similarity value.

[0039] Step S232: Extract the temporal interaction relationship information in the multi-turn dialogue interaction unit, obtain the number of dialogue turn intervals between adjacent intent units, and convert the number of dialogue turn intervals into a standardized time interval parameter.

[0040] The temporal interaction relationship information records the chronological order and dialogue turn information between user input statements and AI digital human response statements, serving as a crucial basis for analyzing the time factors of intent evolution. The dialogue turn interval refers to the number of dialogue turns between two adjacent intent units, reflecting the time span of intent evolution. The standardized time interval parameter is a value obtained by normalizing the dialogue turn interval, making the time intervals comparable across different dialogue scenarios.

[0041] Extracting temporal interaction information from multi-turn dialogue interaction units can be achieved by recording the timestamp of each statement or the dialogue turn number. Then, the number of dialogue turn intervals between adjacent intent units is calculated based on this information. To ensure the comparability of time intervals across different dialogue scenarios, the number of dialogue turn intervals needs to be converted into a standardized time interval parameter. Linear transformations or other normalization methods can be used to map the number of dialogue turn intervals to a fixed range (e.g., [0, 1]).

[0042] Step S233: Based on the semantic similarity value and the standardized time interval parameter, determine the continuity of intent and generate an intent evolution identifier to identify the relationship state between intent units. When the semantic similarity value is higher than the first threshold and the standardized time interval parameter is lower than the second threshold, it is identified as an intent continuity state; otherwise, it is identified as an intent transition state.

[0043] Intent continuity determination uses semantic similarity values ​​and standardized time interval parameters to determine whether the intent relationship between adjacent intent units is continuous or has transitioned. The first and second thresholds are pre-set values ​​used to judge whether the semantic similarity and time interval meet the conditions for intent continuity. The intent evolution identifier is a symbol or numerical value used to mark the relationship state between intent units, clearly indicating whether adjacent intent units are in a state of intent continuity or intent transition.

[0044] The process of determining intent continuity based on semantic similarity and standardized time interval parameters is as follows: First, the calculated semantic similarity value is compared with a first threshold, and the standardized time interval parameter is compared with a second threshold. If the semantic similarity value is higher than the first threshold and the standardized time interval parameter is lower than the second threshold, it indicates that adjacent intent units are semantically similar and have a short time interval, and this is marked as an intent continuity state; otherwise, it is marked as an intent transition state.

[0045] Step S234: Convert the core semantic keyword set of each intent unit into a fixed-dimensional feature vector. Each dimension of the feature vector corresponds to the weighted sum of the word vector components of a keyword in the keyword set.

[0046] Fixed-dimensional feature vectors are vectors with the same dimension, which facilitates subsequent processing and analysis. Weighted summation of word vector components is the process of summing the word vector components of each keyword in the keyword set according to certain weights. These weights can be determined based on the importance of the keywords (such as TF-IDF weights).

[0047] The following method can be used to convert the core semantic keyword set of each intent unit into a fixed-dimensional feature vector. First, determine the dimension of the feature vector, which is usually chosen based on the actual needs and the characteristics of the dataset. Then, convert each keyword into a corresponding word vector and calculate the weighted sum of the word vector components according to their weights. Finally, combine these weighted sums into a fixed-dimensional feature vector.

[0048] Step S235: Arrange the feature vectors according to the time sequence of the dialogue interaction, and perform logical processing between adjacent feature vectors according to the intent evolution identifier: if the identifier is a continuous intent state, then retain the natural transition between adjacent feature vectors in the feature sequence; if the identifier is a transition intent state, then insert a specific transition feature vector representing the intent transition into the feature sequence to obtain a user intent feature sequence containing intent evolution information.

[0049] The temporal order of dialogue interaction refers to the chronological order in which user input statements and AI digital human response statements appear during the dialogue, serving as the basis for arranging feature vectors. Natural transition refers to the maintenance of the original logical relationship and trend of change between adjacent feature vectors in a continuous intent state, without additional processing. A specific transition feature vector is a predefined vector used to represent intent transitions; inserting this vector clearly demonstrates the changes in intent. The user intent feature sequence is a sequence of feature vectors containing intent evolution information, accurately reflecting the development and changes of user intent during the dialogue process.

[0050] The process of arranging feature vectors according to the chronological order of dialogue interactions and performing logical processing based on intent evolution identifiers is as follows: First, the feature vectors of each intent unit are sorted according to the chronological order of dialogue interactions. Then, the intent evolution identifiers between adjacent feature vectors are checked. If the identifier indicates a continuous intent state, the adjacent feature vectors are directly arranged sequentially in the feature sequence; if the identifier indicates a transitional intent state, a specific transitional feature vector representing the intent transition is inserted between adjacent feature vectors.

[0051] Step S240: Perform sentiment lexical recognition on the user input sentence sequence, extract lexical units containing sentiment tendencies and their position weights in the sentence, and calculate the sentence-level sentiment polarity value by combining the grammatical structure components in the linguistic feature map.

[0052] Sentiment lexical identification uses natural language processing (NLP) technology to identify words with emotional connotations from user input. Lexical units containing sentiment connotations are those that express positive, negative, or neutral emotions, such as "happy," "sad," and "satisfied." Positional weight refers to the importance of each sentiment word's location within the sentence, typically determined by its grammatical function and semantic influence. Sentence-level sentiment polarity is a numerical value used to measure the overall sentiment tendency of the sentence, comprehensively considering the sentiment tendency and positional weights of all sentiment words, as well as the influence of grammatical structural components on sentiment expression.

[0053] Step S250: Construct an emotional polarity feature sequence based on the changing trend of statement-level emotional polarity values ​​over time. Each element in the emotional polarity feature sequence corresponds to an emotional state description of a dialogue turn.

[0054] The trend over time refers to the change in sentence-level sentiment polarity values ​​as the number of dialogue rounds increases, reflecting the dynamic evolution of user emotions during the dialogue process. The sentiment polarity feature sequence is a sequence arranged according to the dialogue rounds, where each element corresponds to a sentiment state description for a given dialogue round, visually demonstrating the fluctuations and changes in user emotions during the dialogue.

[0055] The process of constructing an emotional polarity feature sequence based on the changing trend of statement-level emotional polarity values ​​over time is as follows: First, the statement-level emotional polarity value is recorded sequentially according to the dialogue rounds. Then, the changing trend of these emotional polarity values ​​is analyzed, such as whether they show an increase, decrease, or fluctuation. Finally, the emotional state description of each dialogue round (e.g., positive, negative, neutral) is used as elements to construct the emotional polarity feature sequence.

[0056] Step S300: Input the user intent feature sequence and the emotion polarity feature sequence into the preset dialogue development trajectory prediction model to generate the expected interaction trajectory set of the user's current dialogue stage. The expected interaction trajectory set contains multiple possible dialogue development paths and their corresponding path credibility indicators.

[0057] The pre-defined dialogue trajectory prediction model is a trained machine learning or deep learning model that can predict the possible dialogue trajectory of a user at the current stage of the conversation based on the user's intent feature sequence and sentiment polarity feature sequence. The expected interaction trajectory set is a collection of multiple possible dialogue paths, each representing a possible direction of the conversation. The path credibility index is a numerical value used to measure the probability of each dialogue path, reflecting the likelihood of that path occurring in the actual conversation.

[0058] In one implementation, step S300 may include the following steps S310 to S350:

[0059] Step S310: Perform feature dimension alignment and standardization on the user intent feature sequence and sentiment polarity feature sequence. Transform the feature vectors of the two sequences to the same high-dimensional feature space through linear mapping to generate an intent feature matrix and sentiment feature matrix with consistent dimensions.

[0060] Feature dimension alignment refers to adjusting the feature vector dimensions of the user intent feature sequence and the sentiment polarity feature sequence to the same size for subsequent processing and analysis. Standardization is a process of normalizing feature vectors to ensure they have the same scale and range, avoiding model training instability caused by differences in feature scale. Linear mapping is a method of transforming feature vectors from one space to another through linear transformation. A shared high-dimensional feature space refers to a vector space with higher dimensions; transforming the feature vectors of the two sequences into this space facilitates unified processing and comparison. The intent feature matrix and sentiment feature matrix are representations of the user intent feature sequence and sentiment polarity feature sequence in matrix form, with consistent dimensions for convenient subsequent fusion and analysis.

[0061] The following method can be used to align and standardize the feature dimensions of user intent feature sequences and sentiment polarity feature sequences to generate intent feature matrices and sentiment feature matrices with consistent dimensions. First, analyze the feature vector dimensions of the two sequences and find the largest dimension. Then, adjust the feature vector dimensions of the two sequences to the largest dimension by padding or truncation, achieving feature dimension alignment. Next, standardize the feature vectors, for example, using Z-score standardization to adjust the mean of each feature vector to 0 and the standard deviation to 1. Finally, transform the standardized feature vectors to the same high-dimensional feature space through linear mapping (such as matrix multiplication) to generate the intent feature matrix and sentiment feature matrix.

[0062] Step S320: Input the intent feature matrix and the emotion feature matrix into the feature fusion layer of the dialogue development trajectory prediction model, calculate the cross-influence weight between intent and emotion through a bidirectional attention mechanism, and generate a coupled feature matrix that integrates intent and emotion information.

[0063] The feature fusion layer is a crucial component of the dialogue trajectory prediction model. Its main function is to fuse the intent feature matrix and the sentiment feature matrix to extract key information. The bidirectional attention mechanism is an attention mechanism that simultaneously considers the context of the input sequence, calculating the cross-influence weights between intent and sentiment—that is, the degree of influence of intent on sentiment and the degree of influence of sentiment on intent. The coupled feature matrix, which integrates intent and sentiment information, is a matrix obtained by fusing these two data points, combining their features to more comprehensively reflect the user's intent and emotional state.

[0064] In one implementation, step S320 may include the following steps S321 to S326:

[0065] Step S321: Construct an intent-emotion interaction matrix in the feature fusion layer. The row dimension of the matrix corresponds to the sequence length of the intent feature matrix, the column dimension corresponds to the sequence length of the emotion feature matrix, and the matrix element is the dot product similarity of two corresponding feature vectors.

[0066] The intent-sentiment interaction matrix is ​​used to represent the interaction relationship between the intent feature matrix and the sentiment feature matrix. The row dimension of the matrix corresponds to the sequence length of the intent feature matrix, i.e., the number of intent feature vectors; the column dimension corresponds to the sequence length of the sentiment feature matrix, i.e., the number of sentiment feature vectors. Dot product similarity is a method to measure the similarity between two vectors. It is obtained by calculating the dot product of the two vectors; the larger the dot product value, the more similar the two vectors are.

[0067] The intent-emotion interaction matrix can be constructed in the feature fusion layer using the following method. First, obtain the sequence lengths of the intent feature matrix and the emotion feature matrix. Then, create a matrix whose size is the product of the sequence lengths of the intent feature matrix and the emotion feature matrix. Next, iterate through each eigenvector of the intent feature matrix and the emotion feature matrix, calculate their dot product similarity, and fill the result into the corresponding position in the intent-emotion interaction matrix.

[0068] Step S322: Perform row normalization on the intention-emotion interaction matrix to obtain the intention-oriented emotion attention weight distribution, which is used to characterize the degree of influence of each intention unit on the emotional state.

[0069] Row normalization is a process of normalizing each row of the matrix so that the sum of the elements in each row is 1. The intent-driven emotion attention weight distribution is a weight distribution representing the degree of influence of each intent unit on the emotion state, obtained by performing row normalization on the intent-emotion interaction matrix. The larger the value of each element, the greater the influence of the corresponding intent unit on the corresponding emotion state.

[0070] Step S323: Perform column normalization on the intention-emotion interaction matrix to obtain the emotion-oriented intention attention weight distribution, which is used to characterize the moderating coefficient of each emotional state on intention expression.

[0071] Column normalization is a process of normalizing each column of the matrix so that the sum of the elements in each column is 1. The emotion-oriented intention attention weight distribution is a weight distribution representing the moderating coefficient of each emotion state on intention expression, obtained by performing column normalization on the intention-emotion interaction matrix. The larger the value of each element, the greater the moderating effect of the corresponding emotion state on the corresponding intention expression.

[0072] The following steps can be used to normalize the columns of the intent-emotion interaction matrix to obtain the emotion-oriented intent attention weight distribution. First, iterate through each column of the intent-emotion interaction matrix and calculate the sum of the elements in that column. Then, divide each element in that column by the sum to obtain the normalized element values. Finally, use the normalized matrix as the emotion-oriented intent attention weight distribution.

[0073] Step S324: Weight the intention-oriented emotional attention weight distribution with the emotional feature matrix to generate an emotionally enhanced intention feature matrix.

[0074] Weighted summation is the operation of multiplying corresponding elements of two matrices and then summing them. The intention feature matrix for emotion enhancement is a matrix obtained by enhancing the intention feature matrix by considering the influence of intention units on emotional states, thus more accurately reflecting the interaction between intention and emotion.

[0075] The following steps can be used to generate an emotion-enhanced intention feature matrix by weighted summation of the intention-oriented emotion attention weight distribution and the emotion feature matrix. First, ensure that the dimensions of the intention-oriented emotion attention weight distribution and the emotion feature matrix are compatible. Then, multiply each row of the intention-oriented emotion attention weight distribution element-wise with the corresponding column of the emotion feature matrix to obtain an intermediate matrix. Finally, sum each row of the intermediate matrix to obtain the emotion-enhanced intention feature matrix.

[0076] Step S325: Weight the emotion-oriented intention attention weight distribution and the intention feature matrix to generate an intention-enhanced emotion feature matrix.

[0077] Similarly, the weighted summation here involves multiplying corresponding elements of the two matrices and then summing them. The intention-enhanced sentiment feature matrix is ​​based on the sentiment feature matrix, but is enhanced by considering the moderating effect of emotional state on intention expression, resulting in a matrix that better reflects the relationship between emotion and intention.

[0078] The following steps can be used to generate an intent-enhanced sentiment feature matrix by weighted summation of the sentiment-oriented intent attention weight distribution and the intent feature matrix. First, check if the dimensions of the sentiment-oriented intent attention weight distribution and the intent feature matrix match. Then, multiply each column of the sentiment-oriented intent attention weight distribution element-wise with the corresponding row of the intent feature matrix to obtain an intermediate matrix. Finally, sum each column of the intermediate matrix to obtain the intent-enhanced sentiment feature matrix.

[0079] Step S326: Concatenate the intent feature matrix and the sentiment feature matrix of the sentiment enhancement according to the channel dimension, and perform feature dimensionality reduction and information fusion through a 1×1 convolutional layer to generate a coupled feature matrix that integrates intent and sentiment information.

[0080] A 1×1 convolutional layer is a special type of convolutional layer with a kernel size of 1×1. It is mainly used for feature dimensionality reduction and information fusion. By performing a convolution operation on the concatenated matrix, the feature dimensionality can be reduced while extracting key information. The coupled feature matrix that integrates intent and sentiment information is the matrix obtained after feature dimensionality reduction and information fusion. It combines enhanced information from intent and sentiment and can be used more effectively for subsequent dialogue trajectory prediction.

[0081] Step S330: Input the coupled feature matrix into the trajectory generation layer of the dialogue development trajectory prediction model, and perform latent space sampling through the variational autoencoder structure to generate latent vector representations of multiple candidate interaction trajectories.

[0082] The trajectory generation layer is responsible for generating possible dialogue development trajectories based on the input coupled feature matrix. The variational autoencoder structure is a generative model consisting of an encoder and a decoder. It maps input data to a latent space and samples new data from the latent space. The latent space is an abstract vector space where each point represents a possible dialogue development trajectory. The latent vector representation of a candidate interaction trajectory is a set of vectors sampled from the latent space, with each vector corresponding to an abstract representation of a possible candidate interaction trajectory.

[0083] In one implementation, step S330 may include the following steps S331 to S335:

[0084] Step S331: Input the coupled feature matrix into the encoder part of the variational autoencoder, extract the temporal dependent features through a bidirectional LSTM network, and output the latent distribution parameters containing the mean vector and variance vector.

[0085] The encoder part of a variational autoencoder encodes the input data, transforming it into a representation in the latent space. A bidirectional LSTM network is a recurrent neural network that can simultaneously consider the contextual information of the input sequence, effectively extracting temporal dependency features from the coupling feature matrix. Temporal dependency features refer to the interrelationships and patterns of change in data over time; for dialogue trajectory prediction, they reflect the contextual information and development trend of the dialogue. The mean vector and variance vector are two important parameters of the latent distribution, describing the distribution of data in the latent space.

[0086] Step S332: Construct a normal distribution model based on the latent distribution parameters, and generate multiple independent latent vector samples from the normal distribution model through reparameterization techniques. Each latent vector sample corresponds to an abstract representation of a candidate interaction trajectory.

[0087] The normal distribution model is a probability distribution model uniquely determined by its mean and variance vectors. The reparameterization technique addresses the problem of gradients failing to backpropagate during sampling. By introducing an auxiliary variable, it transforms the sampling process into a differentiable operation, enabling effective model training. Independent latent vector samples are a set of vectors sampled from the normal distribution model; each vector represents an abstract representation of a candidate interaction trajectory.

[0088] Constructing a normal distribution model based on latent distribution parameters and generating multiple independent latent vector samples from the normal distribution model using reparameterization techniques can be achieved through the following steps: First, construct a normal distribution model based on the output mean and variance vectors. Then, introduce an auxiliary variable (usually a random vector sampled from the standard normal distribution), and combine the auxiliary variable with the mean and variance vectors using a reparameterization formula to obtain the sampled latent vector samples.

[0089] Step S333: Perform diversity measurement calculation on the latent vector samples. Calculate the similarity between any two latent vector samples using cosine distance. When the similarity is higher than a preset threshold, retain the latent vector sample with the larger corresponding path credibility index value.

[0090] Diversity metrics are calculated to ensure sufficient diversity in the generated candidate interaction trajectories, avoiding a large number of similar trajectories. Cosine distance is a metric used to measure the similarity between two vectors; it determines similarity by calculating the cosine of the angle between the two vectors. The closer the cosine value is to 1, the more similar the two vectors are. The preset threshold is a pre-defined value used to determine if the similarity between two latent vector samples is too high. The path confidence metric is a value calculated during the generation of candidate interaction trajectories, representing the probability of each path.

[0091] Step S334: Input the selected latent vector samples into the feature mapping module of the trajectory generation layer. The latent vectors are mapped from the latent space to the interaction trajectory feature space through a fully connected network to generate a candidate interaction trajectory latent vector representation containing path length information and branch point information.

[0092] The feature mapping module is a component in the trajectory generation layer, whose main function is to map latent vectors from the latent space to the interaction trajectory feature space. A fully connected network is a neural network where each neuron is connected to all neurons in the previous layer. It can perform non-linear transformations on the input latent vectors, achieving the mapping from the latent space to the interaction trajectory feature space. The interaction trajectory feature space is a feature space used to represent candidate interaction trajectories, containing key information such as path length and branch point information. Path length information refers to the length of the candidate interaction trajectory, reflecting the development stage and complexity of the dialogue. Branch point information refers to the positions and situations where branches may occur in the candidate interaction trajectory, reflecting the various possibilities of dialogue development.

[0093] Specifically, the filtered latent vector samples can be input into a fully connected network. The fully connected network processes the input, mapping the latent vectors from the latent space to the interaction trajectory feature space through a series of linear transformations and nonlinear activation functions. Then, path length information and branch point information are extracted from the mapped vectors to generate a candidate interaction trajectory latent vector representation containing this information.

[0094] Step S335: Regularize the latent vector representation of the candidate interaction trajectory to ensure that the feature scale of the latent vectors of different trajectories remains consistent.

[0095] Regularization is a process of normalizing or standardizing data to avoid model training instability caused by differences in feature scale. Maintaining consistent feature scale for latent vectors of different trajectories means adjusting the feature values ​​of all candidate interaction trajectory latent vectors to the same range or with the same statistical properties, enabling the model to treat each trajectory latent vector more fairly and improving the model's training performance and prediction accuracy.

[0096] Various methods can be used to regularize the latent vector representations of candidate interaction trajectories, such as L2 regularization and Z-score standardization. Taking Z-score standardization as an example, firstly, the mean and standard deviation of all candidate interaction trajectory latent vectors are calculated. Then, for each feature value in each latent vector, the mean is subtracted and the result is divided by the standard deviation to obtain the standardized feature value. Through this process, the feature scale of latent vectors from different trajectories will be kept consistent.

[0097] Step S340: Decode the latent vector representation of each candidate interaction trajectory to generate a complete interaction path description containing the possible user input intent sequence and the digital human's suggested response sequence.

[0098] Decoding is the process of converting the latent vector representations of candidate interaction trajectories into concrete text descriptions; it is the output stage of the dialogue development trajectory prediction model. The user's possible input intent sequence refers to a series of intentions and needs that the user might raise within the candidate interaction trajectory. The digital human's suggestion response sequence refers to the corresponding suggestions and responses given by the digital human in response to the user's possible input intent sequence. The complete interaction path description combines the user's possible input intent sequence and the digital human's suggestion response sequence to form a complete dialogue path description, clearly demonstrating the development process and possible directions of the dialogue.

[0099] Specifically, the latent vector representations of candidate interaction trajectories can be input into the decoder of a dialogue development trajectory prediction model. The decoder is typically a language generation model based on a recurrent neural network (such as LSTM, GRU, etc.) or a Transformer architecture, which generates corresponding text sequences based on the latent vector representations. Then, the generated text sequences are parsed to separate the user's possible input intent sequence and the digital human's suggested response sequence. Finally, these two sequences are combined to form a complete description of the interaction path.

[0100] Step S350: Calculate the rationality score of each complete interaction path description through the path evaluation module of the dialogue development trajectory prediction model, calculate the path credibility index by combining the probability distribution in the path generation process, and combine the complete interaction path description with the corresponding path credibility index to obtain the expected interaction trajectory set.

[0101] The path evaluation module is responsible for evaluating the generated complete interaction path descriptions and calculating their reasonableness score. The reasonableness score is a numerical value used to measure the semantic, logical, and contextual reasonableness of a complete interaction path description, reflecting the likelihood and reasonableness of the path appearing in actual dialogue. The probability distribution in the path generation process refers to the probability of each trajectory appearing during the generation of candidate interaction trajectories. The path credibility index is a numerical value obtained by comprehensively considering the reasonableness score and the path generation probability distribution, more accurately reflecting the credibility of each complete interaction path description. The expected interaction trajectory set is a set that combines complete interaction path descriptions with their corresponding path credibility indices, containing multiple possible dialogue development paths and their corresponding credibility information.

[0102] Step S400: Based on the expected interaction trajectory set and the digital human response strategy library, perform matching processing, select the target interaction trajectory that matches the current dialogue scenario according to the path credibility index, and extract the optimal response strategy parameter set corresponding to the target interaction trajectory.

[0103] The expected interaction trajectory set is a collection containing multiple possible dialogue development paths and their corresponding path credibility metrics. The digital human response strategy library is a database that pre-stores digital human response strategies for various dialogue scenarios. Each strategy includes a series of parameters and rules to guide the digital human in making appropriate responses. Matching processing compares and correlates each trajectory in the expected interaction trajectory set with the strategies in the digital human response strategy library to find the most suitable response strategy. The path credibility metric is a numerical value that measures the credibility of each interaction trajectory; based on this metric, more reliable interaction trajectories can be selected. The target interaction trajectory is an interaction trajectory selected from the expected interaction trajectory set that meets the requirements of the current dialogue scenario. The optimal response strategy parameter set is a set of parameters corresponding to the target interaction trajectory; these parameters guide the digital human to generate the most appropriate response.

[0104] In one implementation, step S400 may include the following steps S410-S460:

[0105] Step S410: Perform semantic structure parsing on each trajectory in the expected interaction trajectory set, extract the intent chain sequence, emotional state sequence and semantic dependency path contained in the trajectory, and generate a trajectory semantic structure graph.

[0106] Semantic structure parsing utilizes natural language processing techniques to deeply analyze each trajectory in a set of expected interaction trajectories, revealing their inherent semantic structure and logical relationships. An intent chain sequence refers to the sequence of user intents arranged chronologically within an interaction trajectory, reflecting the evolution and development of user intents during the dialogue. An affective state sequence records the changes in a user's affective state over time during the dialogue, reflecting the user's emotional tendencies at different stages. Semantic dependency paths describe the semantic associations and logical connections between words and sentences within the trajectory, demonstrating the transmission and flow of semantic information within the trajectory. A trajectory semantic structure graph is a graphical representation that intuitively displays the intent chain sequence, affective state sequence, and semantic dependency paths, facilitating subsequent analysis and processing.

[0107] Specifically, first, part-of-speech tagging tools can be used to tag each word in the trajectory to determine its grammatical category. Then, a syntactic analyzer is used to parse the trajectory, identifying its grammatical structure and semantic dependencies. Next, based on the parsing results, intent chain sequences and sentiment state sequences are extracted. For example, by identifying keywords representing intent and words expressing emotion in the trajectory, they are arranged in chronological order to form intent chain sequences and sentiment state sequences. Finally, this information is integrated into a graph to generate a trajectory semantic structure graph.

[0108] Step S420: Extract scene context features from the current dialogue scene, including dialogue domain identifiers, user historical interaction style labels, and current dialogue round statistics, and construct a scene dynamic feature vector.

[0109] Scene context features refer to a series of characteristic information related to the current dialogue scene, reflecting the background of the dialogue, the user's historical behavior, and the current dialogue state. Dialogue domain identifiers clearly define the domain involved in the current dialogue, such as tourism, shopping, and healthcare; different dialogue domains have different semantics and rules. User history interaction style labels are feature labels summarized based on the user's past interaction behavior, such as concise, detailed, and friendly, helping the digital human better understand user preferences and habits. Current dialogue round statistics record the number of rounds already conducted in the current dialogue, reflecting the progress of the dialogue. The scene dynamic feature vector is a vector formed by combining these scene context features, comprehensively representing the characteristic information of the current dialogue scene.

[0110] Step S430: Perform structured matching between the trajectory semantic structure graph and the policy templates in the digital human response policy library, identify the intention response nodes that match the intention chain sequence and the emotion adjustment nodes that match the emotion state sequence in the policy template, and construct a trajectory-policy bidirectional association network. The nodes of the association network are trajectory features or policy nodes, and the edges are the semantic mapping relationship between features and nodes.

[0111] Structured matching compares and analyzes the trajectory semantic structure graph and policy templates in the digital human response policy library to find the correspondence between them. Intent response nodes are the parts of the policy template that respond to a given intent, containing corresponding reply content and rules. Emotional adjustment nodes are the parts of the policy template used to adjust the user's emotional state, such as expressions of comfort or encouragement. The trajectory-policy bidirectional association network is a network structure used to represent the relationship between trajectory features and policy nodes. Nodes can be features such as intent chain sequences and emotional state sequences in the trajectory, or intent response nodes and emotional adjustment nodes in the policy template. Edges represent the semantic mapping relationship between these features and nodes, i.e., their degree of semantic association.

[0112] In one implementation, step S430 may specifically include the following steps S431 to S436:

[0113] Step S431: Parse the structured description of the strategy template in the digital human response strategy library, and extract the intention response node sequence and emotion regulation node sequence contained in the strategy template. Each intention response node contains a core intention identifier and response priority, and each emotion regulation node contains an emotion state identifier and regulation intensity description.

[0114] The structured description of the strategy template is a detailed explanation of the strategy template, recording its various components and rules in a structured manner. The intent response node sequence is a series of intent response nodes arranged in a specific order within the strategy template. Each intent response node has a clear core intent identifier and response priority. The core intent identifier clarifies the specific intent targeted by the node, such as "check product price" or "complain about product quality issues." The response priority indicates the node's priority in responding; when multiple intent response nodes are available, the higher-priority node will be used first. The emotion adjustment node sequence is a series of nodes in the strategy template used to adjust the user's emotional state. Each emotion adjustment node includes an emotion state identifier and a description of adjustment intensity. The emotion state identifier clarifies the emotional state targeted by the node, such as "positive," "negative," or "neutral," while the adjustment intensity description indicates the degree to which the node adjusts the emotional state, such as "mild adjustment," "moderate adjustment," or "severe adjustment."

[0115] Step S432: Decompose the intent chain sequence in the trajectory semantic structure graph into intent units to obtain continuous independent intent units. Each intent unit contains an intent type identifier and core semantic keywords.

[0116] Intent unit decomposition involves dividing the intent chain sequence in the trajectory semantic structure graph according to semantic independence and completeness, breaking it down into multiple consecutive independent intent units. An independent intent unit is the basic unit for refining and classifying user intents; each independent intent unit represents a specific intent of the user during the dialogue. Intent type identifiers clarify the specific type of each intent unit, such as "asking for information," "making a request," or "expressing an opinion." Core semantic keywords are a set of representative keywords extracted from each independent intent unit, accurately summarizing the core semantics of that intent unit.

[0117] For example, based on the semantic and logical relationships of statements in the intent chain sequence, determine which statements can form an independent intent unit. For instance, when there are clear semantic shifts or thematic changes between statements, they are divided into different intent units. Then, assign an intent type identifier to each intent unit. This can be done by analyzing keywords and statement structure within the intent unit, combined with predefined intent type classification rules. Finally, extract core semantic keywords from each intent unit. Keyword extraction algorithms, such as TF-IDF and TextRank, can be used to find representative keywords.

[0118] Step S433: Perform dynamic programming matching between the split intent unit sequence and the intent response node sequence of the policy template, calculate the length of the longest common subsequence between the sequences, mark the intent response nodes in the longest common subsequence as matching nodes, and record the position index of the matching nodes in the policy template.

[0119] Dynamic programming matching is an algorithm used to solve sequence matching problems. It constructs a two-dimensional matrix and uses dynamic programming to progressively calculate the optimal matching result between two sequences. The longest common subsequence length refers to the length of the longest common subsequence between the two sequences, reflecting the similarity between them. A matching node is a node in the intent response node sequence of the policy template that matches the split intent unit sequence. The position index records the specific location of the matching node in the policy template, facilitating subsequent processing and association.

[0120] For example, create a two-dimensional matrix where the rows and columns correspond to the split intent unit sequence and the intent response node sequence of the policy template, respectively. Then, initialize the first row and first column of the matrix to 0. Next, iterate through each element of the matrix, calculating the value of each element according to the recursive formula of dynamic programming. The recursive formula is: if the core semantic keyword of the current intent unit matches the core intent identifier of the current intent response node, the value of the matrix element is the value of the top-left element plus 1; otherwise, the value of the matrix element is the maximum value between the elements above and to the left. Finally, starting from the bottom-right corner of the matrix, backtrack to find the longest common subsequence, mark the intent response nodes in the longest common subsequence as matching nodes, and record their position index in the policy template.

[0121] Step S434: Divide the emotional state sequence in the trajectory semantic structure graph into state units. Each state unit corresponds to the emotional state of a dialogue turn, including the emotional type identifier and the duration of the state.

[0122] State unit partitioning involves dividing the sentiment state sequence in the trajectory semantic structure graph according to dialogue turns, breaking it down into multiple independent state units. Each state unit represents the sentiment state of one dialogue turn, including a sentiment type identifier and state duration. The sentiment type identifier clarifies the sentiment type corresponding to the state unit, such as "positive," "negative," or "neutral." The state duration records the number of turns in which the sentiment state persists in the dialogue, reflecting the stability and persistence of the sentiment state.

[0123] Step S435: Perform sliding window matching between the segmented emotional state sequence and the emotional adjustment node sequence of the strategy template. The window size is the preset number of emotional state units. Calculate the number of matching emotional type identifiers within the window. When the number of matching reaches the preset proportion of the window size, mark the emotional adjustment node within the window as a matching node and record the position index of the matching node in the strategy template.

[0124] Sliding window matching is an algorithm used to find matching subsequences in a sequence. It works by sliding a fixed-size window across the sequence, comparing elements within the window with elements in the target sequence. The preset number of sentiment state units refers to the size of the sliding window, which determines the number of sentiment state units compared each time. The number of matches refers to the number of elements within the sliding window that match the sentiment type identifier with the sentiment state identifier in the sentiment adjustment node sequence of the strategy template. The preset ratio is a pre-defined value used to determine whether a match is successful; a match is considered successful when the number of matches reaches the preset ratio of the window size. A matching node is a node in the sentiment adjustment node sequence of the strategy template that matches a sentiment state unit within the sliding window; its position index records the specific location of these matching nodes within the strategy template.

[0125] For example, determine the preset number and proportion of emotional state units. Then, slide the sliding window starting from the beginning of the emotional state sequence, sliding one emotional state unit at a time. For each window position, compare the emotional type identifier within the window with the corresponding emotional state identifier in the emotional adjustment node sequence of the strategy template, and calculate the number of matches. When the number of matches reaches a preset proportion of the window size, mark the corresponding emotional adjustment node within the window as a matched node and record its position index in the strategy template.

[0126] Step S436: Using the intent unit and sentiment state unit of the trajectory semantic structure graph as source nodes and the matching node of the policy template as target nodes, construct directed edges according to the position index of the matching nodes to generate a trajectory-policy bidirectional association network containing source nodes, target nodes and directed edges. Mark the matching strength value on the directed edges. The matching strength value is the semantic similarity between the intent unit or sentiment state unit and the corresponding node.

[0127] The trajectory-policy bidirectional association network is a network structure used to represent the association between the trajectory semantic structure graph and the policy template. It connects source nodes (intent units and sentiment state units in the trajectory semantic structure graph) and target nodes (matching nodes in the policy template) through directed edges. The matching strength value is an indicator that measures the semantic similarity between the source and target nodes, reflecting the degree of association between them.

[0128] Step S440: Calculate the intent matching completeness and sentiment adaptation coherence of each trajectory based on the trajectory-policy bidirectional association network. The intent matching completeness is the proportion of the continuous matching length between the trajectory intent chain sequence and the policy intent response node to the total length of the trajectory. The sentiment adaptation coherence is the proportion of the consistency of the state transition between the trajectory sentiment state sequence and the policy sentiment adjustment node to the total number of transitions in the sequence.

[0129] Intent matching completeness is an indicator that measures the degree of matching between the trajectory intent chain sequence and the policy intent response nodes, reflecting the extent to which the policy template covers the intent in the trajectory. Continuous matching length refers to the length of the portion of the trajectory intent chain sequence that continuously matches the policy intent response nodes; the total trajectory length is the overall length of the trajectory intent chain sequence. Emotional fit coherence is an indicator that measures the degree of fit between the trajectory emotional state sequence and the policy emotional regulation nodes, focusing on whether the transition of emotional states is consistent with the regulation rules in the policy. State transition consistency refers to the degree of conformity between the transition of emotional states in the trajectory emotional state sequence and the state transition rules specified by the policy emotional regulation nodes; the total number of sequence transitions is the total number of times emotional states transition in the trajectory emotional state sequence.

[0130] Step S450: Intent matching completeness, sentiment adaptation coherence and path credibility indicators are integrated. During the integration process, the weight distribution of each indicator is adjusted in combination with the scene dynamic feature vector. For every preset increase in the dialogue turn statistics value in the scene context features, the weight of intent matching completeness is increased by a preset ratio to generate trajectory-strategy comprehensive adaptation indicators.

[0131] Intent matching completeness, sentiment adaptation coherence, and path credibility metrics each reflect the degree of fit between the trajectory and the strategy from different perspectives. Fusing them yields a comprehensive metric that more fully measures the matching between the trajectory and the strategy. The scene dynamic feature vector includes information such as dialogue domain identifiers, user historical interaction style tags, and current dialogue round statistics, reflecting the characteristics and state of the current dialogue scene. During the fusion process, the weight allocation of each metric is adjusted based on the scene dynamic feature vector, allowing the metric weights to dynamically adjust according to changes in the dialogue scene. Preset quantities and preset ratios are pre-defined values ​​used to specify the increase in the intent matching completeness weight as the dialogue round statistics increase. The trajectory-strategy comprehensive adaptation metric is a comprehensive value after fusion, taking into account multiple factors such as intent matching, sentiment adaptation, and path credibility, enabling a more accurate assessment of the trajectory-strategy adaptability.

[0132] In one implementation, step S450 may include the following steps S451 to S455:

[0133] Step S451: Initialize the weight coefficients for intent matching completeness, emotional adaptation coherence, and path credibility index. The initial weight coefficients are determined based on the correlation analysis between each index and interaction satisfaction in historical dialogue data.

[0134] Historical dialogue data refers to a large number of past dialogue records, containing rich information such as user intent, emotional state, digital human response strategies, and interaction satisfaction. Correlation analysis is used to study the degree of association between two or more variables. By performing correlation analysis on historical dialogue data, the relationships between intent matching completeness, emotional fit coherence, and path credibility indicators and interaction satisfaction can be identified, thereby determining the initial weight coefficients for each indicator. The initial weight coefficients are the weight values ​​assigned to each indicator at the beginning of the fusion process, reflecting the importance of each indicator in the comprehensive evaluation.

[0135] For example, a large amount of historical dialogue data is collected and preprocessed, including data cleaning and feature extraction. Then, correlation analysis methods (such as Pearson correlation coefficient, Spearman correlation coefficient, etc.) are used to calculate the correlation coefficient between each indicator and interaction satisfaction. The larger the correlation coefficient, the stronger the association between the indicator and interaction satisfaction, and the higher its weight should be given in the comprehensive evaluation. Finally, initial weight coefficients are assigned to each indicator based on the magnitude of the correlation coefficient.

[0136] Step S452: Extract the dialogue round statistics from the scene dynamic feature vector. When the dialogue round statistics are greater than the preset round threshold, call the weight dynamic adjustment strategy. For each additional preset number of rounds, the weight coefficient of intent matching completeness increases by a preset ratio, while the weight coefficient of emotion adaptation coherence decreases by the corresponding ratio, and the weight coefficient of path credibility index remains unchanged.

[0137] The scene dynamic feature vector contains a series of feature information of the current dialogue scene, among which the dialogue round statistics record the number of rounds that have been performed in the current dialogue. The preset round threshold is a pre-set value used to determine whether the dialogue has entered a stage where weight adjustments are needed. When the dialogue round statistics are greater than the preset round threshold, it indicates that the dialogue has already gone through a certain number of rounds, and at this point, the weight coefficients of each indicator need to be adjusted according to the progress of the dialogue. The dynamic weight adjustment strategy is a rule that dynamically adjusts the weight coefficients of each indicator based on the dialogue round statistics, specifying how the weight coefficients of each indicator change with each increase of a preset number of rounds. The preset percentage is a pre-set percentage used to specify the increase in the intent matching completeness weight coefficient and the decrease in the sentiment adaptation coherence weight coefficient.

[0138] Step S453: Calculate the weighted values ​​of each indicator after adjustment. The weighted value of intent matching completeness is the product of intent matching completeness and the adjusted intent weight coefficient. The weighted value of sentiment adaptation coherence is the product of sentiment adaptation coherence and the adjusted sentiment weight coefficient. The weighted value of path credibility is the product of path credibility indicator and path weight coefficient.

[0139] The adjusted weighted values ​​for each indicator are calculated based on the adjusted weight coefficients. They reflect the actual contribution of each indicator to the overall evaluation. The intention matching completeness weighted value is the product of intention matching completeness and the adjusted intention weight coefficient, reflecting the importance of intention matching in the overall evaluation. The sentiment fit coherence weighted value is the product of sentiment fit coherence and the adjusted sentiment weight coefficient, reflecting the role of sentiment fit in the overall evaluation. The path credibility weighted value is the product of the path credibility indicator and the path weight coefficient, representing the influence of path credibility in the overall evaluation.

[0140] Step S454: Sum the three weighted values ​​to obtain the initial value of the trajectory-strategy integrated adaptation index.

[0141] The initial value of the trajectory-strategy integrated adaptation index is obtained by summing the weighted values ​​of intent matching completeness, sentiment adaptation coherence, and path credibility. It is a preliminary evaluation result that comprehensively considers multiple factors such as intent matching, sentiment adaptation, and path credibility. By summing the weighted values ​​of each index, information from different aspects can be integrated to obtain a more comprehensive evaluation index.

[0142] Step S455: Standardize the initial values. The standardization process is based on the mean and standard deviation of the comprehensive adaptation index of all strategy templates in the digital human response strategy library to generate a trajectory-strategy comprehensive adaptation index that conforms to the preset distribution range.

[0143] Standardization involves normalizing or standardizing data to ensure it has the same scale and range, facilitating comparison and analysis. The mean and standard deviation of the comprehensive adaptation index for all strategy templates in the digital human response strategy library are numerical values ​​obtained through statistical analysis of the comprehensive adaptation index of all strategy templates, reflecting the overall distribution of the comprehensive adaptation index. The preset distribution range is a pre-defined numerical range, such as [0, 1]. The standardized trajectory-strategy comprehensive adaptation index will fall within this range, making the comprehensive adaptation index of different strategy templates comparable.

[0144] Step S460: Sort the expected interaction trajectories in descending order of trajectory-policy comprehensive adaptation index, select the first ranked expected interaction trajectory as the target interaction trajectory, and extract the set of policy parameters of the corresponding policy template in the trajectory-policy bidirectional association network of the target interaction trajectory as the optimal response policy parameter set.

[0145] The trajectory-policy comprehensive fit index is a metric that comprehensively measures the degree of fit between a trajectory and a policy, taking into account multiple factors such as intent matching, sentiment adaptation, and path credibility. Expected interaction trajectories are sorted from highest to lowest according to the trajectory-policy comprehensive fit index, placing trajectories with higher fit at the top. The target interaction trajectory is selected from the sorted expected interaction trajectories, representing the trajectory with the highest fit. The optimal response policy parameter set is a set of parameters corresponding to the policy template of the target interaction trajectory in the trajectory-policy bidirectional association network. These parameters guide the digital human to generate the most appropriate response.

[0146] Step S500: Construct a structured response sequence for the digital human based on the optimal response strategy parameter set, and convert the structured response sequence into digital human output text.

[0147] The optimal response strategy parameter set is a filtered and determined set of parameters that contains various information needed for the digital human to generate appropriate responses in the current dialogue scenario, such as the content of the intended response, the method of emotion modulation, and the structure of the response. The structured response sequence is a response representation with a clear structure and hierarchy constructed based on the optimal response strategy parameter set. It organizes and arranges the response content according to certain rules, facilitating subsequent processing and transformation. The digital human's output text is the final natural language text presented to the user. It is the result of transforming the structured response sequence and has good readability and fluency.

[0148] For example, a detailed analysis of the optimal response strategy parameter set is performed to understand the meaning and function of each parameter. Then, based on the response structure configuration information in the parameters, the main and auxiliary units of the structured response sequence are divided. Next, the content generation information in the parameters is filled into the main and auxiliary units to form a preliminary structured response sequence. Finally, according to preset text conversion rules, the structured response sequence is converted into natural language text to obtain the digital human's output text.

[0149] In one implementation, step S500 may include the following steps S510-S580:

[0150] Step S510: Analyze the response structure configuration parameters in the optimal response strategy parameter set, and extract the intent response level parameter and the emotion regulation intensity parameter. The intent response level parameter represents the level depth of the user intent chain sequence, and the emotion regulation intensity parameter represents the regulation needs of the emotion state sequence.

[0151] The response structure configuration parameters, part of the optimal response strategy parameter set, define the structure and organization of the digital human's response. The intent response hierarchy parameters clarify the depth of the user's intent chain sequence, reflecting the complexity and hierarchical relationship of the user's intent. For example, a user's intent may have multiple levels, such as "understanding overall product information," "asking about specific product functions," or "comparing different product models." The intent response hierarchy parameters can reflect the number and depth of these levels. The emotion modulation intensity parameter indicates the degree of need for emotional state modulation. Based on the user's emotional state during the dialogue, it determines the required level of emotional modulation. For example, when a user is in a negative emotional state, stronger emotional modulation may be needed; when a user is in a positive emotional state, only mild emotional modulation may be required.

[0152] Step S520: Divide the main unit levels of the structured response sequence based on the intent response level parameters. The number of levels is consistent with the value of the intent response level parameters. Each main unit level corresponds to an intent level in the intent chain sequence. The number of main units in a level is the same as the number of intent units contained in that level.

[0153] The master unit hierarchy in a structured response sequence is a way to organize responses hierarchically, dividing the response content according to the hierarchical relationship of the intent chain sequence. The value of the intent response hierarchy parameter determines the number of master unit levels, with each master unit corresponding to an intent level in the intent chain sequence. The number of master units within a level is the same as the number of intent units contained in that level, ensuring that each intent unit has a corresponding master unit to respond to.

[0154] Step S530: Divide the structured response sequence into auxiliary unit groups based on the emotion regulation intensity parameter. The number of auxiliary unit groups is consistent with the value of the emotion regulation intensity parameter. Each auxiliary unit group corresponds to an emotion regulation interval in the emotion state sequence. The number of auxiliary units in the group is the same as the number of emotion state units contained in the interval.

[0155] The emotion regulation intensity parameter determines the strength and scope of emotion regulation required. Dividing the structured response sequence into auxiliary unit groups based on this parameter allows for better regulation of the user's emotional state. The number of auxiliary unit groups is consistent with the value of the emotion regulation intensity parameter; this means that the greater the emotion regulation intensity, the more auxiliary unit groups are needed. Each auxiliary unit group corresponds to an emotion regulation interval in the emotional state sequence. This interval is defined based on changes in emotional state, and the number of auxiliary units within the group is the same as the number of emotional state units contained in that interval. This ensures that each emotional state unit has a corresponding auxiliary unit for emotion regulation.

[0156] In one implementation, step S530 may include the following steps S531 to S535:

[0157] Step S531: Perform first-order difference calculation on the user's emotional polarity feature sequence to obtain the emotional change rate sequence. The difference result is the difference in emotional polarity value between adjacent emotional state units.

[0158] The user's emotional polarity feature sequence records the changes in emotional polarity during a conversation. Each emotional state unit corresponds to the emotional polarity value of a conversation turn. First-order difference calculation involves subtracting adjacent elements in the sequence to obtain the difference between them. The emotional change rate sequence is obtained by performing first-order difference calculation on the user's emotional polarity feature sequence, reflecting the change in emotional polarity between adjacent conversation turns. The difference result is the difference in emotional polarity values ​​between adjacent emotional state units; a positive value indicates increased emotional polarity, and a negative value indicates decreased emotional polarity.

[0159] Step S532: Set a threshold for dividing the emotion regulation interval. When the absolute value of the difference result in the emotion change rate sequence is greater than the threshold, it is marked as the starting point of the emotion regulation interval. Starting from the starting point, emotion state units are continuously extracted until the absolute value of the difference result is less than the threshold, which is then marked as the end point of the interval, thus obtaining an emotion regulation interval.

[0160] The threshold for defining the emotion regulation interval is a pre-set value used to determine whether emotion changes have reached a level requiring regulation. When the absolute value of the difference in the emotion change rate sequence is greater than this threshold, it indicates a significant change in emotion polarity, and this point is marked as the starting point of the emotion regulation interval. Starting from the starting point, emotion state units are continuously extracted until the absolute value of the difference is less than the threshold, indicating that the emotion change has stabilized, and this point is marked as the end point of the interval. This process yields a complete emotion regulation interval.

[0161] For example, based on historical dialogue data and experience, a suitable threshold for dividing the emotion regulation interval is set. Then, the emotion change rate sequence is traversed, and points where the absolute value of the difference result is greater than the threshold are found and marked as the starting point of the emotion regulation interval. Starting from the starting point, the emotion change rate sequence is traversed again until a point where the absolute value of the difference result is less than the threshold is found and marked as the ending point of the interval. The emotion state units between the starting point and the ending point are recorded to form an emotion regulation interval.

[0162] Step S533: Count the total number of emotion regulation intervals and compare the number with the value of the emotion regulation intensity parameter. When the number is greater than the emotion regulation intensity parameter, retain the N intervals with the largest absolute value of the difference result, where N is the value of the emotion regulation intensity parameter. When the number is less than the emotion regulation intensity parameter, split the longest interval into multiple sub-intervals so that the number of sub-intervals is consistent with the value of the emotion regulation intensity parameter.

[0163] The total number of emotion regulation intervals is obtained by statistically analyzing the previously defined intervals. Comparing this number with the emotion regulation intensity parameter allows for adjustment of the intervals based on the needs of emotion regulation. When the total number of emotion regulation intervals exceeds the emotion regulation intensity parameter, it indicates frequent emotion changes, requiring filtering. The top N intervals with the largest absolute difference, where N is the emotion regulation intensity parameter, are retained. This ensures that adjustment is only applied to intervals with significant emotion changes. When the total number of emotion regulation intervals is less than the emotion regulation intensity parameter, it indicates relatively few emotion changes. The longest interval needs to be split, ensuring the number of its sub-intervals matches the emotion regulation intensity parameter to meet the needs of emotion regulation.

[0164] Step S534: Assign an auxiliary unit group to each emotion regulation interval. The number of auxiliary units in the group is equal to the number of emotion state units contained in the interval. Each auxiliary unit corresponds to one emotion state unit.

[0165] An auxiliary unit group is a set of units used to regulate the emotional regulation intervals. Assigning an auxiliary unit group to each emotional regulation interval ensures that each interval has a corresponding regulation mechanism. The number of auxiliary units in the group is equal to the number of emotional state units contained in the interval, and each auxiliary unit corresponds to one emotional state unit, thus enabling precise regulation of each emotional state unit.

[0166] Step S535: Extract the emotional state identifier of each emotional regulation interval, and match the corresponding emotional regulation parameter template from the digital human response strategy library based on the emotional state identifier. Each auxiliary unit group is associated with an emotional regulation parameter template to guide subsequent parameter filling.

[0167] Emotional state identifiers are labels for the emotional state corresponding to each emotional regulation interval, such as "positive," "negative," and "neutral," clearly indicating the emotional tendency within that interval. The digital human response strategy library stores emotional regulation parameter templates for various emotional states. These templates contain specific parameters and rules for regulating different emotional states. Matching the corresponding emotional regulation parameter template from the digital human response strategy library based on the emotional state identifier provides appropriate regulation guidance for each auxiliary unit group. Each auxiliary unit group is associated with an emotional regulation parameter template; in subsequent parameter filling processes, the corresponding parameters for the auxiliary units will be filled according to this template.

[0168] Step S540: Arrange the main units and auxiliary units according to the strategy of prioritizing the main unit hierarchy and embedding the auxiliary unit groups. The main unit hierarchy is arranged in the temporal order of the intention chain sequence, and the auxiliary unit groups are embedded between the main unit hierarchy corresponding to the emotion regulation interval, thus obtaining the initial framework of the structured response sequence.

[0169] The strategy of prioritizing main unit levels and embedding auxiliary unit groups is a rule for organizing and arranging structured response sequences. It ensures that the response content is first arranged according to the hierarchical relationship of the intent chain sequence, while auxiliary unit groups are logically embedded in their appropriate positions to regulate emotions. The main unit levels are arranged chronologically according to the intent chain sequence, ensuring that the response content aligns with the user's intent sequence, facilitating user comprehension. Auxiliary unit groups are embedded between main unit levels corresponding to the corresponding emotion regulation intervals, enabling the adjustment of the user's emotional state at appropriate times.

[0170] For example, the order of the main unit levels is determined by sorting them according to the temporal order of the intent chain sequence. Then, based on the correspondence between the emotion modulation intervals and the main unit levels, auxiliary unit groups are embedded in their corresponding positions. For instance, for a structured response sequence containing three main unit levels and two auxiliary unit groups, the main unit levels are arranged as layer one, layer two, and layer three according to the intent chain sequence. If the first auxiliary unit group corresponds to the emotion modulation interval between the first and second layers, and the second auxiliary unit group corresponds to the emotion modulation interval between the second and third layers, then the first auxiliary unit group is embedded between the first and second main unit levels, and the second auxiliary unit group is embedded between the second and third main unit levels. Finally, the arranged main units and auxiliary units are combined to obtain the initial framework of the structured response sequence.

[0171] Step S550: Fill the main and auxiliary units of the initial framework with the parameters generated from the optimal response strategy parameter set. Fill the main unit with the core intent response parameters and fill the auxiliary unit with the emotion regulation supplementary parameters to generate a preliminary structured response sequence.

[0172] Content generation parameters are part of the optimal response strategy parameter set, containing specific information used to generate response content. Core intent response parameters are key information for responding to user intent, such as product introductions and Q&A answers; filling these into the main unit ensures the main unit accurately responds to the user's intent. Emotional adjustment supplementary parameters are supplementary information used to adjust the user's emotional state, such as comforting words and encouraging statements; filling these into auxiliary units effectively regulates the user's emotions. The preliminary structured response sequence is the response sequence obtained after filling the main and auxiliary units of the initial framework with content generation parameters; it already contains basic response content and structure.

[0173] In one implementation, step S550 may include the following steps S551-S556:

[0174] Step S551: Analyze the content in the optimal response strategy parameter set to generate parameters. According to the parameter type, they are divided into core intent response parameters and sentiment regulation supplementary parameters. Core intent response parameters include intent type identifier, core semantic keywords and response priority. Sentiment regulation supplementary parameters include sentiment state identifier, regulation direction and supplementary information items.

[0175] Content generation parameters are the specific parameters used to generate response content within the optimal response strategy parameter set, encompassing various types of information. Analyzing these parameters and dividing them into core intent response parameters and sentiment modulation supplementary parameters allows for more targeted parameter filling of the main and auxiliary units. Core intent response parameters are key parameters for responding to user intents, including intent type identifiers, core semantic keywords, and response priorities. The intent type identifier clarifies the type of intent the parameter targets, such as "querying information" or "making a request"; core semantic keywords are key descriptions of the intent, such as "product price" or "after-sales service"; and response priority indicates the parameter's priority in the response. Sentiment modulation supplementary parameters are used to regulate the user's emotional state, including sentiment state identifiers, regulation direction, and supplementary information items. The sentiment state identifier clarifies the sentiment state the parameter targets, such as "positive," "negative," or "neutral"; the regulation direction describes the method of regulating the sentiment state, such as enhancing positive emotions or alleviating negative emotions; and supplementary information items are the specific regulation content, such as comforting words or encouraging statements.

[0176] Step S552: Assign core intent response parameters to each master unit in the master unit hierarchy. The assignment rule is as follows: the hierarchy number corresponds to the intent hierarchy number in the intent chain sequence. Within the same hierarchy, master units are assigned parameters according to the temporal order of intent units. Each master unit is assigned one core intent response parameter. The intent type identifier of the parameter is consistent with the intent unit type corresponding to the master unit.

[0177] The master unit hierarchy is the part of the structured response sequence divided according to the hierarchical relationship of the intent chain sequence. Each master unit corresponds to an intent unit in the intent chain sequence. Assigning core intent response parameters to each master unit ensures that the master unit can accurately respond to the user's intent. The allocation rules specify how the core intent response parameters are allocated. The hierarchy number corresponds to the intent hierarchy number in the intent chain sequence, ensuring consistency between the master unit hierarchy and the hierarchical relationship of the intent chain sequence. Within the same hierarchy, master units are assigned parameters according to the temporal order of intent units, so that the order of response content matches the order of user intent expression. Each master unit is assigned one core intent response parameter, and the intent type identifier of the parameter is consistent with the intent unit type corresponding to the master unit, ensuring the matching between the parameter and the intent.

[0178] For example, the main unit hierarchy is traversed to determine the hierarchy number and order of each main unit within that hierarchy. Then, based on the hierarchy number and order, the corresponding parameter is selected from the core intent response parameters. Specifically, the corresponding intent hierarchy in the intent chain sequence is found based on the hierarchy number, and then the corresponding intent unit is found based on the order of the main units within that hierarchy. Parameters whose intent type identifier matches the type of the intent unit are selected from the core intent response parameters and assigned to that main unit.

[0179] Step S553: ​​Assign supplementary emotion regulation parameters to each auxiliary unit in the auxiliary unit group. The assignment rule is as follows: the group number corresponds to the number of the emotion regulation interval. Within the same group, the auxiliary units are assigned parameters according to the temporal order of the emotion state units. Each auxiliary unit is assigned one supplementary emotion regulation parameter. The emotion state identifier of the parameter is consistent with the emotion state unit type corresponding to the auxiliary unit.

[0180] The auxiliary unit group is the part of the structured response sequence used to regulate the user's emotional state. Each auxiliary unit corresponds to an emotional state unit in the emotional state sequence. Assigning supplementary emotional regulation parameters to each auxiliary unit enables precise regulation of the user's emotional state. The allocation rules specify how these supplementary emotional regulation parameters are assigned; the group number corresponds to the emotional regulation interval number, ensuring a clear correspondence between the auxiliary unit group and the emotional regulation interval. Within the same group, auxiliary units are assigned parameters according to the temporal order of the emotional state units, ensuring that the order of emotional regulation matches the order of emotional state changes. Each auxiliary unit is assigned one supplementary emotional regulation parameter, and the emotional state identifier of the parameter matches the emotional state unit type corresponding to the auxiliary unit, guaranteeing the matching between the parameter and the emotional state.

[0181] Step S554: Calculate the ratio of the number of core semantic keywords in each main unit to the length of the main unit. When the ratio exceeds the range of main unit ratios determined based on historical dialogue data statistics, adjust the number of core semantic keywords by splitting or merging keywords to bring the ratio within the range.

[0182] The ratio of the number of core semantic keywords to the length of the main unit reflects the density of core semantic information within the main unit. The range of this ratio, determined based on historical dialogue data, is a reasonable range, obtained through statistical analysis of the number and length of keywords in main units from a large number of historical dialogues. When the ratio exceeds this range, it indicates that the density of core semantic information in the main unit is too high or too low, which may affect the quality of the response and user comprehension. Adjusting the number of core semantic keywords by splitting or merging keywords can bring the ratio within a reasonable range, improving the rationality and readability of the response.

[0183] Step S555: Calculate the ratio of the number of supplementary information entries in each auxiliary unit to the length of the auxiliary unit. When the ratio exceeds the range of auxiliary unit ratios determined based on historical dialogue data statistics, adjust the number of supplementary information entries by adding or deleting entries to bring the ratio within the range.

[0184] The ratio of the number of supplementary information items to the length of the auxiliary unit reflects the richness of emotion regulation information within the auxiliary unit. The range of this ratio, determined based on historical dialogue data, serves as a reasonable reference standard. It is derived through statistical analysis of the number and length of supplementary information items in a large number of historical dialogues. When the ratio exceeds this range, it means that the emotion regulation information in the auxiliary unit is either too excessive, potentially leading to information redundancy, or too insufficient, failing to achieve effective emotion regulation. Adjusting the ratio by adding or removing supplementary information items can maintain the emotion regulation information in the auxiliary unit at an appropriate level, thereby improving the quality of emotion regulation.

[0185] The calculation of the ratio and adjustment of the number of supplementary information items can be performed as follows: First, count the number of supplementary information items in each auxiliary unit and the length of the auxiliary unit (which can also be measured by the number of characters or words). Then, calculate the ratio between the two. Compare the calculated ratio with the range of auxiliary unit ratios determined based on historical dialogue data. If the ratio is higher than the upper limit of the range, it means there are too many supplementary information items, and some unnecessary items can be removed; if the ratio is lower than the lower limit of the range, it means there are too few supplementary information items, and some relevant adjustment content can be added appropriately.

[0186] Step S556: Fill the adjusted core intent response parameters and emotion modulation supplementary parameters into the corresponding main unit and auxiliary unit to obtain a preliminary structured response sequence containing unit type, parameter content and time sequence position.

[0187] After the adjustments made in the preceding steps, the core intent response parameters and sentiment modulation supplementary parameters are now more reasonable and optimized. Filling these adjusted parameters into the corresponding main and auxiliary units ensures that each part of the structured response sequence has accurate and appropriate content. The preliminary structured response sequence, including unit type, parameter content, and temporal position, is a response representation with a clear structure and content. It clearly shows the type of each unit (main or auxiliary unit), the parameter content it contains, and its temporal position within the entire response sequence, providing a solid foundation for subsequent semantic coherence verification and text generation.

[0188] For example, based on the correspondence between the main unit, auxiliary units, and parameters, the adjusted core intent response parameters are accurately filled into the main unit, and the adjusted emotion modulation supplementary parameters are filled into the auxiliary units. During the filling process, it is essential to ensure that the parameter content matches the intent and emotion modulation needs of the unit. Simultaneously, the type of each unit, its parameter content, and its temporal position in the structured response sequence are recorded.

[0189] Step S560: Perform semantic coherence verification on the preliminary structured response sequence. The verification process includes calculating the semantic similarity and logical correlation between adjacent units. The semantic similarity is calculated by the overlap rate of the core semantic keywords of the units, and the logical correlation is calculated by the length of the semantic dependency path between units.

[0190] Semantic coherence verification is a crucial step in ensuring the semantic and logical consistency of the initial structured response sequence. Semantic similarity between adjacent units reflects the degree of semantic similarity between two units, which can be intuitively measured by calculating the overlap rate of the core semantic keywords of the units. A higher overlap rate indicates that the two units are semantically closer. Logical relevance reflects the logical connection between adjacent units, and is evaluated by calculating the length of the semantic dependency path between units. A shorter semantic dependency path length indicates a stronger logical relationship between the two units. Through the calculation and analysis of semantic similarity and logical relevance, potential semantic incoherence or logical inconsistencies in the initial structured response sequence can be identified.

[0191] Step S570: When the semantic similarity is lower than the preset similarity threshold or the logical correlation is lower than the preset correlation threshold, a preset transition semantic unit is inserted between adjacent units. The transition semantic unit contains logical connectors and semantic bridging phrases that connect the main unit and the auxiliary unit, generating the final structured response sequence.

[0192] The preset similarity threshold and preset relevance threshold are pre-defined standards used to determine whether the semantic similarity and logical relevance between adjacent units are high enough to ensure the coherence of the response. When the semantic similarity is lower than the preset similarity threshold or the logical relevance is lower than the preset relevance threshold, it indicates that the semantic and logical connections between adjacent units are not strong enough, which may cause difficulties for users in understanding the response. The preset transition semantic units are designed to solve this problem. They contain logical connectors (such as "and," "but," "so," etc.) and semantic bridging phrases (such as "about this issue," "on the other hand," etc.) that connect the main unit and the auxiliary unit, enabling the establishment of reasonable semantic and logical connections between adjacent units. By inserting transition semantic units, the response can be made smoother and more coherent, generating the final structured response sequence.

[0193] For example, during semantic coherence verification, adjacent unit pairs with semantic similarity or logical relevance below a preset threshold are recorded. Then, a suitable preset transition semantic unit is selected for each pair of adjacent units. Based on the semantic and logical relationships between adjacent units, appropriate logical connectors and semantic bridging phrases are selected. Finally, the transition semantic unit is inserted between adjacent units to update the initial structured response sequence, resulting in the final structured response sequence.

[0194] Step S580: Convert the content of each unit in the final structured response sequence into natural language text fragments according to preset text conversion rules. The text conversion rules are determined based on the linguistic feature library of the current dialogue domain, including sentence selection rules and vocabulary matching rules, to obtain the digital human output text.

[0195] The pre-defined text conversion rules are a series of rules formulated based on the linguistic feature database of the current dialogue domain. These rules are used to convert the content of each unit in the final structured response sequence from a structured representation into natural and fluent natural language text fragments. The linguistic feature database is a knowledge base obtained by collecting and organizing the linguistic characteristics, sentence structures, and commonly used vocabulary of the current dialogue domain. Sentence selection rules specify how to choose appropriate sentence structures to express the unit content during the conversion process, making the text more consistent with the expression habits of natural language. Lexical adaptation rules ensure that the vocabulary used during conversion is appropriate to the current dialogue domain and context, avoiding inappropriate vocabulary usage. By following these rules, the final structured response sequence is converted into digital human output text, resulting in output text with good readability and professionalism, better meeting the user's communication needs.

[0196] It is understood that the various algorithms involved in the above descriptions of the embodiments of the present invention, such as Euclidean distance algorithm, cosine distance algorithm, conflict resolution algorithm, etc., can all be obtained from relevant content in the prior art. To save space, they will not be elaborated on in the embodiments of the present invention. In addition, those skilled in the art can supplement the details based on common knowledge in the art when implementing the solution of the present invention. For example, they can use normalization to eliminate dimensional conflicts before feature fusion, use interpolation to eliminate dimensional differences, reasonably set the threshold based on historical data, experience or business scenario requirements, train the model based on a general model training method, set the number of layers in the model structure based on actual needs, select the activation function, etc. The present invention will not provide redundant descriptions of overly detailed implementation processes here.

[0197] In one embodiment, a computer system is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the computer system includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores interactive data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a communication session interaction method based on an AI digital human.

[0198] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer system to which the present invention is applied. A specific computer system may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0199] In one embodiment, a computer system is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

Claims

1. A communication session interaction method based on AI digital humans, characterized in that, The method includes: Extract multi-turn dialogue interaction units from the current communication session interaction stream. Each multi-turn dialogue interaction unit includes the user's input statement sequence and the response statement sequence generated by the AI ​​digital human during the continuous dialogue process, as well as the temporal interaction relationship information between the two sequences. Intent-emotion coupling feature extraction is performed on the multi-turn dialogue interaction unit to obtain the user intent feature sequence and emotion polarity feature sequence corresponding to each dialogue interaction unit. The user intent feature sequence is used to characterize the evolution process of the user's intent in the input statement, and the emotion polarity feature sequence is used to characterize the trajectory of the user's emotion polarity change during the dialogue process. The user intent feature sequence and the emotional polarity feature sequence are input into a preset dialogue development trajectory prediction model to generate a set of expected interaction trajectories for the current dialogue stage of the user. The set of expected interaction trajectories includes multiple possible dialogue development paths and their corresponding path credibility indicators. Based on the expected interaction trajectory set and the digital human response strategy library, the target interaction trajectory that matches the current dialogue scenario is selected according to the path credibility index, and the optimal response strategy parameter set corresponding to the target interaction trajectory is extracted. A structured response sequence for the digital human is constructed based on the optimal response strategy parameter set, and the structured response sequence is converted into output text for the digital human.

2. The method according to claim 1, characterized in that, The step of extracting intent-emotion coupling features from the multi-turn dialogue interaction units to obtain the user intent feature sequence and emotion polarity feature sequence corresponding to each dialogue interaction unit includes: The linguistic features of the user input sentence sequence in the multi-turn dialogue interaction unit are analyzed to identify the grammatical structural components and semantic dependencies of each input sentence and generate a linguistic feature map at the sentence level. Based on the linguistic feature map, the boundaries of intent units are divided, and continuous sentence fragments with complete semantic expression are marked as independent intent units. The core semantic keyword set of each intent unit is extracted. A temporal correlation analysis is performed on the core semantic keyword set. The intent evolution coefficient is calculated based on the semantic similarity and time interval between adjacent intent units. A user intent feature sequence is constructed based on the intent evolution coefficient. Each feature vector in the user intent feature sequence corresponds to the semantic representation of an intent unit. The user input sentence sequence is subjected to sentiment lexical identification, and lexical units containing sentiment tendencies and their position weights in the sentence are extracted. The sentence-level sentiment polarity value is calculated by combining the grammatical structure components in the linguistic feature map. An emotional polarity feature sequence is constructed based on the changing trend of the statement-level emotional polarity value over time. Each element in the emotional polarity feature sequence corresponds to the emotional state description of a dialogue turn.

3. The method according to claim 2, characterized in that, The intention unit boundary segmentation based on the linguistic feature map marks continuous sentence fragments with complete semantic expression as independent intention units, and extracts the core semantic keyword set for each intention unit, including: The sentence nodes in the linguistic feature map are traversed, and sentence fragments containing explicit intent markers are identified as the starting boundaries of candidate intent units. The explicit intent markers include interrogative words, imperative verbs, and modal particles. Calculate the semantic correlation between the starting boundary of the candidate intent unit and the subsequent statement node. When the semantic correlation is lower than a preset threshold, it is determined as the ending boundary of the intent unit, and the boundary range label of the complete intent unit is obtained. Dependency parsing is performed on the sentence fragments within the boundary range marking to identify the subject-verb-object core structure and extract the corresponding content words as a preliminary keyword set; The initial keyword set is input into a pre-trained domain dictionary matching model to filter out professional terms related to the current dialogue domain. These terms are then merged with the initial keyword set and deduplicated to generate a core semantic keyword set. Each keyword in the core semantic keyword set is assigned a weight value calculated based on the TF-IDF algorithm, and the core semantic keyword sequence is obtained by arranging them in descending order of weight values.

4. The method according to claim 2, characterized in that, The step of performing temporal correlation analysis on the core semantic keyword set, calculating the intent evolution coefficient based on the semantic similarity and time interval between adjacent intent units, and constructing a user intent feature sequence based on the intent evolution coefficient includes: The core semantic keyword sets of two adjacent intent units are converted into word vector space representations, and the semantic similarity value between the two sets is calculated. Extract the temporal interaction relationship information in the multi-turn dialogue interaction unit, obtain the number of dialogue turn intervals between adjacent intent units, and convert the number of dialogue turn intervals into a standardized time interval parameter; Based on the semantic similarity value and the standardized time interval parameter, the continuity of intent is determined, and an intent evolution identifier is generated to identify the relationship state between intent units. When the semantic similarity value is higher than the first threshold and the standardized time interval parameter is lower than the second threshold, it is identified as an intent continuity state; otherwise, it is identified as an intent transition state. The core semantic keyword set of each intent unit is converted into a feature vector of fixed dimensions, and each dimension of the feature vector corresponds to the weighted sum of the word vector components of a keyword in the keyword set; The feature vectors are arranged in the chronological order of the dialogue interaction, and logical processing is performed between adjacent feature vectors according to the intent evolution identifier: if the identifier is an intent continuity state, the natural transition between adjacent feature vectors is retained in the feature sequence; if the identifier is an intent transition state, a specific transition feature vector representing the intent transition is inserted into the feature sequence to obtain a user intent feature sequence containing intent evolution information.

5. The method according to claim 1, characterized in that, The step of inputting the user intent feature sequence and the emotional polarity feature sequence into a preset dialogue development trajectory prediction model to generate a set of expected interaction trajectories for the user's current dialogue stage includes: The user intent feature sequence and the sentiment polarity feature sequence are aligned and standardized in terms of feature dimensions. The feature vectors of the two sequences are transformed to the same high-dimensional feature space through linear mapping, generating an intent feature matrix and a sentiment feature matrix with consistent dimensions. The intent feature matrix and the emotion feature matrix are input into the feature fusion layer of the dialogue development trajectory prediction model. The cross-influence weight between intent and emotion is calculated through a bidirectional attention mechanism to generate a coupled feature matrix that integrates intent and emotion information. The coupled feature matrix is ​​input into the trajectory generation layer of the dialogue development trajectory prediction model, and latent vector representations of multiple candidate interaction trajectories are generated by latent space sampling through a variational autoencoder structure. The latent vector representation of each candidate interaction trajectory is decoded to generate a complete interaction path description containing the possible user input intent sequence and the digital human's suggested response sequence; The path evaluation module of the dialogue development trajectory prediction model calculates the rationality score of each complete interaction path description, and calculates the path credibility index by combining the probability distribution in the path generation process. The complete interaction path description and the corresponding path credibility index are combined to obtain the expected interaction trajectory set.

6. The method according to claim 5, characterized in that, The step of inputting the intent feature matrix and the emotion feature matrix into the feature fusion layer of the dialogue development trajectory prediction model, calculating the cross-influence weights between intent and emotion through a bidirectional attention mechanism, and generating a coupled feature matrix that fuses intent and emotion information includes: In the feature fusion layer, an intent-emotion interaction matrix is ​​constructed. The row dimension of the matrix corresponds to the sequence length of the intent feature matrix, the column dimension corresponds to the sequence length of the emotion feature matrix, and the matrix element is the dot product similarity of two corresponding feature vectors. The intention-emotion interaction matrix is ​​row normalized to obtain the intention-oriented emotion attention weight distribution, which is used to characterize the degree of influence of each intention unit on the emotional state. The intention-emotion interaction matrix is ​​normalized to obtain the emotion-oriented intention attention weight distribution, which is used to characterize the moderating coefficient of each emotion state on the intention expression. The intention-oriented emotional attention weight distribution is weighted and summed with the emotional feature matrix to generate an emotionally enhanced intention feature matrix. The emotion-oriented intention attention weight distribution is weighted and summed with the intention feature matrix to generate an intention-enhanced emotion feature matrix. The intent feature matrix and the sentiment feature matrix enhanced by the emotion are concatenated along the channel dimension, and feature dimensionality reduction and information fusion are performed through a 1×1 convolutional layer to generate a coupled feature matrix that integrates intent and sentiment information.

7. The method according to claim 6, characterized in that, The step of inputting the coupled feature matrix into the trajectory generation layer of the dialogue development trajectory prediction model, and generating latent vector representations of multiple candidate interaction trajectories through latent space sampling via a variational autoencoder structure includes: The coupled feature matrix is ​​input into the encoder part of the variational autoencoder, and the temporal dependent features are extracted through a bidirectional LSTM network, outputting latent distribution parameters containing mean vector and variance vector; A normal distribution model is constructed based on the latent distribution parameters. Multiple independent latent vector samples are generated from the normal distribution model through reparameterization techniques. Each latent vector sample corresponds to an abstract representation of a candidate interaction trajectory. The diversity measurement of the latent vector samples is calculated, and the similarity between any two latent vector samples is calculated by cosine distance. When the similarity is higher than a preset threshold, the latent vector sample with the larger corresponding path credibility index value is retained. The selected latent vector samples are input into the feature mapping module of the trajectory generation layer. The latent vectors are mapped from the latent space to the interaction trajectory feature space through a fully connected network to generate candidate interaction trajectory latent vector representations containing path length information and branch point information. The latent vector representations of the candidate interactive trajectories are regularized to ensure that the feature scales of the latent vectors of different trajectories remain consistent.

8. The method according to claim 1, characterized in that, The process of matching the expected interaction trajectory set with the digital human response strategy library, filtering out target interaction trajectories that match the current dialogue scenario based on path credibility index, and extracting the optimal response strategy parameter set corresponding to the target interaction trajectory includes: Semantic structure parsing is performed on each trajectory in the expected interaction trajectory set to extract the intent chain sequence, emotional state sequence and semantic dependency path contained in the trajectory, and generate a trajectory semantic structure graph. Extract scene context features from the current dialogue scene, including dialogue domain identifiers, user history interaction style tags, and current dialogue round statistics, and construct a scene dynamic feature vector; The trajectory semantic structure graph is structurally matched with the strategy templates in the digital human response strategy library. The intention response nodes that match the intention chain sequence and the emotion adjustment nodes that match the emotion state sequence in the strategy template are identified. A trajectory-strategy bidirectional association network is constructed. The nodes of the association network are trajectory features or strategy nodes, and the edges are the semantic mapping relationship between features and nodes. Based on the trajectory-policy bidirectional association network, the intent matching completeness and sentiment adaptation coherence of each trajectory are calculated. The intent matching completeness is the proportion of the continuous matching length between the trajectory intent chain sequence and the policy intent response node to the total length of the trajectory. The sentiment adaptation coherence is the proportion of the consistency of the state transition between the trajectory sentiment state sequence and the policy sentiment adjustment node to the total number of sequence transitions. The intent matching completeness, emotion adaptation coherence and path credibility indicators are fused together. During the fusion process, the weight distribution of each indicator is adjusted in combination with the scene dynamic feature vector. For every preset increase in the dialogue round statistics value in the scene context features, the weight of intent matching completeness is increased by a preset ratio to generate a trajectory-strategy comprehensive adaptation indicator. The expected interaction trajectories are sorted in descending order according to the trajectory-policy comprehensive adaptation index. The expected interaction trajectory with the highest ranking is selected as the target interaction trajectory. The set of policy parameters corresponding to the policy template of the target interaction trajectory in the trajectory-policy bidirectional association network is extracted as the optimal response policy parameter set.

9. The method according to claim 8, characterized in that, The step of performing structured matching between the trajectory semantic structure graph and policy templates in the digital human response policy library, identifying intention response nodes matching the intention chain sequence and emotion modulation nodes matching the emotion state sequence in the policy templates, and constructing a trajectory-policy bidirectional association network includes: The structured description of the strategy templates in the digital human response strategy library is analyzed, and the sequence of intent response nodes and emotion regulation nodes contained in the strategy templates are extracted. Each intent response node contains a core intent identifier and response priority, and each emotion regulation node contains an emotion state identifier and regulation intensity description. The intent chain sequence in the trajectory semantic structure graph is split into intent units to obtain continuous independent intent units. Each intent unit contains an intent type identifier and core semantic keywords. Dynamically programmatically match the split intent unit sequence with the intent response node sequence of the policy template, calculate the length of the longest common subsequence between the sequences, mark the intent response nodes in the longest common subsequence as matching nodes, and record the position index of the matching nodes in the policy template. The emotional state sequence in the trajectory semantic structure graph is divided into state units. Each state unit corresponds to the emotional state of a dialogue turn, including the emotional type identifier and the duration of the state. The segmented emotional state sequence is matched with the emotional regulation node sequence of the strategy template using a sliding window. The window size is the preset number of emotional state units. The number of matching emotional type identifiers within the window is calculated. When the number of matching reaches the preset proportion of the window size, the emotional regulation node within the window is marked as a matching node, and the position index of the matching node in the strategy template is recorded. Using the intent units and sentiment state units of the trajectory semantic structure graph as source nodes and the matching nodes of the policy template as target nodes, directed edges are constructed based on the position index of the matching nodes to generate a trajectory-policy bidirectional association network containing source nodes, target nodes, and directed edges. The matching strength value is marked on the directed edges, and the matching strength value is the semantic similarity between the intent unit or sentiment state unit and the corresponding node.

10. A computer system, characterized in that, include: processor; And a memory, wherein the memory stores a computer program that, when run by the processor, causes the processor to perform the method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • AI intelligent customer service response method and system based on remote digital service

    CN119719319A

  • Interactive multi-round dialogue digital human modeling system and method

    CN120279176A