Intelligent dialogue method and device, electronic equipment and storage medium

Through the emotional recognition model and the personalized portrait optimization interaction strategy, the problem of low accuracy of dialogue reply content in the existing technology is solved, and personalized and emotionally adaptable intelligent dialogue reply is achieved.

CN120407748AActive Publication Date: 2025-08-01FIBOCOM WIRELESS

Patent Information

Application Number
CN202510890236.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-01
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing technology cannot better adapt to the dialogue characteristics of young children, and cannot provide personalized reply content for young children with different personalities, resulting in low accuracy of the generated reply content.

Method used

By obtaining the current dialogue content and historical dialogue data of the target object, the pre-trained emotion recognition model is used to determine the emotion label and personality portrait, and optimize the interaction strategy with the preset mapping relationship and reinforcement learning algorithm to generate personalized reply content.

Benefits of technology

It improves the accuracy of the content of intelligent dialogue reply, can better adapt to the personality characteristics and emotional changes of different children, and provides interactive strategies that are more in line with individual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407748A_ABST
    Figure CN120407748A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent dialogue method and device, electronic equipment and a storage medium. The method comprises the steps that dialogue content and historical dialogue data currently input by a target object are acquired; the dialogue content is input into a pre-trained emotion recognition model, an emotion label of the target object is obtained, a personalized portrait of the target object is determined based on historical dialogue data, and the emotion recognition model is obtained through pre-training on the basis of massive infant dialogue samples; the model parameters of the emotion recognition model can be adjusted based on the personalized portrait of the target object; determining a candidate interaction strategy based on the emotion label of the target object and a preset mapping relationship, and performing optimization learning on the candidate interaction strategy based on the emotion label of the target object and the personalized portrait of the target object to determine a target interaction strategy; and generating the target reply content based on the emotion label of the target object, the personalized portrait of the target object and the target interaction strategy, thereby improving the accuracy of the intelligent dialogue reply content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to an intelligent dialogue method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, intelligent dialogue systems for young children have emerged as the times require. They mainly provide dialogue interactions for young children based on auxiliary teaching and emotion management. For example, they are widely used in devices such as educational intelligent toys and intelligent speakers.

[0003] In the related art, usually, the speech content or text content input by a young child is compared with preset emotion keywords, and then a corresponding dialogue reply is selected according to built-in rules. For example, when a young child says "I don't like this", the system may match the "negative" emotion of the young child and output a dialogue reply such as "It doesn't matter. Let's try something else". However, using this method, it is impossible to better adapt to the dialogue characteristics of young children (such as unclear pronunciation or incomplete sentence patterns), and it is impossible to provide personalized reply content for young children with very different personalities, resulting in a relatively low accuracy of the generated reply content. Therefore, how to improve the intelligent dialogue reply content for young children has become a technical problem to be solved urgently. Summary of the Invention

[0004] This application provides an intelligent dialogue method, device, electronic device and storage medium to solve the technical problem that the prior art cannot better adapt to the dialogue characteristics of young children and cannot provide personalized reply content for young children with very different personalities, resulting in a relatively low accuracy of the generated reply content.

[0005] In a first aspect, an embodiment of this application provides an intelligent dialogue method, and the method includes: Obtain the current dialogue content input by a target object and historical dialogue data, where the target object is any young child; Input the dialogue content into a pre-trained emotion recognition model to obtain an emotion label of the target object, and determine a personality portrait of the target object based on the historical dialogue data, where the emotion recognition model is pre-trained based on a large number of young children's dialogue samples, and the model parameters of the emotion recognition model can be adjusted based on the personality portrait of the target object; Determine a candidate interaction strategy based on the emotion label of the target object and a preset mapping relationship, and optimize and learn the candidate interaction strategy based on the emotion label of the target object and the personality portrait of the target object to determine a target interaction strategy, where the preset mapping relationship is used to represent the mapping relationship between the emotion label and the interaction strategy; Generate a target response content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy.

[0006] Optionally, the emotion recognition model includes a speech feature extraction sub-model, a text feature extraction sub-model, and an emotion analysis sub-model; The step of inputting the conversation content into a pre-trained emotion recognition model to obtain the emotion label of the target object includes: Input the speech information in the conversation content into the speech feature extraction sub-model to obtain the spectral features, pitch changes, and volume intensities corresponding to the speech information, and determine the speech feature vector corresponding to the target object based on the spectral features, pitch changes, and volume intensities corresponding to the speech information; Input the text information in the conversation content into the text feature extraction sub-model to obtain the word embeddings, sentence lengths, and context scores corresponding to the text information, and determine the text feature vector corresponding to the target object based on the word embeddings, sentence lengths, and context scores corresponding to the text information; Input the speech feature vector and the text feature vector into the emotion analysis sub-model to obtain the emotion label of the target object.

[0007] Optionally, the step of inputting the speech feature vector and the text feature vector into the emotion analysis sub-model to obtain the emotion label of the target object includes: Input the speech feature vector and the text feature vector into the emotion analysis sub-model, and use the emotion analysis sub-model to perform feature fusion according to the formula to obtain the emotion label of the target object; where represents the emotion label of the target object, represents the dynamic attention weight, represents the normalization function, represents the weight matrix of the speech feature, represents the speech feature vector, represents the weight matrix of the text feature, represents the text feature vector.

[0008] Optionally, the step of determining the personality portrait of the target object based on the historical conversation data includes: Based on the historical conversation data, calculate the activity level, expression tendency, and emotional stability corresponding to the target object respectively, where the activity level is used to characterize the degree of activity of the target object in participating in the conversation, the expression tendency is used to characterize the relative dependence degree of the target object on tone and vocabulary in language expression, and the emotional stability is used to characterize the degree of emotional fluctuation of the target object; Based on the activity level, expression tendency, and emotional stability corresponding to the target object, determine the personality portrait of the target object.

[0009] Optionally, after determining the personality portrait of the target object based on the historical conversation data, the method further includes: Based on the personality portrait of the target object, adjust the first parameter or the second parameter in the emotion recognition model, so that the emotion label output by the adjusted emotion recognition model is closer to the true emotion label of the target object and matches the personality portrait of the target object, where the first parameter is the relevant parameter corresponding to the weight matrix of voice features, and the second parameter is the relevant parameter corresponding to the weight matrix of text features.

[0010] Optionally, the optimizing and learning the candidate interaction strategies based on the emotion label of the target object and the personality portrait of the target object to determine the target interaction strategy includes: Using a preset reinforcement learning algorithm, optimize and learn the candidate interaction strategies, and determine the target interaction strategy from the candidate interaction strategies, where the reinforcement learning algorithm uses the emotion label of the target object and the personality portrait of the target object as states, the interaction strategies in the candidate interaction strategies as actions, and the degree of emotion relief as rewards, and selects the target interaction strategy by continuously iteratively calculating the values of different state-action pairs, and the target interaction strategy is the interaction strategy corresponding to the optimal value.

[0011] Optionally, after generating the target reply content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy, the method further includes: Obtain the emotion trend factor and the emotional fluctuation range of the target object; When the emotion trend factor of the target object is less than the first preset threshold, trigger a warning notification to warn of the emotional change of the target object; When the emotional fluctuation range of the target object is greater than the second preset threshold, generate dynamic suggestion content based on the emotion label of the target object, the personality portrait of the target object, and the conversation context content of the target object in the current conversation scenario; and, Generate a visualization score based on the emotion label of the target object, the personality portrait of the target object, and the emotion trend factor of the target object.

[0012] In a second aspect, an embodiment of the present application further provides an intelligent dialogue device, where the device includes: A first acquisition module, configured to acquire the current dialogue content input by a target object and historical dialogue data, where the target object is any young child; A first determination module, configured to input the dialogue content into a pre-trained emotion recognition model to obtain the emotion label of the target object, and determine the personality portrait of the target object based on the historical dialogue data, where the emotion recognition model is pre-trained based on a large number of young child dialogue samples, and the model parameters of the emotion recognition model can be adjusted based on the personality portrait of the target object; A second determination module, configured to determine a candidate interaction strategy based on the emotion label of the target object and a preset mapping relationship, and optimize and learn the candidate interaction strategy based on the emotion label of the target object and the personality portrait of the target object to determine a target interaction strategy, where the preset mapping relationship is used to represent the mapping relationship between the emotion label and the interaction strategy; A first generation module, configured to generate a target reply content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy.

[0013] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor is configured to implement the intelligent dialogue method described in any embodiment of the first aspect when executing the program stored in the memory.

[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the intelligent dialogue method described in any embodiment of the first aspect when executed by a processor.

[0015] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: The method provided by the embodiment of the present application obtains the current conversation content and historical conversation data input by the target object, where the target object is any child; inputs the conversation content into a pre-trained emotion recognition model to obtain the emotion label of the target object, and determines the personality portrait of the target object based on the historical conversation data. The emotion recognition model is pre-trained based on a large number of child conversation samples, and the model parameters of the emotion recognition model can be adjusted based on the personality portrait of the target object; determines candidate interaction strategies based on the emotion label of the target object and a preset mapping relationship, and optimizes and learns the candidate interaction strategies based on the emotion label of the target object and the personality portrait of the target object to determine the target interaction strategy, where the preset mapping relationship is used to represent the mapping relationship between the emotion label and the interaction strategy; generates a target reply content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy. By the above method, the model parameters of the emotion recognition model can be adjusted using the personality portrait of the target object, so that the emotion recognition model can comprehensively consider the conversation characteristics and personality portrait of the target object to improve the accuracy of emotion recognition, thereby improving the accuracy of the intelligent conversation reply content for children; moreover, by combining the emotion label of the target object and the personality portrait of the target object, the candidate interaction strategies are optimized and learned, so as to select the target interaction strategy that best suits the target user from the candidate interaction strategies, thereby further improving the accuracy of the target reply content. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a flowchart of an intelligent conversation method provided by an embodiment of the present application; Figure 2 It is a flowchart of another intelligent conversation method provided by an embodiment of the present application; Figure 3 It is a structural diagram of an intelligent conversation device provided by an embodiment of the present application; Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the scope of protection of this application.

[0020] See Figure 1 , Figure 1 , which is a schematic flowchart of an intelligent dialogue method provided by an embodiment of this application. As Figure 1 shown, the intelligent dialogue method may include the following steps: Step S101: Obtain the current dialogue content input by the target object and the historical dialogue data, where the target object is any young child.

[0021] Specifically, the above-mentioned dialogue content may refer to the voice information currently input by the target object, or the text information currently input by the target object, or the voice information and text information currently input by the target object. This application does not make specific limitations. As an optional implementation manner, the above-mentioned dialogue content is the voice information currently input by the target object, and then through speech-to-text processing, the text information corresponding to the voice information is obtained, and then subsequent analysis and processing are performed. The above-mentioned historical dialogue data refers to the dialogue data generated by the target object before inputting the current dialogue content.

[0022] Step S102: Input the dialogue content into a pre-trained emotion recognition model to obtain the emotion label of the target object, and based on the historical dialogue data, determine the personality portrait of the target object, where the emotion recognition model is pre-trained based on a large number of young child dialogue samples, and the model parameters of the emotion recognition model can be adjusted based on the personality portrait of the target object.

[0023] Specifically, the above-mentioned emotion recognition model is pre-trained based on a large number of young child dialogue samples, so it can learn dialogue features such as the fuzzy pronunciation or incomplete sentences of young children during the training process; and the model parameters of the emotion recognition model can be adjusted based on the personality portrait of the target object, so that during the process of recognizing emotions, the emotion recognition model can comprehensively recognize based on the personality portrait of the target object, improving the accuracy of emotion recognition. The emotion label of the above-mentioned target object is the output result of the emotion recognition model, which is used to represent the current emotional state of the target object, such as "happy", "sad", etc. The personality portrait of the above-mentioned target object refers to the personality indicators of the target object obtained by analyzing the historical dialogue data of the target object, such as activity level, expression tendency, and emotional stability.

[0024] Step S103: Based on the emotion label of the target object and the preset mapping relationship, determine the candidate interaction strategies, and optimize and learn the candidate interaction strategies based on the emotion label of the target object and the personality portrait of the target object to determine the target interaction strategy, where the preset mapping relationship is used to represent the mapping relationship between the emotion label and the interaction strategy.

[0025] Specifically, the above preset mapping relationship is used to represent the mapping relationship between the emotion label and the interaction strategy. Usually, each emotion label can correspond to one or more interaction strategies. For example, when the emotion label is "sad", the interaction strategies can be "comforting conversation", "playing music", etc. The above candidate interaction strategies refer to one or more interaction strategies that have a mapping relationship with the emotion label of the target object. When optimizing and learning the candidate interaction strategies based on the emotion label of the target object and the personality portrait of the target object, optimization can be performed based on a preset optimization learning algorithm, so as to determine the optimal target interaction strategy from the candidate interaction strategies. The optimization learning algorithm here can be the Q-Learning algorithm, the online reinforcement learning algorithm based on temporal difference learning (State-Action-Reward-State-Action, abbreviated as SARSA), the Deep Deterministic Policy Gradient algorithm (abbreviated as DDPG), the Proximal Policy Optimization algorithm (abbreviated as PPO), etc.

[0026] Step S104: Generate the target reply content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy.

[0027] After determining the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy, the dialogue context content of the previous rounds of the target object can be obtained, and the context relevance score can be calculated. Then, using the dialogue generation model, learn the dialogue features of the target object, the personality portrait of the target object, the target interaction strategy, and the context relevance score, and comprehensively obtain the final target reply content and output it. Among them, the dialogue generation model here can be pre-trained in advance based on training samples annotated with children's dialogue features, children's personality portraits, interaction strategies, and context relevance scores, or can be fine-tuned using existing large speech models. The present application does not make specific limitations.

[0028] In this embodiment, the model parameters of the emotion recognition model can be adjusted using the personality portrait of the target object, so that the emotion recognition model can comprehensively consider the conversation characteristics and personality portrait of the target object to improve the accuracy of emotion recognition, thereby improving the accuracy of the intelligent conversation reply content for young children. Moreover, by combining the emotion labels of the target object and the personality portrait of the target object, the candidate interaction strategies are optimized and learned, so as to select the target interaction strategy that best suits the target user from the candidate interaction strategies, thereby further improving the accuracy of the target reply content.

[0029] In an alternative embodiment, the above step S102, inputting the conversation content into a pre-trained emotion recognition model to obtain the emotion label of the target object, includes: Input the voice information in the conversation content into the voice feature extraction sub-model to obtain the spectrum feature, pitch change, and volume intensity corresponding to the voice information, and determine the voice feature vector corresponding to the target object based on the spectrum feature, pitch change, and volume intensity corresponding to the voice information; Input the text information in the conversation content into the text feature extraction sub-model to obtain the word embedding, sentence length, and context score corresponding to the text information, and determine the text feature vector corresponding to the target object based on the word embedding, sentence length, and context score corresponding to the text information; Input the voice feature vector and the text feature vector into the emotion analysis sub-model to obtain the emotion label of the target object.

[0030] Specifically, the above emotion recognition model may include a voice feature extraction sub-model, a text feature extraction sub-model, and an emotion analysis sub-model. Among them, the voice feature extraction sub-model can be implemented using an improved unsupervised speech pre-training model, the Wav2Vec 2.0 model (that is, a young child voice enhancement layer is added to the original Wav2Vec 2.0 model) or other models. The text feature extraction sub-model can be implemented using a lightweight DistilBERT model (with a parameter quantity of approximately 66M) or other models. The emotion analysis sub-model can be implemented using a transfome model or other models.

[0031] When inputting the conversation content into a pre-trained emotion recognition model to obtain the emotion label of the target object, the voice information in the conversation content can be input into the voice feature extraction sub-model to obtain the spectral features, pitch changes, and volume intensities corresponding to the voice information, and based on the spectral features, pitch changes, and volume intensities corresponding to the voice information, determine the voice feature vector corresponding to the target object. The voice information here can be the voice conversation content directly input by the target object or the voice information converted from the text conversation content directly input by the target object. At the same time, the text information in the conversation content can also be input into the text feature extraction sub-model to obtain the word embeddings, sentence lengths, and context scores corresponding to the text information, and based on the word embeddings, sentence lengths, and context scores corresponding to the text information, determine the text feature vector corresponding to the target object. The text information here can be the text conversation content directly input by the target object or the text information converted from the voice conversation content directly input by the target object. Then, the voice feature vector and the text feature vector can be input into the emotion analysis sub-model for feature fusion and sentiment analysis to obtain the emotion label of the target object.

[0032] In this embodiment, the voice feature extraction sub-model and the text feature extraction sub-model in the emotion recognition model can be used to extract the voice features and text features of the target object respectively, so that the voice features and text features of the children's conversation can be fully considered during emotion recognition, improving the accuracy of emotion recognition.

[0033] In an alternative embodiment, the above step of inputting the voice feature vector and the text feature vector into the emotion analysis sub-model to obtain the emotion label of the target object includes: Input the voice feature vector and the text feature vector into the emotion analysis sub-model, and use the emotion analysis sub-model to perform feature fusion according to the formula to obtain the emotion label of the target object; where, represents the emotion label of the target object, represents the dynamic attention weight, represents the normalization function, represents the weight matrix of the voice features, represents the voice feature vector, represents the weight matrix of the text features, represents the text feature vector.

[0034] Specifically, when using the emotion analysis sub-model to obtain the emotion label of the target object, the voice feature vector and the text feature vector can be input into the emotion analysis sub-model, and the emotion analysis sub-model can perform feature fusion according to the formula to obtain the emotion label of the target object. Among them, Represents the emotion label of the target object, that is, the current emotional state of the target object (such as the numerical value corresponding to "sadness", etc.). Represents the dynamic attention weight, whose value range is between 0 and 1, which determines the relative importance of speech features and text features. The initial value can be 0.6 and is dynamically adjusted according to the conversation frequency. Represents the normalization function to ensure that the output value is between 0 and 1 for easy classification. Represents the weight matrix of speech features, which is used to map speech features to the emotion space, and its size can depend on the number of layers of the improved Wav2Vec 2.0 model (such as 128×64, etc.). Represents the speech feature vector, which can include spectral features (represented by MFCC), pitch variation (represented by ΔPitch), and volume intensity (represented by Volume). Represents the weight matrix of text features, which is used to adjust the emotional relevance of text information, and its size can match the output of the DistilBERT model (such as 768×64, etc.). Represents the text feature vector, which can include word embeddings (represented by Word_emb), sentence length (represented by Sent_len), and context score (represented by Context_score).

[0035] In this embodiment, by using the attention mechanism to dynamically fuse the speech features and text features of the target object for extraction, the speech features and text features of the conversation of young children can be fully considered during emotion recognition, improving the accuracy of emotion recognition.

[0036] In an alternative embodiment, in step S102 above, based on historical conversation data, determine the personality portrait of the target object, including: Based on historical conversation data, calculate the activity level, expression tendency, and emotional stability corresponding to the target object respectively. Among them, the activity level is used to characterize the active degree of the target object's participation in the conversation, the expression tendency is used to characterize the relative dependence degree of the target object on tone and vocabulary in language expression, and the emotional stability is used to characterize the degree of emotional fluctuation of the target object. Based on the activity level, expression tendency, and emotional stability corresponding to the target object, determine the personality portrait of the target object.

[0037] Specifically, based on historical conversation data, calculate the activity level, expression tendency, and emotional stability corresponding to the target object respectively. Among them, the activity level can be used to characterize the active degree of the target object's participation in the conversation, and it can be calculated using the following formula: ; Among them, Represents the activity level, Represents the total interaction duration, and represents the total number of interactions of the target object within . Activity level High indicates that the target object talks frequently, showing an extroverted and proactive communication tendency, and may like activities with strong interactivity. Activity level Low indicates that the target object talks less, showing introverted, cautious or quiet characteristics, and may be more suitable for low-stimulus interaction strategies. By calculating the activity level , it can guide strategy selection. For example, when the activity level is low, the system is prompted to select non-forcing and quiet interaction strategies (such as telling stories instead of high-interaction games, etc.). It can also reflect the interaction willingness of the target object, help the system judge the participation degree of the target object, and thus adjust the interaction rhythm or strategy intensity.

[0038] The expression tendency can be used to characterize the relative dependence degree of the target object on tone and vocabulary in language expression, and it can be calculated by the following formula: ; where represents the expression tendency, represents the tone weight, which can be calculated based on voice features (such as pitch change, speech rate, volume, etc.) in historical conversation data, and reflects the intensity of the target object expressing emotions through tone. represents the vocabulary weight, which can be calculated based on the complexity of historical conversation data, the usage frequency of emotional vocabulary, etc., and reflects the degree of the target object expressing emotions through vocabulary. Expression tendency High indicates that the target object relies more on tone (such as pitch, speech rate) to express emotions, uses fewer words or the emotional vocabulary is not obvious, which may reflect introverted or implicit emotional expression characteristics. Expression tendency Low indicates that the target object relies more on vocabulary content to express emotions, with less tone change, which may reflect an extroverted or direct expression style. By calculating the expression tendency , the emotion recognition result can be optimized. For example, when the E value is high, the system is prompted to pay more attention to voice features (such as pitch, speech rate, etc.) rather than text features to improve the accuracy of emotion recognition; and, the model parameters of the emotion recognition model can also be adjusted personalized. For example, the system reduces the tone detection threshold (Pitchmin = 0.05) for child A to capture subtle pitch changes and adapt to its high expression tendency.

[0039] Emotional stability can be used to characterize the degree of emotional fluctuation of the target object, and it can be calculated by the following formula: ; where represents the degree of emotional fluctuation, and respectively represent the emotion values in two consecutive interactions at time t-1 and time t (for example, the emotion intensity or category probability estimated by the emotion recognition model, etc.). represents the variance function, which calculates the emotion value and the degree of fluctuation over time. The degree of emotion fluctuation being high indicates that the target object has small emotional changes and shows a stable emotional state, which may reflect characteristics such as introversion or better emotional control. The degree of emotion fluctuation being low indicates that the target object has large emotional fluctuations and is easily affected by external stimuli, showing emotional instability. By calculating the degree of emotion fluctuation , the emotion recognition sensitivity can be adjusted. For example, when the S value is low, the system is prompted to reduce attention to minor emotional fluctuations and focus on the main emotional state (such as "sadness") to improve the robustness of recognition; moreover, it can also guide strategy selection. For example, a young child with a stable emotion may require a milder interaction strategy, while a young child with large emotional fluctuations may require a more dynamic interaction strategy.

[0040] Next, based on the activity level, expression tendency, and emotional stability corresponding to the target object, a personality portrait of the target object is determined .

[0041] In this embodiment, based on the historical conversation data, the personality portrait of the target object can be accurately determined, which is convenient for subsequently adjusting the model parameters of the emotion recognition model based on the personality portrait of the target object and outputting the target reply content that conforms to the personality portrait of the target object.

[0042] In an alternative embodiment, after the step S102 of determining the personality portrait of the target object based on the historical conversation data, the method further includes: Based on the personality portrait of the target object, the first parameter or the second parameter in the emotion recognition model is adjusted so that the emotion label output by the adjusted emotion recognition model is closer to the true emotion label of the target object and matches the personality portrait of the target object, where the first parameter is the relevant parameter corresponding to the weight matrix of the voice feature, and the second parameter is the relevant parameter corresponding to the weight matrix of the text feature.

[0043] Specifically, after determining the personality portrait of the target object, the first parameter or the second parameter in the emotion recognition model can be adjusted based on the personality portrait of the target object. For example, for an introverted young child, the relevant parameter corresponding to the weight matrix of the voice feature (such as the tone threshold ) can be reduced; for an extroverted young child, the relevant parameter corresponding to the weight matrix of the text feature (such as the word weight ). This can make the emotion labels output by the adjusted emotion recognition model closer to the true emotion labels of the target object and match the personality portrait of the target object. When adjusting the model parameters in the emotion recognition model, continuous updates can be performed through online learning, and the update method is as follows: ; Among them, represents the updated model parameters, which are used for subsequent emotion recognition and dialogue generation. represents the model parameters before the update and serves as the basis for the update. represents the learning rate, whose value range is from 0 to 1 (such as 0.01, etc.), and it controls the parameter update step size to avoid overfitting. represents the predicted emotion value, which is output by the current emotion recognition model. represents the true emotion value, which is usually obtained through annotation or secondary inference of the model. represents the personality portrait of the target object, which is used to adjust the loss weight to make the model pay more attention to personality characteristics. represents the mean squared error loss, which is used to measure the gap between the predicted emotion value and the true emotion value.

[0044] Through the above method, based on the personality portrait of the target object generated each time, the first parameter or the second parameter in the emotion recognition model can be continuously adjusted, so that the emotion labels output by the adjusted emotion recognition model are closer to the true emotion labels of the target object and match the personality portrait of the target object.

[0045] In an optional embodiment, the above step S103, optimizing and learning the candidate interaction strategies based on the emotion labels of the target object and the personality portrait of the target object to determine the target interaction strategy, includes: Using a preset reinforcement learning algorithm to optimize and learn the candidate interaction strategies, and determining the target interaction strategy from the candidate interaction strategies. Among them, the reinforcement learning algorithm uses the emotion labels of the target object and the personality portrait of the target object as states, the interaction strategies in the candidate interaction strategies as actions, and the degree of emotion alleviation as the reward, and selects the target interaction strategy by continuously iteratively calculating the value of different state-action pairs. The target interaction strategy is the interaction strategy corresponding to the optimal value.

[0046] Specifically, the Q-Learning algorithm can be used as the preset reinforcement learning algorithm to optimize and learn the candidate interaction strategies, and determine the target interaction strategy from the candidate interaction strategies. Among them, the state is , the action is , and the reward is the degree of emotion alleviation . The value Q update formula of the state-action pair is as follows: ; Among them, is the value of the state-action pair, representing the long-term return of selecting strategy in state . represents the current state, which is composed of the emotion label of the target object and the personality portrait of the target object (such as ). represents the current action, such as "telling a story", etc. represents the learning rate, whose value range is from 0 to 1 (such as 0.1, etc.), and it is used to control the Q-value update speed. represents the immediate reward, which is calculated by the degree of emotion alleviation (such as R = 0.8 means the emotion is improved by 80%). represents the discount factor, whose value range is from 0 to 1 (such as 0.9, etc.), and it is used to balance the immediate and future rewards. represents the optimal Q-value of all possible actions in the next state , which is used to predict the return of the future optimal strategy.

[0047] In the above way, the preset reinforcement learning algorithm can be used to optimize and learn the candidate interaction strategies, so as to determine the optimal target interaction strategy from the candidate interaction strategies.

[0048] In an optional embodiment, after the above step S104, generating the target reply content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy, the method further includes: Obtaining the emotion trend factor and the emotion fluctuation range of the target object; Triggering a warning notification when the emotion trend factor of the target object is less than the first preset threshold, so as to give a warning about the emotion change of the target object; When the emotion fluctuation range of the target object is greater than the second preset threshold, generating dynamic suggestion content based on the emotion label of the target object, the personality portrait of the target object, and the dialogue context content of the target object in the current dialogue scenario; and, Generating a visualization score based on the emotion label of the target object, the personality portrait of the target object, and the emotion trend factor of the target object.

[0049] Specifically, the above first preset threshold and the above second preset threshold can be set according to actual needs, and no specific limitation is made here. The above emotion trend factor refers to the emotion change trend of the target object in a relatively short period of time currently, such as changing from "happy" to "anxious" within 5 minutes, etc. This emotion trend factor can use the exponential smoothing method for the emotion historical data is obtained through calculation. As an optional implementation, when the emotional trend factor is satisfied, a warning notification (such as "The child's emotion may decline, it is recommended to pay attention", etc.) can be triggered, and the refresh rate can be 0.5 seconds / time. The above emotional fluctuation range refers to the degree of emotional change of the target object. As an optional implementation, when the emotional fluctuation range is satisfied, based on the emotional label E of the target object, the personality portrait P of the target object, and the conversation context content of the target object in the current conversation scenario , through the matching matrix , dynamic suggestion content (such as "Feeling down after studying, it is recommended to praise the effort", etc.) can be generated. Of course, the following formula can also be used to generate a visual score in real time based on the emotional label E of the target object, the personality portrait P of the target object, and the emotional trend factor : ; Among them, represents the visual score, which synthesizes three parameters: the emotional label E of the target object, the personality portrait P of the target object, and the emotional trend factor , and the value range is from 0 to 1, which is used for dashboard display. represents the personality weight, and its value range is from 0 to 1 (such as 0.5, etc.), which is used to adjust the influence of personality on the score. represents the trend weight, and its value range is from 0 to 1 (such as 0.4, etc.), which is used to highlight the importance of trend analysis. represents , and the maximum value among the three, which is used for normalization processing. represents the emotional trend factor, and its calculation formula is ). represents the emotional value at time t (such as 0.8). represents the emotional value at time t - 1, and the initial value is 0. represents the smoothing coefficient, and its value range is from 0 to 1 (such as 0.3, etc.), which is used to balance the proportion of current and historical data.

[0050] Through the above methods, the emotions of the target object can be deeply analyzed, and warning notifications can be triggered and dynamic suggestion content can be generated accordingly, providing accurate monitoring and intervention support for parents.

[0051] In an optional embodiment, a smart speaker can be used as a smart conversation device to provide emotional management and educational support for children aged 3 to 6. The smart conversation device may include an emotion recognition module, a personality adaptive fine-tuning module, an interaction strategy matching module, a context adaptive generation module, and a parent smart support module. The specific implementation steps are as follows: Step S201: The emotion recognition module collects the voice and text inputs of the child, extracts the voice feature vector and text feature vector respectively using the improved Wav2Vec 2.0 model and the lightweight DistilBERT model, and generates an emotion label through the attention mechanism.

[0052] Device preparation: Use a smart speaker equipped with a microphone (such as model XYZ-200, sampling rate 16kHz), run an embedded Linux system, and install the software of this system.

[0053] Specific actions: 1. The child says "I don't like kindergarten" in front of the smart speaker, and the microphone collects the voice signal with a frame length of 10ms to generate an original audio file (WAV format).

[0054] 2. The voice feature extraction sub-module in the emotion recognition module loads the improved Wav2Vec 2.0 model (with about 50M parameters), calculates the spectral features (MFCC, 40 dimensions), pitch change ( , unit Hz / s), and volume intensity (Volume, unit dB), and outputs the voice feature vector .

[0055] 3. The text feature extraction sub-module in the emotion recognition module converts the audio to text "I don't like kindergarten" through speech-to-text, loads the lightweight DistilBERT model (with about 66M parameters), and extracts the word embedding ( , 768 dimensions), sentence length ( , 4 words), and context score ( , calculated based on the previous sentence "Went to kindergarten today" dialogue content), and outputs the text feature vector .

[0056] Result: The emotion analysis sub-module in the emotion recognition module fuses and the text feature vector , through the formula , and outputs an emotion label of "sad" with a confidence of 0.85.

[0057] Step S202: The personality adaptive fine-tuning module analyzes the child's dialogue data, constructs a personality portrait, and dynamically adjusts the recognition threshold and dialogue style according to the personality portrait.

[0058] Device Preparation: The smart speaker is connected to the cloud database, which stores the historical conversation data of children (such as 10 interaction records).

[0059] Specific Actions: 1. Extract the conversation records of the child from the database (such as 5 short sentences, slow speech rate), calculate the activity level , expression tendency , emotional stability , generate a personality profile , and determine that the child is an "introverted type".

[0060] 2. The adaptive fine-tuning mechanism adjusts the parameters according to P: lower the tone threshold , reduce the dependence on pitch changes; update the model parameters through the formula ( ), and optimize the emotion recognition model.

[0061] Result: After adjusting the model parameters, the confidence level of the system in recognizing "sadness" in children can be increased from 0.85 to 0.90.

[0062] Step S203: The interaction strategy matching module optimizes the output interaction strategy based on the emotion label and personality profile using the Q-Learning optimization algorithm.

[0063] Device Preparation: The smart speaker loads the emotion-strategy mapping matrix (5×20 dimensions, and the pre-training data can be obtained based on Piaget's theory).

[0064] Specific Actions: 1. Input the emotion label "sadness" and the personality profile , query the matrix , and obtain candidate interaction strategies, such as "comforting conversation" and "playing music".

[0065] 2. Q-Learning Optimization: State , action a = "comforting conversation", reward (based on subsequent emotion improvement), update the Q value .

[0066] 3. Output Strategy: The system selects "comforting conversation" as the target interaction strategy and generates candidate response content, such as "Are you feeling a little unhappy?" etc.

[0067] Result: After the strategy is executed, the child's emotion improves to "calm", and the adjustment time is 1 minute.

[0068] Step S204: The context adaptation generation module combines context adaptation generation and lightweight Transformer context enhancement to generate the target response content.

[0069] Specific actions: 1. Input dialogue features and personality portraits , and dynamically adjust the generation strategy: for short sentence inputs (such as "Don't want to play"), prefer concise responses (such as "Okay, let's change one"); for long sentences (such as "I painted pictures in kindergarten today"), generate detailed responses (such as "That's great. What did you paint?"). Combine the current multi-turn context of the child, such as , and use a lightweight Transformer encoder to calculate the context relevance score , and incorporate it into the generation process.

[0070] 2. According to the input short sentence "Don't want to play", through the formula , generate the target response content, such as "Okay, let's change one". Among them, represents the final dialogue response content. represents the candidate response content, generated by the decoder. is the conditional probability, indicating the fitness of the response content under the dialogue features , context sequence and personality portrait P. represents the dialogue features, which fuse the speech feature vector and the text feature vector . represents the context sequence, such as the previous 3 rounds of dialogue features and emotion labels (such as . represents the personality portrait, which guides the tone adjustment. represents the decoder weight matrix, with a size matching the vocabulary (such as 512×vocab_size, etc.). represents the context relevance weight, with a value range from 0 to 1 (such as 0.2), used to adjust the influence of the context on the response. represents the context relevance score, which is calculated by the Transformer. represents selecting the response with the highest comprehensive probability.

[0071] 3. Adjust the tone to a gentle type (suitable for an introverted personality), and output the audio (speech rate 1.2 times normal, volume -3dB).

[0072] Result: Generate the response content "Okay".

[0073] Step S205: The parental intelligent support module analyzes the emotional trend through the exponential smoothing method, calculates the trend factor, generates trigger warnings and scenario suggestions, and simultaneously generates a growth report to support parental intervention.

[0074] Device preparation: The smart speaker is connected to the parental mobile application to synchronize data in real time.

[0075] Specific actions: 1. The emotional trend analysis module processes historical emotional data, such as , calculates the trend factor (where ), and obtains .

[0076] 2. Check the trigger conditions: When the emotional fluctuation , according to the emotional label "sad", the personality portrait and the dialogue context content in the current dialogue scenario, generate dynamic suggestion content, such as "worthy of praise", etc., and push it to the parental mobile application.

[0077] 3. Generate a weekly report: Such as emotional stability , participation .

[0078] Result: The parents receive the suggestions and implement them, and the emotions of the young children are improved, with an acceptance rate of 85%.

[0079] Thus, by improving the Wav2Vec 2.0 model and the lightweight DistilBERT model, extracting speech features and text features, and using the attention mechanism for fusion, the recognition accuracy is improved. This adaptability solves the problem of insufficient understanding of the dialogue features of young children in the prior art and lays a precise foundation for subsequent educational intervention. Moreover, by constructing a personality portrait of young children (such as indicators like activity level, expression tendency, emotional stability, etc.), the emotional recognition threshold and dialogue style are dynamically adjusted. For example, the tone threshold is lowered for introverted young children, and the vocabulary weight is increased for extroverted young children, etc., to enhance individual adaptability. This online learning mechanism ensures that the model parameters can be optimized with interactions, surpassing traditional unified models, and significantly enhancing the personalization and educational effect of the interaction. In addition, by introducing an emotional trend analysis and scenario trigger mechanism, the exponential smoothing method is used to predict emotional changes, and warnings and suggestions are triggered when the emotional fluctuation exceeds 0.2 or the emotional trend is lower than -0.15, and precise intervention strategies (such as "worthy of praise") can be generated in combination with the scenario context (such as "after learning"), which is more intelligent than static suggestions, enables parents to better understand the emotional changes of young children, and effectively improves the user experience.

[0080] See Figure 3 , Figure 3 which is a schematic structural diagram of an intelligent dialogue device provided by an embodiment of the present application. As Figure 3As shown in the figure, the intelligent dialogue device 300 includes: A first acquisition module 301, configured to acquire the current dialogue content and historical dialogue data input by a target object, where the target object is any child; A first determination module 302, configured to input the dialogue content into a pre-trained emotion recognition model to obtain an emotion label of the target object, and determine a personality portrait of the target object based on the historical dialogue data. The emotion recognition model is pre-trained based on a large number of child dialogue samples, and the model parameters of the emotion recognition model can be adjusted based on the personality portrait of the target object; A second determination module 303, configured to determine a candidate interaction strategy based on the emotion label of the target object and a preset mapping relationship, and optimize and learn the candidate interaction strategy based on the emotion label of the target object and the personality portrait of the target object to determine a target interaction strategy, where the preset mapping relationship is used to represent the mapping relationship between the emotion label and the interaction strategy; A first generation module 304, configured to generate a target reply content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy.

[0081] Further, the emotion recognition model includes a speech feature extraction sub-model, a text feature extraction sub-model, and an emotion analysis sub-model; the first determination module 302 includes: A first input sub-module, configured to input the speech information in the dialogue content into the speech feature extraction sub-model to obtain the spectral features, pitch changes, and volume intensities corresponding to the speech information, and determine a speech feature vector corresponding to the target object based on the spectral features, pitch changes, and volume intensities corresponding to the speech information; A second input sub-module, configured to input the text information in the dialogue content into the text feature extraction sub-model to obtain the word embeddings, sentence lengths, and context scores corresponding to the text information, and determine a text feature vector corresponding to the target object based on the word embeddings, sentence lengths, and context scores corresponding to the text information; A third input sub-module, configured to input the speech feature vector and the text feature vector into the emotion analysis sub-model to obtain the emotion label of the target object.

[0082] Further, the third input sub-module is specifically configured to: Input the speech feature vector and the text feature vector into the emotion analysis sub-model, and use the emotion analysis sub-model to perform feature fusion according to the formula to obtain the emotion label of the target object; where represents the emotion label of the target object, represents the dynamic attention weight, represents the normalization function, The weight matrix representing the speech features, The speech feature vector, The weight matrix representing the text features, The text feature vector.

[0083] Further, the first determination module 302 further includes: A calculation sub-module, configured to calculate, based on historical conversation data, the activity level, expression tendency, and emotional stability corresponding to the target object respectively, where the activity level is used to characterize the active degree of the target object participating in the conversation, the expression tendency is used to characterize the relative dependence degree of the target object on tone and vocabulary in language expression, and the emotional stability is used to characterize the emotional fluctuation degree of the target object; A determination sub-module, configured to determine the personality portrait of the target object based on the activity level, expression tendency, and emotional stability corresponding to the target object.

[0084] Further, the intelligent conversation device 300 further includes: An adjustment module, configured to adjust the first parameter or the second parameter in the emotion recognition model based on the personality portrait of the target object, so that the emotion label output by the adjusted emotion recognition model is closer to the true emotion label of the target object and matches the personality portrait of the target object, where the first parameter is the relevant parameter corresponding to the weight matrix of the speech features, and the second parameter is the relevant parameter corresponding to the weight matrix of the text features.

[0085] Further, the second determination module 303 includes: An optimization learning sub-module, configured to use a preset reinforcement learning algorithm to optimize and learn the candidate interaction strategies, and determine the target interaction strategy from the candidate interaction strategies, where the reinforcement learning algorithm uses the emotion label of the target object and the personality portrait of the target object as states, the interaction strategies in the candidate interaction strategies as actions, and the emotion alleviation degree as the reward, and selects the target interaction strategy by continuously iteratively calculating the values of different state-action pairs, and the target interaction strategy is the interaction strategy corresponding to the optimal value.

[0086] Further, the intelligent conversation device 300 further includes: A second acquisition module, configured to acquire the emotion trend factor and the emotion fluctuation range of the target object; A trigger module, configured to trigger a warning notification to warn of the emotion change of the target object when the emotion trend factor of the target object is less than a first preset threshold; A second generation module, configured to generate dynamic suggestion content based on the emotion label of the target object, the personality portrait of the target object, and the conversation context content of the target object in the current conversation scenario when the emotion fluctuation range of the target object is greater than a second preset threshold; and, A third generation module, configured to generate a visualization score based on the emotion label of the target object, the personality portrait of the target object, and the emotion trend factor of the target object.

[0087] It should be noted that the intelligent dialogue device 300 can implement the steps of the intelligent dialogue method provided in any of the foregoing method embodiments, and can achieve the same technical effects, which will not be elaborated herein one by one.

[0088] As Figure 4 shown, an embodiment of the present application further provides an electronic device, including a processor 411, a communication interface 412, a memory 413, and a communication bus 414. Among them, the processor 411, the communication interface 412, and the memory 413 communicate with each other through the communication bus 414. The memory 413 is used to store a computer program. In an embodiment of the present application, when the processor 411 executes the program stored on the memory 413, it implements the intelligent dialogue method provided in any of the foregoing method embodiments.

[0089] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the intelligent dialogue method provided in any of the foregoing method embodiments.

[0090] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0091] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. An intelligent dialogue method, characterized in that, The method includes: Obtaining the current conversation content and historical conversation data input by the target object, where the target object is any child; Inputting the conversation content into a pre-trained emotion recognition model to obtain the emotion label of the target object, and determining the personality portrait of the target object based on the historical conversation data, where the emotion recognition model is pre-trained based on a large number of child conversation samples, and the model parameters of the emotion recognition model can be adjusted based on the personality portrait of the target object; Determining candidate interaction strategies based on the emotion label of the target object and a preset mapping relationship, and optimizing and learning the candidate interaction strategies based on the emotion label of the target object and the personality portrait of the target object to determine the target interaction strategy, where the preset mapping relationship is used to represent the mapping relationship between the emotion label and the interaction strategy; Generating target response content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy.

2. The intelligent dialogue method according to claim 1, wherein The emotion recognition model includes a speech feature extraction sub-model, a text feature extraction sub-model, and an emotion analysis sub-model; The step of inputting the conversation content into a pre-trained emotion recognition model to obtain the emotion label of the target object includes: Inputting the speech information in the conversation content into the speech feature extraction sub-model to obtain the spectral features, pitch changes, and volume intensities corresponding to the speech information, and determining the speech feature vector corresponding to the target object based on the spectral features, pitch changes, and volume intensities corresponding to the speech information; Inputting the text information in the conversation content into the text feature extraction sub-model to obtain the word embeddings, sentence lengths, and context scores corresponding to the text information, and determining the text feature vector corresponding to the target object based on the word embeddings, sentence lengths, and context scores corresponding to the text information; Inputting the speech feature vector and the text feature vector into the emotion analysis sub-model to obtain the emotion label of the target object.

3. The intelligent dialogue method according to claim 2, wherein The step of inputting the speech feature vector and the text feature vector into the emotion analysis sub-model to obtain the emotion label of the target object includes: Input the speech feature vector and the text feature vector into the emotion analysis sub-model, and use the emotion analysis sub-model to perform feature fusion according to the formula to obtain the emotion label of the target object; Among them, represents the emotion label of the target object, represents the dynamic attention weight, represents the normalization function, represents the weight matrix of the speech feature, represents the speech feature vector, represents the weight matrix of the text feature, represents the text feature vector.

4. The intelligent dialogue method according to claim 1, wherein, The step of determining the personality portrait of the target object based on the historical conversation data includes: Calculating the activity level, expression tendency, and emotional stability corresponding to the target object based on the historical conversation data, where the activity level is used to represent the active degree of the target object's participation in the conversation, the expression tendency is used to represent the relative dependence degree of the target object on tone and vocabulary in language expression, and the emotional stability is used to represent the degree of emotional fluctuation of the target object; Determining the personality portrait of the target object based on the activity level, expression tendency, and emotional stability corresponding to the target object.

5. The intelligent dialogue method according to claim 4, wherein After determining the personality portrait of the target object based on the historical conversation data, the method further includes: Based on the personality portrait of the target object, adjust the first parameter or the second parameter in the emotion recognition model, so that the emotion label output by the adjusted emotion recognition model is closer to the true emotion label of the target object and matches the personality portrait of the target object, where the first parameter is the relevant parameter corresponding to the weight matrix of the speech feature, and the second parameter is the relevant parameter corresponding to the weight matrix of the text feature.

6. The intelligent dialogue method according to claim 1, wherein Optimally learning the candidate interaction strategies based on the emotion label of the target object and the personality portrait of the target object to determine the target interaction strategy, including: Using a preset reinforcement learning algorithm to optimally learn the candidate interaction strategies and determine the target interaction strategy from the candidate interaction strategies, where the reinforcement learning algorithm uses the emotion label of the target object and the personality portrait of the target object as states, the interaction strategies in the candidate interaction strategies as actions, and the degree of emotion alleviation as the reward, and selects the target interaction strategy by continuously iteratively calculating the values of different state-action pairs, and the target interaction strategy is the interaction strategy corresponding to the optimal value.

7. The intelligent dialogue method according to claim 1, wherein After generating the target reply content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy, the method further includes: Obtaining the emotion trend factor and the emotion fluctuation range of the target object; When the emotion trend factor of the target object is less than a first preset threshold, triggering a warning notification to warn of the emotion change of the target object; When the emotion fluctuation range of the target object is greater than a second preset threshold, generating dynamic suggestion content based on the emotion label of the target object, the personality portrait of the target object, and the conversation context content of the target object in the current conversation scenario; and Generating a visualization score based on the emotion label of the target object, the personality portrait of the target object, and the emotion trend factor of the target object.

8. An intelligent dialogue device, characterized in that, The device includes: A first acquisition module, configured to acquire the conversation content currently input by the target object and the historical conversation data, where the target object is any young child; A first determination module, configured to input the conversation content into a pre-trained emotion recognition model to obtain the emotion label of the target object, and determine the personality portrait of the target object based on the historical conversation data, where the emotion recognition model is pre-trained based on a large number of young child conversation samples, and the model parameters of the emotion recognition model can be adjusted based on the personality portrait of the target object; A second determination module, configured to determine candidate interaction strategies based on the emotion label of the target object and a preset mapping relationship, and optimally learn the candidate interaction strategies based on the emotion label of the target object and the personality portrait of the target object to determine the target interaction strategy, where the preset mapping relationship is used to represent the mapping relationship between the emotion label and the interaction strategy; A first generation module, configured to generate target reply content based on the emotion label of the target object, the personality portrait of the target object, and the target interaction strategy.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is configured to implement the intelligent dialogue method according to any one of claims 1-7 when executing the program stored on the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the intelligent dialogue method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Dynamic interaction method, server, electronic equipment and storage medium

    CN112148850A

  • Emotional support man-machine conversation method and system based on emotional strategy matching

    CN116415596A

  • Intelligent interaction method and device, electronic equipment and readable storage medium

    CN119166766A

Cited By

  • Virtual user simulation method and system for dialogue test

    CN121766451A

  • Emotion recognition method, system and device for mental health of children and storage medium

    CN121790020A