Scene-based emotional interactive accompanying doll system and method
Through multimodal sensing data analysis and user portrait construction, emotional evolution, scene interaction and physiological regulation units are designed, which solves the shortcomings of existing emotional interaction dolls in identifying user emotions and scene types, realizes personalized emotional interaction, and improves interactive experience and efficiency.
Patent Information
- Application Number
- CN202510276780.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Existing emotional interaction dolls find it difficult to accurately identify user emotional states and scene types, and the interaction strategy lacks personalization and dynamicity, and cannot effectively achieve continuous emotional companionship.
Through multimodal sensing data acquisition and feature analysis, emotional feature vectors and scene feature vectors are generated, user portraits are constructed, emotional evolution units, scene interaction units and physiological adjustment units are designed, and personalized emotional interaction models are realized.
It realizes personalized and accurate response of emotional interaction dolls, accurately identify users' emotional state and scene needs, dynamically adjusts interaction strategies, provides a natural and smooth interactive experience, and improves the execution efficiency of emotional guidance.
Smart Images

Figure CN120179069A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a scenario-based emotional interactive companion doll system and method. Background Art
[0002] Currently, most intelligent doll products on the market have basic voice interaction and motion control functions, but there are still some deficiencies in emotional interaction. Existing emotional interaction dolls usually use a single sensor to collect user data, making it difficult to comprehensively perceive the user's emotional state and usage scenario, resulting in a lack of pertinence and adaptability in the interaction process. At the same time, the interaction strategies of these dolls are often preset fixed patterns and cannot be dynamically adjusted according to the user's personality characteristics and usage habits, making it difficult to form a continuous and effective emotional companionship effect.
[0003] In addition, existing emotional interaction dolls mainly use simple feature extraction and fusion methods when processing multi-modal data, and the analysis of emotional features and scene features is relatively rough. In terms of interaction decision-making, there is a lack of in-depth analysis and prediction of the user's emotional change law, making it difficult to achieve targeted emotional guidance. At the same time, the existing system rarely considers the influence of the user's physiological rhythm characteristics on the interaction process, affecting the naturalness and comfort of the interaction experience. Summary of the Invention
[0004] In view of the problems existing in the existing emotional interactive dolls, the present invention is proposed.
[0005] Therefore, the problem to be solved by the present invention is how to accurately identify the user's emotional state and scene type based on multi-modal sensing data, and realize intelligent emotional companionship functions including emotional evolution, scene interaction, and physiological regulation by constructing a personalized emotional interaction model.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, an embodiment of the present invention provides a scenario-based emotional interactive companion doll system, which includes a data acquisition module for collecting an interaction data group and a scenario data group to generate an extended data matrix; a feature analysis module for extracting features from the extended data matrix, generating an emotional feature vector and a scenario feature vector using a multi-modal feature network, determining the current emotional state of the user based on the emotional feature vector, and identifying the current scenario type based on the scenario feature vector; a portrait construction module for extracting the user's long-term usage habit features, emotional preference features, physiological rhythm features, and scenario preference features from an emotional memory library to construct a user portrait; an interaction decision module including an emotion evolution unit, a scenario interaction unit, and a physiological regulation unit; the emotion evolution unit analyzes the emotion change law based on the emotional preference features and the historical emotional state sequence, predicts the emotion change trend, and determines an emotion guidance target sequence; the scenario interaction unit generates an instruction screening strategy including an instruction priority sequence and a modality combination rule based on the emotion guidance target sequence and the current scenario type, screens the instructions corresponding to the current scenario according to the instruction priority sequence, and combines the screened instructions according to the modality combination rule to form a basic instruction sequence; the physiological regulation unit adjusts the basic instruction sequence according to the physiological rhythm features to generate an execution instruction sequence.
[0008] As a preferred solution of the scenario-based emotional interactive companion doll system of the present invention, it further includes an execution control module for controlling the voice module, expression module, and action module of the doll to complete emotional interaction behaviors based on the execution instruction sequence; a database module, where the database module includes an emotional memory library and a scenario-based emotional interaction instruction library, the emotional memory library stores a historical emotional feature vector sequence, an emotional state sequence, and a scenario type sequence, and the scenario-based emotional interaction instruction library stores a voice instruction set, an expression instruction set, and an action instruction set for different scenarios; the interaction data group includes facial images, heart rate data, stress data, and voice data, and the scenario data group includes environmental parameters, time features, location features, and surrounding item features.
[0009] As a preferred solution of the scenario-based emotional interactive companion doll system of the present invention, the generation of emotional feature vectors and scenario feature vectors using a multi-modal feature network includes the following steps: preprocess the extended data matrix to obtain a standardized data matrix, and divide the standardized data matrix into an interaction sub-matrix and a scenario sub-matrix; process the interaction sub-matrix and the scenario sub-matrix through a dual-branch feature extraction network, where the dual-branch feature extraction network includes an interaction feature branch and a scenario feature branch; among them, the interaction feature branch includes an emotional modality adaptation layer and an emotional fusion network, and the scenario feature branch includes a scenario encoder and a scenario semantic understanding network; use the emotional modality adaptation layer in the interaction feature branch to extract and transform features of the interaction sub-matrix, and use the emotional fusion network to integrate the transformed features to output emotional representation features; use the scenario encoder in the scenario feature branch to encode features of the scenario sub-matrix, and perform combined processing through the scenario semantic understanding network to output scenario semantic representation features; establish a spatio-temporal alignment module to connect the interaction feature branch and the scenario feature branch, and the spatio-temporal alignment module adopts an adaptive time window and an emotional delay compensation mechanism to input the emotional representation features and the scenario semantic representation features into an emotional mapping layer for feature mapping; input the mapped emotional representation features and scenario semantic representation features into an emotional feature generator and a scenario feature generator respectively, and the emotional feature generator and the scenario feature generator use a multi-head attention network for deep feature processing to generate emotional feature vectors and scenario feature vectors.
[0010] As a preferred solution of the scenario-based emotional interactive companion doll system of the present invention, the determination of the user's current emotional state and the recognition of the current scenario type include the following steps: establish an emotional state recognition model, where the emotional state recognition model includes an emotional type classifier and an emotional intensity calculation unit, and input the emotional feature vector into the emotional state recognition model; perform classification operations on the emotional feature vector through the emotional type classifier to obtain an emotional type probability vector; based on the comparison result between the maximum probability value in the emotional type probability vector and the emotional determination threshold, and in combination with the historical emotional state sequence, determine the basic emotional type; input the emotional feature vector into the emotional intensity calculation unit, and calculate the emotional intensity value based on the preset emotional intensity quantization standard; input the basic emotional type and the emotional intensity value into the emotional state mapping network to generate an emotional state feature curve, and determine the user's current emotional state based on the peak position and waveform characteristics of the emotional state feature curve; establish a scenario recognition model, input the scenario feature vector into the scenario recognition model, and calculate the matching degree between the scenario feature vector and each scenario type template based on the built-in scenario type template library; based on the comparison result between the maximum matching degree and the scenario matching threshold, and in combination with the historical scenario data, determine the current scenario type.
[0011] As a preferred embodiment of the scene-based emotional interactive companion doll system of the present invention, the working process of the portrait construction module is as follows: Read the historical emotional feature vector sequence, historical emotional state sequence, and historical scene type sequence in the emotional memory library, and splice the three sequences with the user's basic information to construct an original feature matrix; Perform time series decomposition on the historical emotional feature vector sequence in the original feature matrix, calculate the emotional state transition probability matrix, and obtain emotional preference features; Map the historical scene type sequence to the scene association matrix, calculate the scene transition probability and scene duration distribution, and obtain scene preference features; Perform frequency domain transformation on the historical emotional feature vector sequence, combine with the physiological parameters in the user's basic information, calculate the periodic feature parameters, and obtain physiological rhythm features; Align the historical emotional state sequence and the historical scene type sequence in time, calculate the joint distribution matrix, and obtain long-term usage habit features; Construct a three-layer user portrait based on emotional preference features, scene preference features, physiological rhythm features, and long-term usage habit features, and establish an inter-layer connection relationship through feature indexing.
[0012] As a preferred embodiment of the scene-based emotional interactive companion doll system of the present invention, the working process of the emotional evolution unit is as follows: Obtain the historical emotional state sequence and emotional preference features in the emotional memory library, and construct a hierarchical emotional state map; Perform time series feature analysis on the historical emotional state sequence, combine with the hierarchical emotional state map, extract emotional periodic patterns and conversion critical points, and form an emotional change feature set; Establish a two-channel emotional prediction framework, where the first channel predicts the state transition trend based on the emotional state transition probability matrix, and the second channel predicts the emotional change law based on the emotional change feature set. The output results of the two channels are weighted and fused to generate a basic emotional prediction vector; Design a scene modulation sub-unit to convert the current scene type into a scene context vector, and modulate the basic emotional prediction vector through a cross-attention network to generate a context-adaptive emotional prediction result; Design a steady-state evaluation sub-unit to construct an emotional steady-state interval based on the user's current emotional state, calculate the deviation degree between the context-adaptive emotional prediction result and the emotional steady-state interval, and generate an emotional deviation vector; Obtain at least two candidate emotional guidance directions from the preset doll companion therapy knowledge base according to the emotional deviation vector, and select the optimal emotional guidance direction based on the children's emotional development theory; Compare the optimal emotional guidance direction with the user's current emotional state, calculate the emotional guidance gradient, and output an emotional guidance target sequence with a time series control parameter.
[0013] As a preferred embodiment of the scene-based emotional interactive companion doll system of the present invention, the working process of the scene interaction unit is as follows: convert the timing control parameters in the emotional guidance target sequence into an interaction timing matrix, construct a scene constraint model based on the interaction timing matrix and the current scene type, and output an interaction constraint vector; use the interaction constraint vector to segment the emotional guidance target sequence, construct a sub-goal mapping network for each time window, dynamically associate the emotional guidance target with the scene restriction conditions, and generate a phased interaction strategy; perform timing combination and hierarchical analysis on the phased interaction strategy in sequence, extract the interaction intention features and timing constraint relationships, construct a priority mapping matrix and a modality collaboration graph, and generate an instruction screening strategy, where the instruction screening strategy includes an instruction priority sequence and a modality combination rule; retrieve the scene-based emotional interaction instruction library according to the instruction priority sequence to obtain a candidate instruction set, fuse the long-term usage habit features and scene preference features into an interaction evaluation vector, and use the interaction evaluation vector to evaluate and screen the candidate instruction set to obtain an optimal instruction set; construct an instruction execution dependency graph based on the modality combination rule; reorganize the optimal instruction set according to the instruction execution dependency graph, and control the timing relationship of instruction execution through a recursive gated network to generate a basic instruction sequence.
[0014] As a preferred embodiment of the scene-based emotional interactive companion doll system of the present invention, the working process of the physiological regulation unit is as follows: extract the cycle feature parameters from the physiological rhythm features, and divide the basic instruction sequence into corresponding execution cycle units in combination with the cycle feature parameters; generate an intensity adjustment matrix based on the physiological rhythm features, and perform intensity mapping on the execution parameters of the voice module, expression module, and action module in each execution cycle unit; adjust the execution parameters in the execution cycle unit according to the intensity adjustment matrix to obtain an adaptive instruction group; reorganize the adaptive instruction group according to the execution cycle unit to generate an execution instruction sequence.
[0015] Second aspect, embodiments of the present invention provide a scenario-based emotional interactive companion doll method, which includes collecting an interaction data set and a scenario data set through a multi-modal sensing system of the doll to generate an extended data matrix; performing feature extraction on the extended data matrix, using a multi-modal feature network to generate an emotional feature vector and a scenario feature vector, determining the current emotional state of the user based on the emotional feature vector, and identifying the current scenario type based on the scenario feature vector; extracting the long-term usage habit features, emotional preference features, physiological rhythm features, and scenario preference features of the user based on the user's basic information, the historical emotional feature vector sequence, historical emotional state sequence, and historical scenario type sequence stored in the emotional memory library, and constructing a user profile; constructing an interaction decision module according to the user profile, the user's current emotional state, and the current scenario type, where the interaction decision module includes an emotional evolution unit, a scenario interaction unit, and a physiological regulation unit; the emotional evolution unit analyzes the emotional change law based on the emotional preference features and the historical emotional state sequence, and combines the current emotional state and the current scenario type to predict the emotional change trend and determine the emotional guidance target sequence; the scenario interaction unit generates an instruction screening strategy including an instruction priority sequence and a modality combination rule based on the emotional guidance target sequence and the current scenario type, screens the multi-modal interaction instructions corresponding to the current scenario from the scenario-based emotional interaction instruction library according to the priority sequence, and combines the screened instructions according to the modality combination rule to form a basic instruction sequence; the physiological regulation unit adjusts the execution parameters of each instruction in the basic instruction sequence according to the physiological rhythm features to generate an execution instruction sequence; controlling the voice module, expression module, and action module of the doll to complete the emotional interaction behavior based on the execution instruction sequence, and storing the emotional feature vector, emotional state, and scenario type in the interaction scenario in the emotional memory library.
[0016] The beneficial effects of the present invention are as follows: The present invention realizes personalized and accurate responses of emotional interaction dolls. Based on the comprehensive analysis of multi-modal sensing and scenario perception, it can accurately identify the emotional state and scenario needs of users. Through the design of a dual-channel emotional prediction framework and a scenario modulation mechanism, the system can more accurately predict the emotional change trend and dynamically adjust the interaction strategy according to the scenario features. Based on the hierarchical emotional state map and the progressive emotional guidance scheme, the system can provide a natural and smooth interaction experience while ensuring the safety of emotional regulation. Through the modular multi-modal interaction design and the collaborative control of the recursive gated network, the system realizes the precise coordination of various interaction methods such as voice, expression, and action, and improves the execution efficiency of emotional guidance. Brief Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0018] Figure 1 It is a module connection diagram of a scenario-based emotional interactive companion doll system.
[0019] Figure 2 It is a multi-modal feature network processing flowchart of a scenario-based emotional interactive companion doll system.
[0020] Figure 3 It is a working flowchart of the emotional evolution unit of a scenario-based emotional interactive companion doll system.
[0021] Figure 4 It is a working flowchart of the scenario interaction unit of a scenario-based emotional interactive companion doll system. Specific Embodiments
[0022] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification.
[0023] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from this description. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0024] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selectively exclusive embodiment from other embodiments.
[0025] Embodiment 1, referring to Figures 1 to 4 , which is the first embodiment of the present invention. This embodiment provides a scenario-based emotional interactive companion doll system, and the module connection diagram is as shown in Figure 1 , including
[0026] A data acquisition module, which is used to collect an interaction data group and a scenario data group through the multi-modal sensing system of the doll. The interaction data group includes facial images, heart rate data, pressure data, and voice data, and the scenario data group includes environmental parameters, time characteristics, location characteristics, and surrounding item characteristics, and generates an extended data matrix.
[0027] It should be noted that the superiority of the interaction data group and the scenario data group lies in that: the facial image captures the user's emotional micro-expressions to provide visual emotional cues, the heart rate data reflects the changes in physiological states to provide objective emotional indicators, the stress data records the intimate interaction modes such as hugs and strokes, and the voice data contains emotional intonation information; at the same time, the environmental parameters affect the emotional baseline level, the time characteristics are associated with the daily emotional cycle, the location characteristics distinguish the different scenario requirements such as home and school, and the surrounding item characteristics provide context-based interaction references. This multi-dimensional data collection method breaks through the limitation of the single-trigger feedback of traditional dolls, constructs a comprehensive emotional perception basis by comprehensively analyzing physiological, behavioral, and environmental factors, and thus realizes the differentiated emotional companionship function for different scenarios.
[0028] A feature analysis module, which is used to extract and fuse features from the extended data matrix, generate an emotional feature vector and a scenario feature vector by using an improved multi-modal feature network, determine the current emotional state of the user based on the emotional feature vector, identify the current scenario type based on the scenario feature vector, and store the emotional feature vector, the emotional state, and the scenario type in the emotional memory bank;
[0029] Specifically, the processing flow chart of the multi-modal feature network is as Figure 2 shown. Generating an emotional feature vector and a scenario feature vector by using an improved multi-modal feature network includes the following steps: performing normalization processing and noise filtering on the extended data matrix to obtain a standardized data matrix, and dividing the standardized data matrix into an interaction sub-matrix and a scenario sub-matrix based on a time window and data types; processing the interaction sub-matrix and the scenario sub-matrix through a dual-branch feature extraction network, and the dual-branch feature extraction network includes an interaction feature branch and a scenario feature branch; wherein the interaction feature branch includes an emotional modality adaptation layer and an emotional fusion network, and the scenario feature branch includes a scenario encoder and a scenario semantic understanding network; using the emotional modality adaptation layer in the interaction feature branch to extract and transform features from the interaction sub-matrix, and then using the emotional fusion network to perform weight integration on the transformed features to output emotional representation features; using the scenario encoder in the scenario feature branch to perform feature encoding on the scenario sub-matrix, and performing combined processing through the scenario semantic understanding network to output scenario semantic representation features; establishing a spatio-temporal alignment module to connect the interaction feature branch and the scenario feature branch, and the spatio-temporal alignment module adopts an adaptive time window and an emotional delay compensation mechanism to input the emotional representation features and the scenario semantic representation features into a personalized emotional mapping layer for feature mapping; inputting the mapped emotional representation features and scenario semantic representation features into an emotional feature generator and a scenario feature generator respectively, and the emotional feature generator and the scenario feature generator adopt a multi-head attention network to perform deep feature processing, and generate an emotional feature vector and a scenario feature vector through an attention weight distribution and a feature selection mechanism.
[0030] Among them, the emotion modality adaptation layer includes a facial image feature analysis unit, a heart rate feature analysis unit, a stress feature analysis unit, and a speech feature analysis unit. Each unit processes the corresponding features and converts them into a unified dimension representation. The emotion fusion network uses an emotion context-aware attention mechanism to process the converted features, calculates the emotion feature weight coefficients, and generates emotion representation features. Correspondingly, the scene encoder includes an environmental parameter analysis unit, a time feature analysis unit, a space feature analysis unit, and an item feature analysis unit. Each unit extracts and processes the corresponding features respectively; the scene semantic understanding network combines the features through a feature fusion mechanism to generate scene semantic representation features.
[0031] It should be noted that the adaptive time window of the spatio-temporal alignment module dynamically adjusts the window parameters according to the state change rate; the emotion delay compensation mechanism predicts the current emotion trend by analyzing the historical emotion state sequence to achieve the temporal alignment of the emotion representation features and the scene semantic representation features.
[0032] It should be pointed out that the multi-head attention network in the emotion feature generator and the scene feature generator contains multiple attention heads. Each attention head independently learns the importance weights of different feature dimensions; the key features are highlighted through the attention weight allocation, and the most representative feature combinations are screened through the feature selection mechanism. The finally generated emotion feature vector includes dimensions such as emotion type, emotion intensity, emotion persistence, and emotion variability, and the scene feature vector includes dimensions such as scene type, scene similarity, scene duration, and scene change probability.
[0033] Secondly, determining the user's current emotion state based on the emotion feature vector and identifying the current scene type based on the scene feature vector includes the following steps: establishing an emotion state recognition model, which includes an emotion type classifier and an emotion intensity calculation unit, and inputting the emotion feature vector into the emotion state recognition model; classifying the emotion feature vector through the emotion type classifier to obtain an emotion type probability vector; based on the comparison result between the maximum probability value in the emotion type probability vector and the emotion determination threshold, combined with the historical emotion state sequence to determine the basic emotion type. The specific steps are as follows: when the maximum probability value in the emotion type probability vector is greater than the emotion determination threshold, select the type corresponding to the maximum probability value in the emotion type probability vector as the basic emotion type; when the maximum probability value in the emotion type probability vector is less than or equal to the emotion determination threshold, obtain the emotion state sequence within the preset time window from the emotion memory bank, and use the emotion type with the highest frequency of occurrence in the emotion state sequence as the basic emotion type.
[0034] Further, input the emotional feature vector into the emotional intensity calculation unit to calculate the emotional intensity value based on a preset emotional intensity quantization standard; wherein, the preset emotional intensity quantization standard refers to a quantization mechanism calculated based on the feature weights and emotional scoring threshold systems of different perceptual modalities, in combination with the physiological signal fluctuation range and age-related standardized reference values.
[0035] Further, input the basic emotional type and the emotional intensity value into the emotional state mapping network. The emotional state mapping network uses a non-linear conversion algorithm to comprehensively consider the emotional type weight coefficient and the intensity influence factor to generate an emotional state feature curve, and determines the user's current emotional state based on the peak position and waveform characteristics of the emotional state feature curve; establish a scene recognition model, input the scene feature vector into the scene recognition model, and calculate the matching degree between the scene feature vector and each scene type template based on the built-in scene type template library; based on the comparison result between the maximum matching degree and the scene matching threshold, combined with historical scene data to determine the current scene type. The specific steps are as follows: when the maximum value of the matching degree is greater than the scene matching threshold, select the type corresponding to the scene type template with the highest matching degree as the current scene type; when the maximum value of the matching degree is less than or equal to the scene matching threshold, retrieve the historical scene record with the highest similarity from the emotional memory library based on the scene feature vector, and determine the scene type corresponding to the historical scene record as the current scene type.
[0036] It should be noted that the reason for comparing the maximum probability value in the emotional type probability vector with the emotional determination threshold is that the maximum probability value in the probability vector output by the emotional type classifier represents the best matching degree between the current emotional feature and this type. When the maximum probability value is lower than the threshold, it indicates that the matching degree between the current emotional feature and all emotional types is not ideal enough, and further judgment needs to be made in combination with the historical emotional state sequence. The reason for comparing the maximum matching degree with the scene matching threshold is that the scene type corresponding to the maximum matching degree is the most likely current scene. When the maximum matching degree is not sufficient to meet the threshold requirement, it indicates that the current scene does not match the known templates well enough, and it is necessary to assist in the judgment by retrieving historical scene data.
[0037] Preferably, the above technical solution can effectively address the issues of unstable emotion and scene recognition faced by emotional interactive companion dolls in actual usage scenarios by introducing a comparison mechanism between the maximum probability value / match degree and the threshold, combined with the method of assisting judgment with historical data. Since the user may be a child, whose emotional expressions and behavior patterns are often not standardized enough, relying solely on real-time feature vectors may lead to large fluctuations in the recognition results. This solution uses a dual guarantee mechanism of threshold judgment and historical data reference, which not only ensures the fast response ability when features are obvious but also utilizes historical data to provide stable and reliable recognition results when features are not obvious, thereby enhancing the emotional interaction experience of the doll. At the same time, this solution can also gradually accumulate emotion and scene data, continuously enrich the historical database, enabling the doll to better adapt to the emotional expressions and usage habits of specific users and realizing personalized emotional companionship functions.
[0038] A portrait construction module, configured to extract the long-term usage habit features, emotional preference features, physiological rhythm features, and scene preference features of the user based on the user's basic information and the emotional memory library, and construct a user portrait.
[0039] Specifically, the working process of the portrait construction module is as follows: Read the historical emotional feature vector sequence, historical emotional state sequence, and historical scene type sequence in the emotional memory library, and splice the three sequences with the user's basic information to construct an original feature matrix; perform time series decomposition on the historical emotional feature vector sequence in the original feature matrix, calculate the emotional state transition probability matrix, where the emotional state transition probability matrix represents the migration law between emotional states, extract the principal eigenvector by performing eigenvalue decomposition on this matrix, and combine it with the emotional state residence duration distribution to form emotional preference features reflecting the user's emotional preference pattern; map the historical scene type sequence to a scene association matrix, calculate the scene transition probability and scene duration distribution, where the scene transition probability and scene duration distribution respectively characterize the scene switching tendency and scene interaction persistence, and integrate the two into a low-dimensional representation through a multidimensional scaling algorithm to form scene preference features representing the user's scene selection pattern.
[0040] Furthermore, perform a frequency-domain transformation on the historical emotional feature vector sequence, combine the physiological parameters in the user's basic information, calculate the periodic feature parameters, and obtain the physiological rhythm features. The specific steps are as follows: Segment the historical emotional feature vector sequence at equal time intervals, apply the fast Fourier transform to each segment sequence to obtain a frequency-domain feature matrix; Extract the main frequency components and energy distribution ratios in the frequency-domain feature matrix to construct a spectrum feature vector; Extract physiological parameters (such as age, gender, average heart rate, average sleep duration, etc.) from the user's basic information to construct a physiological basis feature vector; Input the spectrum feature vector and the physiological basis feature vector into an adaptive weighted fusion network to calculate the circadian cycle coefficient, emotional fluctuation cycle coefficient, and energy change cycle coefficient; Based on the above cycle coefficients, construct a multi-scale periodic feature parameter set including daily cycle features, weekly cycle features, and monthly cycle features; Perform a non-linear mapping on the multi-scale periodic feature parameter set to generate physiological rhythm features representing the user's physiological activity, emotional sensitivity, and interaction acceptance.
[0041] Furthermore, align the historical emotional state sequence and the historical scene type sequence in time, calculate the joint distribution matrix, where the joint distribution matrix represents the co-occurrence frequency of emotional states and scene types, group the matrix through hierarchical clustering methods, extract the high-frequency co-occurrence patterns and their time distribution characteristics, and form long-term usage habit features representing the user's context preferences; Construct a three-layer user portrait based on emotional preference features, scene preference features, physiological rhythm features, and long-term usage habit features, and establish an inter-layer connection relationship through feature indexing; Among them, the first layer stores the user's basic information, the second layer stores the long-term usage habit features and scene preference features, the third layer stores the emotional preference features and physiological rhythm features, and the three-layer structure establishes a two-way indexing relationship through feature IDs.
[0042] The interaction decision module is used to construct a personalized emotional interaction model according to the user portrait, the user's current emotional state, and the current scene type. The personalized emotional interaction model includes an emotional evolution unit, a scene interaction unit, and a physiological regulation unit; The emotional evolution unit analyzes the emotional change law based on the emotional preference features and the historical emotional state sequence, combines the current emotional state and the current scene type to predict the emotional change trend, and determines the target direction of emotional guidance; The scene interaction unit generates an interaction strategy including instruction priorities and combination rules based on the target direction and the current scene type, screens the multi-modal interaction instructions corresponding to the current scene from the scene-based emotional interaction instruction library according to the instruction priorities, and combines the screened instructions according to the combination rules, considering the long-term usage habit features and scene preference features, to form a basic instruction sequence; The physiological regulation unit adjusts the execution parameters of each instruction in the basic instruction sequence according to the physiological rhythm features to generate an execution instruction sequence.
[0043] Specifically, the workflow diagram of the emotional evolution unit is as Figure 3As shown, it includes obtaining the historical emotional state sequence and emotional preference characteristics in the emotional memory library, constructing a hierarchical emotional state map, which hierarchically organizes emotional states according to intensity and type; performing time series feature analysis on the historical emotional state sequence, and combining with the hierarchical emotional state map, extracting emotional periodic patterns and conversion critical points to form an emotional change feature set; among them, the emotional change feature set includes emotional duration period, emotional conversion rate, and emotional fluctuation intensity.
[0044] Furthermore, a dual-channel emotional prediction framework is established. The first channel predicts the state transition trend based on the emotional state transition probability matrix, and the second channel predicts the emotional change law based on the emotional change feature set. The output results of the two channels are weighted and fused to generate a basic emotional prediction vector. The specific steps are as follows: The historical emotional state sequence is divided into multiple subsequence units according to a preset time window, and a Markov prediction model is constructed using the emotional state transition probability matrix. Forward-backward reasoning is performed on each subsequence unit based on the Markov prediction model to calculate the state transition chain of each subsequence unit, forming the first prediction component; clustering the emotional periodic patterns in the emotional change feature set, constructing a support vector machine classifier based on the conversion critical point, and inputting the clustering result and the output result of the support vector machine classifier into a long short-term memory network to generate the second prediction component; using the attention mechanism to align the features of the first prediction component and the second prediction component, constructing a confidence evaluation unit based on child development psychology indicators, and calculating the weight coefficient according to the confidence evaluation unit; fusing the first prediction component and the second prediction component by weighted summation to generate a basic emotional prediction vector.
[0045] Furthermore, a scene modulation sub-unit is designed to convert the current scene type into a scene context vector, and modulate the basic emotional prediction vector through a cross-attention network to generate a context-adaptive emotional prediction result; among them, the modulation process is as follows: The scene type is used to obtain the scene description vector by querying the scene feature library, and the scene description vector is decomposed into scene element sub-vectors based on the hierarchical scene semantic decomposer. The scene element sub-vectors include environmental elements, time elements, and activity elements; the basic emotional prediction vector is copied to construct a prediction feature sequence, and the multi-head attention mechanism is used to calculate the correlation intensity matrix between the prediction feature sequence and the scene element sub-vectors; each feature vector in the prediction feature sequence is weighted based on the correlation intensity matrix, and the weighted result is input into a bidirectional gated recurrent unit to generate a modulation vector; a linear interpolation combination of the modulation vector and the basic emotional prediction vector is performed to obtain a context-adaptive emotional prediction result.
[0046] Furthermore, a steady-state evaluation sub-unit is designed to construct an emotional steady-state interval based on the user's current emotional state, calculate the deviation degree between the situation-adaptive emotional prediction result and the emotional steady-state interval, and generate an emotional deviation vector, which is specifically as follows: Input the user's current emotional state into the emotional steady-state model. The emotional steady-state model constructs an emotional fluctuation threshold based on the children's psychological development standard, combines the current emotional state to generate upper and lower boundary curves, and forms an emotional steady-state interval; Calculate the difference between the situation-adaptive emotional prediction result and the upper and lower boundary curves of the emotional steady-state interval, and construct a time-series deviation feature map based on the difference sequence; Use a sliding window to scan the time-series deviation feature map, extract local fluctuation features and global trend features, input the local fluctuation features and global trend features into a Gaussian mixture model, and generate a deviation metric vector; Based on the deviation metric vector, adaptively adjust the emotional steady-state interval and output an emotional deviation vector.
[0047] Furthermore, obtain multiple candidate emotional guidance directions from the preset doll companionship treatment knowledge base according to the emotional deviation vector, and select the optimal emotional guidance direction based on the children's emotional development theory; Compare the optimal emotional guidance direction with the user's current emotional state, calculate the emotional guidance gradient, and output an emotional guidance target sequence with time-series control parameters. The specific steps are as follows: Construct a multi-dimensional emotional space, map the optimal emotional guidance direction and the user's current emotional state to the multi-dimensional emotional space to form an emotional direction vector and the current emotional point; Calculate the Euclidean distance between the emotional direction vector and the current emotional point to obtain an emotional distance value; Expand the emotional direction vector along the time axis to form an emotional evolution path curve, and perform piecewise linearization on the emotional evolution path curve to generate a piecewise guidance curve; Calculate the slope and curvature of the piecewise guidance curve in each time period to form an emotional change rate matrix; Based on the emotional distance value and the emotional change rate matrix, use the gradient descent algorithm to calculate the optimal emotional transition path and generate an emotional guidance gradient; Calibrate the emotional guidance gradient with the user's emotional tolerance parameter to adjust the gradient size to ensure that the emotional guidance process is gentle and acceptable; Generate time-series control parameters according to the emotional guidance gradient, including transition duration parameters, intensity adjustment parameters, and rhythm control parameters, and associate them with the emotional guidance target to form an emotional guidance target sequence.
[0048] Through the design of the above-mentioned emotion evolution unit, the present invention realizes more accurate prediction of emotion change trends and more natural emotion guidance effects. Through the hierarchical emotion state map and the dual-channel prediction framework, the accuracy of emotion prediction is improved; based on the scene modulation cross-attention mechanism, the scene adaptability of emotion guidance is enhanced; through the steady-state evaluation and progressive emotion guidance scheme under the standards of child psychology, the safety and acceptance of emotion regulation are improved; by adopting a modular design and standardized interfaces, the system has good scalability and maintainability. This emotion evolution unit is particularly suitable for the application scenario of emotion interactive companion dolls, and can provide children with more intelligent, natural and safe emotion companionship services, fully meeting the needs of continuous, stable and personalized emotion guidance during long-term companionship.
[0049] Specifically, the workflow diagram of the scene interaction unit is as Figure 4 shown, including converting the timing control parameters in the emotion guidance target sequence into an interaction timing matrix, constructing a scene constraint model based on the interaction timing matrix and the current scene type, and outputting an interaction constraint vector, where the interaction constraint vector includes time window parameters and scene limitation conditions; using the interaction constraint vector to segment the emotion guidance target sequence, constructing a sub-goal mapping network for each time window, dynamically associating the emotion guidance target with the scene limitation conditions, and generating a phased interaction strategy. The specific steps are as follows: dividing the emotion guidance target sequence into multiple time windows, calculating the constraint density of each time window based on the interaction constraint vector, and generating a window constraint feature map; constructing a two-layer sub-goal mapping network, where the two-layer sub-goal mapping network includes an emotion mapping layer and a scene adaptation layer; the emotion mapping layer receives the emotion guidance target and the emotion change gradient within each time window, and generates an emotion target representation vector through non-linear transformation; the scene adaptation layer receives the window constraint feature map and the current scene type parameters, extracts the scene limitation factors, and outputs a scene constraint representation vector through an adaptive gating unit; designing an emotion-scene fusion matrix to perform tensor fusion on the emotion target representation vector and the scene constraint representation vector to form a joint feature space; establishing a dynamic weight allocation mechanism in the joint feature space, setting different weight coefficients for different scene types, and generating a time-sequential interaction intention vector through weighted fusion; using the time-sequential interaction intention vector to drive the interaction strategy generation network, and initializing the interaction strategy generation network with the parameters of the child cognitive development model, and adjusting the strategy generation parameters according to the cognitive characteristics of children of different ages; dividing the output result of the interaction strategy generation network into three types of strategy templates, namely, soothing type, companionship type, and guidance type, according to the doll interaction mode, selecting the most matching strategy template according to the current emotion guidance target, filling in the strategy parameters, and generating a phased interaction strategy.
[0050] Furthermore, perform temporal combination and hierarchical analysis on the phased interaction strategies in sequence, extract interaction intention features and temporal constraint relationships, construct a priority mapping matrix and a modality collaboration graph, and generate an instruction screening strategy. The instruction screening strategy includes an instruction priority sequence and a modality combination rule. The specific steps are as follows: Combine the phased interaction strategies of multiple time windows through a temporal fusion algorithm to form a complete interaction strategy chain; perform hierarchical decomposition on the interaction strategy chain, divide the strategy content into a core intention layer, an interaction behavior layer, and a modality expression layer, extract the feature vectors of each layer, and generate a three-layer feature map; extract interaction intention features based on the three-layer feature map, including emotion regulation intention, companionship interaction intention, and cognitive guidance intention, and calculate the intention intensity value; analyze the temporal constraint relationships in the interaction strategy chain, identify and quantify the conditions for preferential execution, mutually exclusive execution, and synchronous execution; establish an intention-instruction mapping relationship library, map the extracted interaction intention features to the instruction type space, and generate a priority mapping matrix based on the intention intensity value; analyze the temporal constraint relationships in the interaction strategy chain, identify and quantify the conditions for preferential execution, mutually exclusive execution, and synchronous execution; establish an intention-instruction mapping relationship library, map the extracted interaction intention features to the instruction type space, and generate a priority mapping matrix based on the intention intensity value; combine the temporal control parameters in the emotion guidance target sequence to construct a voice-expression-action three-modal association network, and generate a modality collaboration graph; according to the priority mapping matrix and the modality collaboration graph, calculate the execution priorities and modality combination parameters of various instructions to form an instruction screening strategy. The instruction screening strategy includes an instruction priority sequence and a modality combination rule.
[0051] Furthermore, retrieve the candidate instruction set from the scenario-based emotion interaction instruction library according to the instruction priority sequence, fuse the long-term usage habit features and scenario preference features into an interaction evaluation vector, and use the interaction evaluation vector to evaluate and screen the candidate instruction set to obtain the optimal instruction set; construct an instruction execution dependency graph based on the modality combination rule. The instruction execution dependency graph is used to determine the parallel execution units and sequential execution units of the instructions, and label the time window constraints of each execution unit; reorganize the optimal instruction set according to the instruction execution dependency graph, and control the temporal relationship of instruction execution through a recursive gated network. The recursive gated network is used to ensure the coordinated cooperation of multi-modal instructions.
[0052] Through the design of the above-mentioned scenario interaction unit, the present invention realizes the precise mapping and efficient execution from the emotional guidance goal to the specific interaction instructions. Based on the double-layer sub-goal mapping network, the adaptability of the interaction strategy to the scenario constraints is improved; through the emotion-scenario fusion matrix and the dynamic weight allocation mechanism, the precise expression of the interaction intention is realized; the instruction screening strategy constructed by the hierarchical analysis and the time-series combination method ensures the orderly execution of the multi-modal interaction instructions; the interaction evaluation mechanism that combines the usage habit characteristics and the scenario preference characteristics enhances the personalization degree of the interaction process; the instruction execution control based on the recursive gated network ensures the coordination and consistency of the multi-modal interactions such as speech, expression, and action. This scenario interaction unit is particularly suitable for the precise interaction requirements of the emotion interactive companion doll in complex scenarios, and can transform the abstract emotional guidance goal into a specific and executable multi-modal interaction instruction sequence, improving the execution efficiency of emotional guidance while ensuring the naturalness of the interaction.
[0053] Furthermore, the working process of the physiological regulation unit is as follows: extract the cycle feature parameters from the physiological rhythm characteristics, and divide the basic instruction sequence into corresponding execution cycle units in combination with the cycle feature parameters; generate an intensity regulation matrix based on the physiological rhythm characteristics, and perform intensity mapping on the execution parameters of the voice module, expression module, and action module in each execution cycle unit; adjust the execution parameters in the execution cycle unit according to the intensity regulation matrix to obtain an adaptive instruction group; reorganize the adaptive instruction group according to the execution cycle unit to generate an execution instruction sequence.
[0054] The execution control module is used to control the voice module, expression module, and action module of the doll to complete the emotional interaction behavior based on the execution instruction sequence.
[0055] The database module, the database module includes an emotional memory library and a scenario-based emotional interaction instruction library, where the emotional memory library stores the historical emotional feature vector sequence, emotional state sequence, and scenario type sequence, and the scenario-based emotional interaction instruction library stores the voice instruction set, expression instruction set, and action instruction set for different scenarios.
[0056] Furthermore, this embodiment also provides a scenario-based emotional interactive companion doll method, which includes collecting an interaction data set and a scenario data set through the multi-modal sensing system of the doll to generate an extended data matrix; extracting features from the extended data matrix, using a multi-modal feature network to generate an emotional feature vector and a scenario feature vector, determining the current emotional state of the user based on the emotional feature vector, and identifying the current scenario type based on the scenario feature vector; extracting the long-term usage habit features, emotional preference features, physiological rhythm features, and scenario preference features of the user based on the user's basic information, the historical emotional feature vector sequence, the historical emotional state sequence, and the historical scenario type sequence stored in the emotional memory library to construct a user profile; constructing an interaction decision module according to the user profile, the user's current emotional state, and the current scenario type, where the interaction decision module includes an emotional evolution unit, a scenario interaction unit, and a physiological regulation unit; the emotional evolution unit analyzes the emotional change law based on the emotional preference features and the historical emotional state sequence, and combines the current emotional state and the current scenario type to predict the emotional change trend and determine the emotional guidance target sequence; the scenario interaction unit generates an instruction screening strategy including an instruction priority sequence and a modality combination rule based on the emotional guidance target sequence and the current scenario type, screens the multi-modal interaction instructions corresponding to the current scenario from the scenario-based emotional interaction instruction library according to the priority sequence, and combines the screened instructions according to the modality combination rule to form a basic instruction sequence; the physiological regulation unit adjusts the execution parameters of each instruction in the basic instruction sequence according to the physiological rhythm features to generate an execution instruction sequence; controlling the voice module, the expression module, and the action module of the doll to complete the emotional interaction behavior based on the execution instruction sequence, and storing the emotional feature vector, the emotional state, and the scenario type in the interaction scenario in the emotional memory library.
[0057] In summary, the present invention realizes the personalized and accurate response of the emotional interaction doll. Based on the comprehensive analysis of multi-modal sensing and scenario perception, it can accurately identify the emotional state and scenario needs of the user. Through the design of a dual-channel emotional prediction framework and a scenario modulation mechanism, the system can more accurately predict the emotional change trend and dynamically adjust the interaction strategy according to the scenario features. Based on the hierarchical emotional state map and the progressive emotional guidance scheme, the system can provide a natural and smooth interaction experience while ensuring the safety of emotional regulation. Through the modular multi-modal interaction design and the collaborative control of the recursive gated network, the system realizes the precise coordination of various interaction methods such as voice, expression, and action, improving the execution efficiency of emotional guidance.
[0058] Example 2, referring to Figures 1 to 4 , which is the second embodiment of the present invention. This embodiment provides a scenario-based emotional interactive companion doll system. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0059] To verify the effectiveness of the technical solution of the present invention in actual applications, an interactive treatment scenario for children aged 3 to 6 years in a certain children's rehabilitation center was selected for experimental verification. The experimental subjects were 40 children (20 boys and 20 girls), and the experimental duration was 12 weeks. The experimental environment included three typical scenarios: a rehabilitation training room, a children's activity room, and a personal rest area. Each scenario was equipped with standard lighting equipment (400 - 600 lux) and a constant temperature air conditioner (24 ± 1 °C) to ensure environmental consistency.
[0060] In the data collection session, a multi-modal sensing system was used to monitor the children. Facial image collection was performed using an in-built high-definition camera (30 fps, 1080P); heart rate data was collected through a wearable photoplethysmogram sensor (sampling rate 100 Hz); pressure data was obtained using a distributed thin-film pressure sensing array (16×16 dot matrix, response time < 50 ms); voice data was recorded via an omnidirectional microphone array (sampling rate 48 kHz). For scene data collection, environmental parameters including temperature, humidity, light, noise, etc. were monitored in real time through an environmental sensor network; time features were recorded based on the system clock; location features were tracked through a UWB positioning system; and peripheral item features were obtained using an RFID tag identification system.
[0061] During the experiment, the researchers first configured the parameters of the doll system. The emotion type classifier adopted an improved ResNet architecture and was pre-trained using a public emotion dataset; the emotion state determination threshold was set to 0.75; the scene matching threshold was set to 0.8. The emotion intensity quantization standard was set with different baseline values according to the age groups of children: 3 - 4-year-old group (baseline value 0.6), 4 - 5-year-old group (baseline value 0.7), 5 - 6-year-old group (baseline value 0.8). The time window parameters were set as follows: the basic analysis window was 5 seconds, the emotion state evaluation window was 30 seconds, and the scene persistence judgment window was 3 minutes. In terms of interaction strategies, the basic weight ratios of the three types of strategies, namely the soothing type, the accompanying type, and the guiding type, were set to 3:4:3 and dynamically adjusted according to the scene type.
[0062] Through 12 weeks of experimental operation, the system accumulated approximately 3600 hours of interaction data. Data analysis showed that: the accuracy rate of emotion state recognition gradually increased from the initial 76.5% to 94.2%; the scene adaptability score increased from the initial 72.3 points to 91.8 points (full score 100 points); in the children's acceptance evaluation, 90.5% of the test subjects showed a willingness for positive interaction, and the average single interaction duration extended from the initial 8.3 minutes to 23.7 minutes. The evaluation of the emotion regulation effect showed that the average relief time of negative emotions was shortened by 46.3%, and the duration of positive emotions was extended by 52.8%. Heart rate variability analysis indicated that, accompanied by the doll, the children's stress index decreased by an average of 37.2%, and this effect was particularly obvious in the rehabilitation training scenario.
[0063] Based on the above experimental data, the research team conducted a comparative analysis of the present invention and the prior art, as shown in Table 1.
[0064] Table 1 Performance Comparison between the Present Invention and the Prior Art
[0065] Evaluation metrics Traditional doll method Simple emotional interaction system System of the present invention Accuracy of emotion recognition 65.3% 82.7% 94.2% Scene adaptability score No scene awareness 75.6 91.8 Average interaction duration 5.8 min 12.4 min 23.7h Negative emotion alleviation efficiency Baseline value Improved by 23.5% Improved by 46.3% Degree of improvement in stress index Baseline value Decreased by 18.9% Decreased by 37.2% Personalization level None Low High Multimodal synergy Simple feedback Partial synergy Full synergy Number of scenes covered 1 - 2 types 3 - 4 types More than 8 types Continuous learning ability None Weak Strong
[0066] From the comparison data in Table 1, it can be seen that the present invention has achieved significant improvements in key indicators such as emotion recognition accuracy, scene adaptability, and interaction persistence. Especially in terms of personalization and continuous learning ability, through the innovative multi-modal feature network and portrait construction module, the accurate grasp and dynamic adjustment of children's emotional needs are realized. The experimental results show that the present invention not only outperforms the prior art in technical indicators but also demonstrates obvious advantages in the actual application effect, providing a new technical solution for the field of children's emotional companionship.
[0067] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A scene-based emotional interactive companion doll system, characterized in that: include, A data collection module is used to collect interaction data groups and scene data groups to generate an extended data matrix; The feature analysis module is used to extract features from the extended data matrix, generate emotion feature vectors and scene feature vectors using a multimodal feature network, determine the user's current emotion state, and identify the current scene type; The portrait construction module is used to extract the user's long-term usage habit characteristics, emotional preference characteristics, physiological rhythm characteristics, and scene preference characteristics based on the emotional memory library to construct a user portrait; Interactive decision-making module, including emotion evolution unit, scene interaction unit and physiological regulation unit; The emotion evolution unit predicts the emotion change trend based on the emotion preference characteristics and the historical emotion state sequence, and determines the emotion guidance target sequence; the scene interaction unit generates an instruction screening strategy including an instruction priority sequence and a modal combination rule based on the emotion guidance target sequence and the current scene type, screens the instructions corresponding to the current scene according to the instruction priority sequence, and combines the screened instructions into a basic instruction sequence according to the modal combination rule; The physiological regulation unit adjusts the basic instruction sequence according to the physiological rhythm characteristics to generate an execution instruction sequence.
2. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: Also includes, An execution control module, used to control the voice module, expression module, and action module of the doll to complete emotional interaction behaviors based on the execution instruction sequence; A database module, the database module includes an emotion memory library and a scenario-based emotion interaction instruction library, wherein the emotion memory library stores a historical emotion feature vector sequence, an emotion state sequence, and a scenario type sequence, and the scenario-based emotion interaction instruction library stores a voice instruction set, an expression instruction set, and an action instruction set for different scenarios; The interaction data group includes facial images, heart rate data, pressure data, and voice data; the scene data group includes environmental parameters, time features, location features, and surrounding object features.
3. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: The method of generating the emotion feature vector and the scene feature vector by using the multimodal feature network comprises the following steps: Preprocessing the extended data matrix to obtain a standardized data matrix, and dividing the standardized data matrix into an interaction sub-matrix and a scene sub-matrix; Processing the interaction sub-matrix and the scene sub-matrix through a dual-branch feature extraction network, wherein the dual-branch feature extraction network includes an interaction feature branch and a scene feature branch; wherein the interaction feature branch includes an emotion modality adaptation layer and an emotion fusion network, and the scene feature branch includes a scene encoder and a scene semantic understanding network; The emotion modality adaptation layer in the interaction feature branch is used to extract and transform the interaction sub-matrix, and the emotion fusion network is used to integrate the transformed features to output the emotion representation features; Using the scene encoder in the scene feature branch to perform feature encoding on the scene sub-matrix, and combining and processing through the scene semantic understanding network to output scene semantic representation features; Establishing a spatiotemporal alignment module to connect the interactive feature branch and the scene feature branch, wherein the spatiotemporal alignment module adopts an adaptive time window and an emotion delay compensation mechanism to input the emotion representation features and the scene semantic representation features into the emotion mapping layer for feature mapping; The mapped emotion representation features and scene semantic representation features are respectively input into an emotion feature generator and a scene feature generator, and the emotion feature generator and the scene feature generator use a multi-head attention network to perform deep feature processing to generate an emotion feature vector and a scene feature vector.
4. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: Determining the user's current emotional state and identifying the current scene type comprises the following steps: Establishing an emotion state recognition model, the emotion state recognition model includes an emotion type classifier and an emotion intensity calculation unit, and inputting the emotion feature vector into the emotion state recognition model; The emotion type classifier performs classification operation on the emotion feature vector to obtain an emotion type probability vector; Based on the comparison result of the maximum probability value in the emotion type probability vector and the emotion judgment threshold, the basic emotion type is determined in combination with the historical emotion state sequence; Inputting the emotion feature vector into the emotion intensity calculation unit, and calculating the emotion intensity value based on a preset emotion intensity quantification standard; Inputting the basic emotion type and the emotion intensity value into an emotion state mapping network to generate an emotion state characteristic curve, and determining the user's current emotion state based on the peak position and waveform characteristics of the emotion state characteristic curve; Establishing a scene recognition model, inputting the scene feature vector into the scene recognition model, and calculating the matching degree between the scene feature vector and each scene type template based on a built-in scene type template library; Based on the comparison result of the maximum matching degree and the scene matching threshold, the current scene type is determined in combination with the historical scene data.
5. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: The workflow of the portrait construction module is as follows: Read the historical emotion feature vector sequence, historical emotion state sequence and historical scene type sequence in the emotion memory library, and concatenate the three sequences with the user basic information to construct the original feature matrix; Performing time series decomposition on the historical emotion feature vector sequence in the original feature matrix, calculating the emotion state transition probability matrix, and obtaining the emotion preference feature; Mapping the historical scene type sequence to a scene association matrix, calculating the scene transition probability and scene duration distribution, and obtaining scene preference features; Performing frequency domain transformation on the historical emotion feature vector sequence, combining the physiological parameters in the user basic information, calculating the period feature parameters, and obtaining the physiological rhythm feature; Aligning the historical emotion state sequence with the historical scene type sequence in time, calculating the joint distribution matrix, and obtaining long-term usage habit features; A three-layer user portrait is constructed based on the emotion preference feature, the scene preference feature, the physiological rhythm feature, and the long-term usage habit feature, and an inter-layer connection relationship is established through feature indexing.
6. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: The workflow of the emotion evolution unit is as follows: Obtain the historical emotional state sequence and emotional preference characteristics in the emotional memory library and construct a hierarchical emotional state map; Performing time series feature analysis on the historical emotional state sequence, combining the hierarchical emotional state map, extracting emotional periodic patterns and transition critical points, and forming an emotional change feature set; A dual-channel emotion prediction framework is established, in which the first channel predicts the state transition trend based on the emotion state transition probability matrix, and the second channel predicts the emotion change law based on the emotion change feature set. The output results of the two channels are weighted and fused to generate the basic emotion prediction vector. Design a scene modulation subunit to convert the current scene type into a scene context vector, modulate the basic emotion prediction vector through a cross-attention network, and generate a context-adaptive emotion prediction result; Design a steady-state evaluation subunit, construct an emotional steady-state interval based on the user's current emotional state, calculate the degree of deviation between the situation-adaptive emotional prediction result and the emotional steady-state interval, and generate an emotional deviation vector; Obtaining at least two candidate emotion guidance directions from a preset doll companion therapy knowledge base according to the emotion deviation vector, and selecting the optimal emotion guidance direction based on the child emotion development theory; The optimal emotion guidance direction is compared with the user's current emotion state, the emotion guidance gradient is calculated, and an emotion guidance target sequence with timing control parameters is output.
7. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: The workflow of the scene interaction unit is as follows: Converting the timing control parameters in the emotion guidance target sequence into an interaction timing matrix, constructing a scene constraint model based on the interaction timing matrix and the current scene type, and outputting an interaction constraint vector; The emotion guidance target sequence is segmented by using the interaction constraint vector, a sub-target mapping network is constructed for each time window, the emotion guidance target is dynamically associated with the scene constraint condition, and a phased interaction strategy is generated; The phased interaction strategies are sequentially combined and hierarchically analyzed to extract interaction intention features and timing constraint relationships, build a priority mapping matrix and a modal collaboration diagram, and generate an instruction screening strategy, which includes an instruction priority sequence and a modal combination rule; According to the instruction priority sequence, the scenario-based emotional interaction instruction library is retrieved to obtain a candidate instruction set, long-term usage habit features and scenario preference features are integrated into an interaction evaluation vector, and the candidate instruction set is evaluated and screened using the interaction evaluation vector to obtain an optimal instruction set; Constructing an instruction execution dependency graph based on the modal combination rule; The optimal instruction set is reorganized according to the instruction execution dependency graph, and the timing relationship of instruction execution is controlled through a recursive gating network to generate a basic instruction sequence.
8. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: The workflow of the physiological regulation unit is as follows: Extracting period characteristic parameters from the physiological rhythm characteristics, and dividing the basic instruction sequence into corresponding execution period units in combination with the period characteristic parameters; Generate an intensity adjustment matrix based on the physiological rhythm characteristics, and perform intensity mapping on the execution parameters of the voice module, the expression module, and the action module in each execution cycle unit; adjusting the execution parameters in the execution cycle unit according to the strength adjustment matrix to obtain an adaptive instruction group; The adaptive instruction group is reorganized according to the execution cycle unit to generate an execution instruction sequence.
9. A method for using a scene-based emotional interactive companion doll system, based on the scene-based emotional interactive companion doll system according to any one of claims 1 to 8, characterized in that: Also includes, The interactive data set and the scene data set are collected through the multimodal sensing system of the doll to generate an extended data matrix; Performing feature extraction on the extended data matrix, generating an emotion feature vector and a scene feature vector using a multimodal feature network, determining a current emotion state of the user based on the emotion feature vector, and identifying a current scene type based on the scene feature vector; Based on the user's basic information, the historical emotional feature vector sequence, historical emotional state sequence, and historical scene type sequence stored in the emotional memory library, the user's long-term usage habit characteristics, emotional preference characteristics, physiological rhythm characteristics, and scene preference characteristics are extracted to build a user portrait; Constructing an interactive decision module according to the user portrait, the user's current emotional state, and the current scene type, wherein the interactive decision module includes an emotional evolution unit, a scene interaction unit, and a physiological regulation unit; The emotion evolution unit analyzes the emotion change rules based on the emotion preference characteristics and the historical emotion state sequence, and predicts the emotion change trend in combination with the current emotion state and the current scene type, and determines the emotion guidance target sequence; The scene interaction unit generates an instruction screening strategy including an instruction priority sequence and a modal combination rule based on the emotion guidance target sequence and the current scene type, screens the multimodal interaction instructions corresponding to the current scene from the scene-based emotion interaction instruction library according to the priority sequence, and combines the screened instructions into a basic instruction sequence according to the modal combination rule; The physiological regulation unit adjusts the execution parameters of each instruction in the basic instruction sequence according to the physiological rhythm characteristics to generate an execution instruction sequence; Based on the execution instruction sequence, the speech module, expression module and action module of the doll are controlled to complete the emotional interaction behavior, and the emotional feature vector, emotional state and scene type in the interactive scene are stored in the emotional memory library.
Citation Information
Patent Citations
Information processing method and mobile terminal
CN105681900A
Method and system for processing multimedia scene interaction data
CN118260380A
Mental health state assessment method based on big data
CN119324060A
Real-time emotion recognition and regulation method and device combined with brain-computer interface
CN119357767A
Cited By
Intelligent pet robot system and control method thereof
CN120516705A
Multi-source data fused building settlement trend prediction method, equipment and medium
CN120744721A
Robot dynamic control method and system combined with multi-mode emotion interaction
CN121061899A
Robot dynamic control method and system combined with multi-modal emotional interaction
CN121061899B
Expression recognition and interaction method based on edge vision and multi-modal model
CN121281118A