A scene-based emotional interactive companion doll system and method
Through the comprehensive analysis of multimodal sensing and scene perception, a personalized emotional interaction model is constructed, which solves the problem of insufficient emotional interaction doll recognition in the emotional state and scene, and realizes accurate emotional companionship and natural interaction experience.
Patent Information
- Application Number
- CN202510276780.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Existing emotional interaction dolls find it difficult to fully perceive the user's emotional state and usage scenarios, the interaction strategy lacks personalized adjustment, and lacks in multimodal data processing and physiological rhythm analysis, resulting in a lack of targetedness and comfort in the interaction process.
The multimodal sensing system is used to collect data, and the emotional feature vector and scene feature vector are generated through the multimodal feature network. Combining user portraits and physiological rhythm characteristics, a personalized emotional interaction model is built, including emotional evolution, scene interaction and physiological regulation units to achieve accurate emotional companionship functions.
It realizes personalized and accurate response of emotional interaction dolls, can accurately identify users' emotional state and scene needs, dynamically adjust interaction strategies, provide a natural and smooth interactive experience, and improve the execution efficiency of emotional guidance.
Smart Images

Figure CN120179069B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a scene-based emotional interactive companion doll system and method. Background Art
[0002] Currently, most smart dolls on the market offer basic voice interaction and motion control capabilities, but they still have some shortcomings in terms of emotional interaction. Existing emotionally interactive dolls typically use a single sensor to collect user data, making it difficult to fully perceive the user's emotional state and usage scenarios, resulting in a lack of targeted and adaptable interaction. Furthermore, these dolls' interaction strategies are often pre-set, fixed patterns that cannot be dynamically adjusted based on the user's personality traits and usage habits, making it difficult to achieve sustained and effective emotional companionship.
[0003] Furthermore, existing emotionally interactive dolls primarily use simple feature extraction and fusion methods to process multimodal data, resulting in a relatively crude analysis of emotional and scene characteristics. In terms of interactive decision-making, they lack in-depth analysis and prediction of user emotional dynamics, making it difficult to achieve targeted emotional guidance. Furthermore, existing systems rarely consider the impact of users' physiological rhythms on the interaction process, compromising the naturalness and comfort of the interactive experience. Summary of the Invention
[0004] In view of the problems existing in the existing emotional interactive dolls, the present invention is proposed.
[0005] Therefore, the problem to be solved by the present invention is how to accurately identify the user's emotional state and scene type based on multimodal sensing data, and realize intelligent emotional companionship functions including emotional evolution, scene interaction and physiological regulation by constructing a personalized emotional interaction model.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In the first aspect, an embodiment of the present invention provides a scene-based emotional interactive companion doll system, which includes a data acquisition module for collecting interaction data groups and scene data groups to generate an extended data matrix; a feature analysis module for extracting features from the extended data matrix, generating emotion feature vectors and scene feature vectors using a multimodal feature network, determining the user's current emotional state based on the emotion feature vector, and identifying the current scene type based on the scene feature vector; a portrait construction module for extracting the user's long-term usage habit features, emotion preference features, physiological rhythm features, and scene preference features based on the emotion memory library to construct a user portrait; an interactive decision module including an emotion evolution unit, a scene interaction unit, and a physiological regulation unit; the emotion evolution unit analyzes the emotion change law based on the emotion preference features and the historical emotion state sequence, predicts the emotion change trend, and determines the emotion guidance target sequence; the scene interaction unit generates an instruction screening strategy including an instruction priority sequence and a modal combination rule based on the emotion guidance target sequence and the current scene type, screens the instructions corresponding to the current scene according to the instruction priority sequence, and combines the screened instructions to form a basic instruction sequence according to the modal combination rule; the physiological regulation unit adjusts the basic instruction sequence according to the physiological rhythm features to generate an execution instruction sequence.
[0008] As a preferred solution of the scene-based emotional interactive companion doll system described in the present invention, it also includes: an execution control module for controlling the doll's voice module, expression module, and action module to complete emotional interaction behavior based on the execution instruction sequence; a database module, the database module includes an emotional memory library and a scenario-based emotional interaction instruction library, wherein the emotional memory library stores historical emotional feature vector sequences, emotional state sequences, and scene type sequences, and the scenario-based emotional interaction instruction library stores voice instruction sets, expression instruction sets, and action instruction sets for different scenarios; the interaction data group includes facial images, heart rate data, pressure data, and voice data, and the scene data group includes environmental parameters, time features, location features, and surrounding object features.
[0009] As a preferred solution of the scene-based emotional interactive companion doll system described in the present invention, the method comprises the following steps: generating emotional feature vectors and scene feature vectors using a multimodal feature network: preprocessing the extended data matrix to obtain a standardized data matrix, dividing the standardized data matrix into an interaction sub-matrix and a scene sub-matrix; processing the interaction sub-matrix and the scene sub-matrix through a dual-branch feature extraction network, the dual-branch feature extraction network including an interaction feature branch and a scene feature branch; wherein the interaction feature branch includes an emotional modality adaptation layer and an emotional fusion network, and the scene feature branch includes a scene encoder and a scene semantic understanding network; utilizing the emotional modality adaptation layer in the interaction feature branch to extract and transform the interaction sub-matrix, and using the emotional modality adaptation layer to extract and transform the interaction sub-matrix, and utilizing ... generate the emotional modality. The emotion fusion network integrates the converted features and outputs the emotion representation features; the scene encoder in the scene feature branch is used to encode the scene sub-matrix, and the scene semantic understanding network is used for combination processing to output the scene semantic representation features; a spatiotemporal alignment module is established to connect the interaction feature branch and the scene feature branch. The spatiotemporal alignment module adopts an adaptive time window and emotion delay compensation mechanism to input the emotion representation features and the scene semantic representation features into the emotion mapping layer for feature mapping; the mapped emotion representation features and scene semantic representation features are input into the emotion feature generator and the scene feature generator respectively. The emotion feature generator and the scene feature generator adopt a multi-head attention network for deep feature processing to generate emotion feature vectors and scene feature vectors.
[0010] As a preferred solution of the scene-based emotional interactive companion doll system of the present invention, the following steps are included: determining the user's current emotional state and identifying the current scene type: establishing an emotional state recognition model, the emotional state recognition model includes an emotional type classifier and an emotional intensity calculation unit, and inputting the emotional feature vector into the emotional state recognition model; performing classification operations on the emotional feature vector through the emotional type classifier to obtain an emotional type probability vector; determining the basic emotional type based on the comparison result of the maximum probability value in the emotional type probability vector and the emotional judgment threshold in combination with the historical emotional state sequence; inputting the emotional feature vector into the emotional intensity calculation unit, and calculating the emotional intensity value based on a preset emotional intensity quantification standard; inputting the basic emotional type and the emotional intensity value into the emotional state mapping network to generate an emotional state characteristic curve, and determining the user's current emotional state based on the peak position and waveform characteristics of the emotional state characteristic curve; establishing a scene recognition model, inputting the scene feature vector into the scene recognition model, and calculating the matching degree between the scene feature vector and each scene type template based on the built-in scene type template library; determining the current scene type based on the comparison result of the maximum matching degree and the scene matching threshold in combination with historical scene data.
[0011] As a preferred solution of the scene-based emotional interactive companion doll system of the present invention, the workflow of the portrait construction module is as follows: read the historical emotional feature vector sequence, historical emotional state sequence and historical scene type sequence in the emotional memory library, and splice the three sequences with the user basic information to construct the original feature matrix; perform time series decomposition on the historical emotional feature vector sequence in the original feature matrix, calculate the emotional state transition probability matrix, and obtain the emotional preference feature; map the historical scene type sequence to the scene association matrix, calculate the scene transition probability and scene duration distribution, and obtain the scene preference feature; perform frequency domain transformation on the historical emotional feature vector sequence, combine the physiological parameters in the user basic information, calculate the periodic characteristic parameters, and obtain the physiological rhythm feature; align the historical emotional state sequence with the historical scene type sequence in time, calculate the joint distribution matrix, and obtain the long-term usage habit feature; construct a three-layer user portrait based on the emotional preference feature, scene preference feature, physiological rhythm feature, and long-term usage habit feature, and establish the inter-layer connection relationship through feature indexing.
[0012] As a preferred solution of the scene-based emotional interactive companion doll system of the present invention, the workflow of the emotional evolution unit is as follows: obtain the historical emotional state sequence and emotional preference characteristics in the emotional memory library, and construct a hierarchical emotional state map; perform time series feature analysis on the historical emotional state sequence, combine it with the hierarchical emotional state map, extract emotional periodic patterns and conversion critical points, and form an emotional change feature set; establish a dual-channel emotional prediction framework, in which the first channel predicts the state transition trend based on the emotional state transition probability matrix, and the second channel predicts the emotional change law based on the emotional change feature set, and the output results of the two channels are weighted fused to generate a basic emotional prediction vector; design the scene The modulation subunit converts the current scene type into a scene context vector, modulates the basic emotion prediction vector through the cross-attention network, and generates a context-adaptive emotion prediction result; designs a steady-state evaluation subunit, constructs an emotion steady-state interval based on the user's current emotion state, calculates the degree of deviation between the context-adaptive emotion prediction result and the emotion steady-state interval, and generates an emotion deviation vector; obtains at least two candidate emotion guidance directions from the preset doll companion therapy knowledge base according to the emotion deviation vector, and selects the optimal emotion guidance direction based on the child emotion development theory; compares the optimal emotion guidance direction with the user's current emotion state, calculates the emotion guidance gradient, and outputs an emotion guidance target sequence with timing control parameters.
[0013] As a preferred solution of the scene-based emotional interactive companion doll system of the present invention, the workflow of the scene interaction unit is as follows: converting the timing control parameters in the emotional guidance target sequence into an interaction timing matrix, constructing a scene constraint model based on the interaction timing matrix and the current scene type, and outputting an interaction constraint vector; using the interaction constraint vector to segment the emotional guidance target sequence, constructing a sub-goal mapping network for each time window, dynamically associating the emotional guidance target with the scene constraint conditions, and generating a phased interaction strategy; performing temporal combination and hierarchical analysis on the phased interaction strategy in turn, extracting interaction intention features and timing constraint relationships, constructing a priority mapping matrix and a modal collaboration graph, and generating an instruction screening strategy, which includes an instruction priority sequence and a modal combination rule; searching the scenario-based emotional interaction instruction library according to the instruction priority sequence to obtain a candidate instruction set, fusing long-term usage habit features and scene preference features into an interaction evaluation vector, and using the interaction evaluation vector to evaluate and screen the candidate instruction set to obtain the optimal instruction set; constructing an instruction execution dependency graph based on the modal combination rule; reorganizing the optimal instruction set according to the instruction execution dependency graph, controlling the timing relationship of instruction execution through a recursive gating network, and generating a basic instruction sequence.
[0014] As a preferred solution of the scene-based emotional interactive companion doll system of the present invention, the workflow of the physiological regulation unit is as follows: extracting periodic characteristic parameters from physiological rhythm characteristics, and dividing the basic instruction sequence into corresponding execution cycle units in combination with the periodic characteristic parameters; generating an intensity regulation matrix based on physiological rhythm characteristics, and performing intensity mapping on the execution parameters of the voice module, expression module, and action module in each execution cycle unit; adjusting the execution parameters in the execution cycle unit according to the intensity regulation matrix to obtain an adaptive instruction group; and reorganizing the adaptive instruction group according to the execution cycle unit to generate an execution instruction sequence.
[0015] In the second aspect, an embodiment of the present invention provides a scene-based emotional interactive companion doll method, which includes collecting interaction data groups and scene data groups through the doll's multimodal sensing system to generate an extended data matrix; performing feature extraction on the extended data matrix, generating emotion feature vectors and scene feature vectors using a multimodal feature network, determining the user's current emotion state based on the emotion feature vector, and identifying the current scene type based on the scene feature vector; extracting the user's long-term usage habit characteristics, emotion preference characteristics, physiological rhythm characteristics and scene preference characteristics based on the user's basic information, the historical emotion feature vector sequence, the historical emotion state sequence, and the historical scene type sequence stored in the emotion memory library, and constructing a user portrait; constructing an interactive decision module based on the user portrait, the user's current emotion state, and the current scene type, the interactive decision module including an emotion evolution unit, a scene interaction unit and a physiological regulation unit. The emotion evolution unit analyzes the emotion change rules based on the emotion preference characteristics and historical emotion state sequence, and predicts the emotion change trend in combination with the current emotion state and the current scene type, and determines the emotion guidance target sequence; the scene interaction unit generates an instruction screening strategy containing an instruction priority sequence and a modal combination rule based on the emotion guidance target sequence and the current scene type, and screens the multimodal interaction instructions corresponding to the current scene from the scene-based emotion interaction instruction library according to the priority sequence, and combines the screened instructions to form a basic instruction sequence according to the modal combination rule; the physiological regulation unit adjusts the execution parameters of each instruction in the basic instruction sequence according to the physiological rhythm characteristics, and generates an execution instruction sequence; based on the execution instruction sequence, the voice module, expression module, and action module of the puppet are controlled to complete the emotion interaction behavior, and the emotion feature vector, emotion state, and scene type in the interaction scene are stored in the emotion memory library.
[0016] The beneficial effects of the present invention are: the present invention realizes the personalized and precise response of the emotional interactive doll, and based on the comprehensive analysis of multimodal sensing and scene perception, it can accurately identify the user's emotional state and scene requirements. Through the design of a dual-channel emotion prediction framework and a scene modulation mechanism, the system can make more accurate predictions on the trend of emotional changes and dynamically adjust the interaction strategy according to the characteristics of the scene. Based on a hierarchical emotional state map and a progressive emotion guidance scheme, the system can provide a natural and smooth interactive experience while ensuring the safety of emotional regulation. Through modular multimodal interaction design and collaborative control of recursive gating networks, the system realizes the precise coordination of multiple interaction methods such as voice, expression, and action, and improves the execution efficiency of emotional guidance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is the module connection diagram of the scene-based emotional interactive companion doll system.
[0019] Figure 2 Flowchart of the multimodal feature network processing for a scene-based emotional interactive companion doll system.
[0020] Figure 3 This is the workflow diagram of the emotion evolution unit of the scene-based emotional interactive companion doll system.
[0021] Figure 4 This is the workflow diagram of the scene interaction unit of the scene-based emotional interactive companion doll system. DETAILED DESCRIPTION
[0022] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0025] Example 1, reference Figures 1 to 4 , which is the first embodiment of the present invention, provides a scene-based emotional interactive companion doll system, the module connection diagram is as follows Figure 1 Shown, including,
[0026] The data acquisition module is used to collect interaction data groups and scene data groups through the doll's multimodal sensing system, where the interaction data group includes facial images, heart rate data, pressure data, and voice data; the scene data group includes environmental parameters, time characteristics, location characteristics, and surrounding object characteristics to generate an extended data matrix.
[0027] It should be noted that the interaction and scenario data sets are preferred because: facial images capture the user's emotional micro-expressions and provide visual emotional cues; heart rate data reflects changes in physiological state and provides objective emotional indicators; pressure data records intimate interaction patterns such as hugging and caressing; and voice data contains emotional intonation information. Furthermore, environmental parameters influence emotional baseline levels, time features correlate with daily emotional cycles, location features distinguish between different scenarios such as home and school, and surrounding object features provide contextual interaction references. This multi-dimensional data collection method breaks through the limitations of traditional dolls' single trigger feedback. By comprehensively analyzing physiological, behavioral, and environmental factors, it builds a comprehensive emotional perception foundation, thereby realizing differentiated emotional companionship functions for different scenarios.
[0028] A feature analysis module is used to extract and fuse features from the expanded data matrix, generate emotion feature vectors and scene feature vectors using an improved multimodal feature network, determine the user's current emotion state based on the emotion feature vector, identify the current scene type based on the scene feature vector, and store the emotion feature vector, emotion state, and scene type in an emotion memory bank;
[0029] Specifically, the multimodal feature network processing flow chart is as follows Figure 2 As shown, the generation of emotion feature vectors and scene feature vectors using the improved multimodal feature network includes the following steps: normalizing and filtering the extended data matrix to obtain a standardized data matrix, dividing the standardized data matrix into an interaction sub-matrix and a scene sub-matrix based on the time window and the data type; processing the interaction sub-matrix and the scene sub-matrix through a dual-branch feature extraction network, the dual-branch feature extraction network includes an interaction feature branch and a scene feature branch; the interaction feature branch includes an emotion modality adaptation layer and an emotion fusion network, and the scene feature branch includes a scene encoder and a scene semantic understanding network; using the emotion modality adaptation layer in the interaction feature branch to extract and transform the interaction sub-matrix, and then using the emotion fusion network to weight the transformed features Integration, output emotion representation features; use the scene encoder in the scene feature branch to encode the scene sub-matrix, and combine and process it through the scene semantic understanding network to output the scene semantic representation features; establish a spatiotemporal alignment module to connect the interaction feature branch and the scene feature branch. The spatiotemporal alignment module adopts an adaptive time window and emotion delay compensation mechanism to input the emotion representation features and scene semantic representation features into the personalized emotion mapping layer for feature mapping; the mapped emotion representation features and scene semantic representation features are input into the emotion feature generator and scene feature generator respectively. The emotion feature generator and scene feature generator use a multi-head attention network for deep feature processing, and generate emotion feature vectors and scene feature vectors through attention weight allocation and feature selection mechanism.
[0030] The emotion modality adaptation layer includes facial image feature analysis units, heart rate feature analysis units, stress feature analysis units, and speech feature analysis units. Each unit processes corresponding features and converts them into a unified dimensional representation. The emotion fusion network uses an emotional context-aware attention mechanism to process the converted features, calculates emotion feature weight coefficients, and generates emotion representation features. Correspondingly, the scene encoder includes environmental parameter analysis units, temporal feature analysis units, spatial feature analysis units, and object feature analysis units. Each unit extracts and processes corresponding features. The scene semantic understanding network combines these features through a feature fusion mechanism to generate scene semantic representation features.
[0031] It is worth noting that the adaptive time window of the spatiotemporal alignment module dynamically adjusts the window parameters according to the state change rate; the emotional delay compensation mechanism predicts the current emotional trend by analyzing the historical emotional state sequence, thereby achieving temporal alignment of emotional representation features and scene semantic representation features.
[0032] It should be pointed out that the multi-head attention network in the emotion feature generator and scene feature generator contains multiple attention heads, and each attention head independently learns the importance weights of different feature dimensions; key features are highlighted through attention weight distribution, and the most representative feature combinations are screened through the feature selection mechanism. The final generated emotion feature vector contains the dimensions of emotion type, emotion intensity, emotion persistence and emotion variability, and the scene feature vector contains the dimensions of scene type, scene similarity, scene duration and scene change probability.
[0033] Secondly, determining the user's current emotional state based on the emotional feature vector and identifying the current scene type based on the scene feature vector include the following steps: establishing an emotional state recognition model, the emotional state recognition model includes an emotional type classifier and an emotional intensity calculation unit, and inputting the emotional feature vector into the emotional state recognition model; performing classification operations on the emotional feature vector through the emotional type classifier to obtain an emotional type probability vector; based on the comparison result of the maximum probability value in the emotional type probability vector and the emotional judgment threshold, combining the historical emotional state sequence to determine the basic emotional type, the specific steps are as follows: when the maximum probability value in the emotional type probability vector is greater than the emotional judgment threshold, the type corresponding to the maximum probability value in the emotional type probability vector is selected as the basic emotional type; when the maximum probability value in the emotional type probability vector is less than or equal to the emotional judgment threshold, the emotional state sequence within the preset time window is obtained from the emotional memory bank, and the emotional type with the highest frequency in the emotional state sequence is used as the basic emotional type.
[0034] Furthermore, the emotion feature vector is input into the emotion intensity calculation unit, and the emotion intensity value is calculated based on a preset emotion intensity quantification standard; wherein, the preset emotion intensity quantification standard refers to a quantification mechanism based on the feature weights and emotion scoring threshold system of different perception modalities, combined with the physiological signal fluctuation range and age-related standardized reference values for calculation.
[0035] Furthermore, the basic emotion type and emotion intensity value are input into the emotion state mapping network. The emotion state mapping network adopts a nonlinear conversion algorithm to comprehensively consider the emotion type weight coefficient and intensity influencing factor to generate an emotion state characteristic curve. The user's current emotion state is determined based on the peak position and waveform characteristics of the emotion state characteristic curve; a scene recognition model is established, the scene feature vector is input into the scene recognition model, and the matching degree between the scene feature vector and each scene type template is calculated based on the built-in scene type template library; based on the comparison result of the maximum matching degree and the scene matching threshold, the current scene type is determined in combination with the historical scene data. The specific steps are as follows: when the maximum matching degree is greater than the scene matching threshold, the type corresponding to the scene type template with the highest matching degree is selected as the current scene type; when the maximum matching degree is less than or equal to the scene matching threshold, the historical scene record with the highest similarity is retrieved from the emotion memory library based on the scene feature vector, and the scene type corresponding to the historical scene record is determined as the current scene type.
[0036] It should be noted that the maximum probability value in the emotion type probability vector is selected for comparison with the emotion judgment threshold because the maximum probability value in the probability vector output by the emotion type classifier represents the best match between the current emotion feature and the type. When the maximum probability value is lower than the threshold, it means that the matching degree between the current emotion feature and all emotion types is not ideal, and further judgment is needed in combination with the historical emotion state sequence. The maximum matching degree is selected for comparison with the scene matching threshold because the scene type corresponding to the maximum matching degree is the most likely current scene. When the maximum matching degree is not enough to meet the threshold requirement, it means that the current scene does not match the known template well enough, and it is necessary to retrieve historical scene data to assist in the judgment.
[0037] Preferably, the above technical solution can effectively address the problem of unstable emotion and scene recognition faced by emotionally interactive companion dolls in actual usage scenarios by introducing a comparison mechanism of the maximum probability value / matching degree and the threshold, combined with a method of auxiliary judgment using historical data. Since the user may be a child, their emotional expression and behavior patterns are often not standardized, and relying solely on real-time feature vectors may lead to large fluctuations in recognition results. This solution uses a dual guarantee mechanism of threshold judgment and historical data reference to ensure rapid response when the features are obvious, and can use historical data to provide stable and reliable recognition results when the features are not obvious, thereby enhancing the emotional interaction experience of the doll. At the same time, the solution can also gradually accumulate emotion and scene data, continuously enriching the historical database, so that the doll can better adapt to the emotional expression and usage habits of specific users, and realize personalized emotional companionship functions.
[0038] The portrait construction module is used to extract the user's long-term usage habit characteristics, emotional preference characteristics, physiological rhythm characteristics and scenario preference characteristics based on the user's basic information and emotional memory library to construct a user portrait.
[0039] Specifically, the workflow of the portrait construction module is as follows: read the historical emotional feature vector sequence, historical emotional state sequence and historical scene type sequence in the emotional memory library, and splice the three sequences with the user's basic information to construct the original feature matrix; perform time series decomposition on the historical emotional feature vector sequence in the original feature matrix, and calculate the emotional state transition probability matrix, where the emotional state transition probability matrix represents the migration law between emotional states, and extract the main feature vector by performing feature decomposition on the matrix, and combine it with the distribution of emotional state residence time to form an emotional preference feature reflecting the user's emotional preference pattern; map the historical scene type sequence to the scene association matrix, calculate the scene transition probability and scene duration distribution, where the scene transition probability and scene duration distribution represent the scene switching tendency and scene interaction persistence respectively, and integrate the two into a low-dimensional representation through a multi-dimensional scaling algorithm to form a scene preference feature that represents the user's scene selection pattern.
[0040] Furthermore, the historical emotional feature vector sequence is transformed in the frequency domain, and the periodic feature parameters are calculated in combination with the physiological parameters in the user's basic information to obtain the physiological rhythm features. The specific steps are as follows: the historical emotional feature vector sequence is segmented according to equal time intervals, and each segment of the sequence is subjected to fast Fourier transform to obtain the frequency domain feature matrix; the main frequency components and energy distribution ratios in the frequency domain feature matrix are extracted to construct the spectrum feature vector; physiological parameters (such as age, gender, average heart rate, average sleep duration, etc.) are extracted from the user's basic information to construct the physiological basic feature vector; the spectrum feature vector and the physiological basic feature vector are input into the adaptive weighted fusion network to calculate the circadian cycle coefficient, the emotional fluctuation cycle coefficient, and the energy change cycle coefficient; based on the above cycle coefficients, a multi-scale periodic feature parameter set including daily cycle features, weekly cycle features, and monthly cycle features is constructed; the multi-scale periodic feature parameter set is nonlinearly mapped to generate physiological rhythm features that characterize the user's physiological activity, emotional sensitivity, and interaction acceptance.
[0041] Furthermore, the historical emotional state sequence and the historical scene type sequence are aligned in time, and the joint distribution matrix is calculated, where the joint distribution matrix represents the co-occurrence frequency of emotional state and scene type. The matrix is grouped by a hierarchical clustering method, and high-frequency co-occurrence patterns and their time distribution characteristics are extracted to form long-term usage habit features that represent user situational preferences; a three-layer user portrait is constructed based on emotional preference features, scene preference features, physiological rhythm features, and long-term usage habit features, and the inter-layer connection relationship is established through feature indexing; among them, the first layer stores user basic information, the second layer stores long-term usage habit features and scene preference features, and the third layer stores emotional preference features and physiological rhythm features. The three-layer structure establishes a bidirectional index relationship through feature ID.
[0042] The interactive decision-making module is used to build a personalized emotional interaction model based on the user portrait, the user's current emotional state, and the current scene type. The personalized emotional interaction model includes an emotional evolution unit, a scene interaction unit, and a physiological regulation unit; the emotional evolution unit analyzes the law of emotional changes based on emotional preference characteristics and historical emotional state sequences, and predicts the emotional change trend in combination with the current emotional state and the current scene type to determine the target direction of emotional guidance; the scene interaction unit generates an interaction strategy containing instruction priority and combination rules based on the target direction and the current scene type, and filters the multimodal interaction instructions corresponding to the current scene from the scenario-based emotional interaction instruction library according to the instruction priority. According to the combination rules and taking into account the long-term usage habit characteristics and scene preference characteristics, the filtered instructions are combined to form a basic instruction sequence; the physiological regulation unit adjusts the execution parameters of each instruction in the basic instruction sequence according to the physiological rhythm characteristics to generate an execution instruction sequence.
[0043] Specifically, the workflow diagram of the emotion evolution unit is as follows: Figure 3As shown, it includes obtaining the historical emotional state sequence and emotional preference characteristics in the emotional memory library, constructing a hierarchical emotional state map, and organizing the emotional states in layers according to intensity and type; performing time series feature analysis on the historical emotional state sequence, combining the hierarchical emotional state map, extracting the emotional periodic pattern and conversion critical point, and forming an emotional change feature set; among which, the emotional change feature set includes the emotional duration period, emotional conversion rate and emotional fluctuation intensity.
[0044] Furthermore, a dual-channel emotion prediction framework is established, in which the first channel predicts the state transition trend based on the emotion state transition probability matrix, and the second channel predicts the emotion change law based on the emotion change feature set. The output results of the two channels are weighted and fused to generate a basic emotion prediction vector. The specific steps are as follows: the historical emotion state sequence is divided into multiple subsequence units according to the preset time window, the emotion state transition probability matrix is used to construct a Markov prediction model, forward-backward reasoning is performed on each subsequence unit based on the Markov prediction model, and the state transition chain of each subsequence unit is calculated to form a first prediction component; the emotion periodic patterns in the emotion change feature set are clustered, and a support vector machine classifier is constructed based on the transition critical point. The clustering results and the output results of the support vector machine classifier are input into the long short-term memory network to generate the second prediction component; the attention mechanism is used to align the features of the first prediction component and the second prediction component, a confidence assessment unit is constructed based on child development psychology indicators, and the weight coefficient is calculated according to the confidence assessment unit; the first prediction component and the second prediction component are fused by weighted summation to generate a basic emotion prediction vector.
[0045] Furthermore, a scene modulation sub-unit is designed to convert the current scene type into a scene context vector, and modulate the basic emotion prediction vector through a cross-attention network to generate a context-adaptive emotion prediction result; wherein, the modulation process is as follows: the scene type is obtained by querying the scene feature library to obtain a scene description vector, and the scene description vector is split into scene element sub-vectors based on a hierarchical scene semantic decomposer, and the scene element sub-vectors include environmental elements, time elements and activity elements; the basic emotion prediction vector is copied to construct a prediction feature sequence, and the multi-head attention mechanism is used to calculate the correlation strength matrix between the prediction feature sequence and the scene element sub-vectors; each feature vector in the prediction feature sequence is weighted based on the correlation strength matrix, and the weighted result is input into a bidirectional gated recurrent unit to generate a modulation vector; the modulation vector and the basic emotion prediction vector are linearly interpolated and combined to obtain a context-adaptive emotion prediction result.
[0046] Furthermore, a steady-state evaluation subunit is designed to construct an emotional steady-state interval based on the user's current emotional state, calculate the degree of deviation between the situational adaptability emotional prediction result and the emotional steady-state interval, and generate an emotional deviation vector, as follows: the user's current emotional state is input into the emotional steady-state model, the emotional steady-state model constructs an emotional fluctuation threshold based on the child psychology development standard, and generates upper and lower boundary curves in combination with the current emotional state to form an emotional steady-state interval; the difference between the situational adaptability emotional prediction result and the upper and lower boundary curves of the emotional steady-state interval is calculated, and a time series deviation feature map is constructed based on the difference sequence; the time series deviation feature map is scanned using a sliding window to extract local fluctuation features and global trend features, and the local fluctuation features and global trend features are input into a Gaussian mixture model to generate a deviation measurement vector; the emotional steady-state interval is adaptively adjusted based on the deviation measurement vector, and the emotional deviation vector is output.
[0047] Furthermore, according to the emotion deviation vector, multiple candidate emotion guidance directions are obtained from the preset doll companion therapy knowledge base, and the optimal emotion guidance direction is selected based on the child emotion development theory; the optimal emotion guidance direction is compared with the user's current emotion state, the emotion guidance gradient is calculated, and the emotion guidance target sequence with timing control parameters is output. The specific steps are as follows: construct a multidimensional emotion space, map the optimal emotion guidance direction and the user's current emotion state to the multidimensional emotion space, form an emotion direction vector and a current emotion point; calculate the Euclidean distance between the emotion direction vector and the current emotion point to obtain the emotion distance value; expand the emotion direction vector along the time axis to form Emotional evolution path curve, the emotional evolution path curve is piecewise linearized to generate a piecewise guidance curve; the slope and curvature of the piecewise guidance curve in each time period are calculated to form an emotional change rate matrix; based on the emotional distance value and the emotional change rate matrix, the gradient descent algorithm is used to calculate the optimal emotional transition path to generate an emotional guidance gradient; the emotional guidance gradient is calibrated with the user's emotional tolerance parameter, and the gradient size is adjusted to ensure that the emotional guidance process is smooth and acceptable; according to the emotional guidance gradient, timing control parameters are generated, including transition duration parameters, intensity adjustment parameters and rhythm control parameters, and associated with the emotional guidance target to form an emotional guidance target sequence.
[0048] Through the design of the above-mentioned emotion evolution unit, the present invention achieves more accurate predictions of emotion change trends and more natural emotion guidance effects. The accuracy of emotion prediction is improved through the hierarchical emotion state map and dual-channel prediction framework; the scene adaptability of emotion guidance is enhanced through the cross-attention mechanism based on scene modulation; the safety and acceptability of emotion regulation are improved through steady-state assessment and progressive emotion guidance scheme under child psychology standards; the use of modular design and standardized interfaces makes the system have good scalability and maintainability. The emotion evolution unit is particularly suitable for the application scenario of emotional interactive companion dolls, and can provide children with more intelligent, natural and safe emotional companionship services, fully meeting the needs of continuous, stable and personalized emotional guidance during long-term companionship.
[0049] Specifically, the workflow diagram of the scene interaction unit is as follows: Figure 4 As shown, it includes converting the timing control parameters in the emotion guidance target sequence into an interactive timing matrix, building a scene constraint model based on the interactive timing matrix and the current scene type, and outputting an interactive constraint vector, which includes time window parameters and scene constraints; using the interactive constraint vector to segment the emotion guidance target sequence, building a sub-target mapping network for each time window, dynamically associating the emotion guidance target with the scene constraints, and generating a phased interaction strategy. The specific steps are as follows: dividing the emotion guidance target sequence into multiple time windows, calculating the constraint density of each time window based on the interactive constraint vector, and generating a window constraint feature map; building a two-layer sub-target mapping network, which includes an emotion mapping layer and a scene adaptation layer; the emotion mapping layer receives the emotion guidance target and the emotion change gradient in each time window, and generates an emotion target representation vector through nonlinear transformation; the scene adaptation layer The response layer receives the window constraint feature map and the current scene type parameters, extracts the scene restriction factors, and outputs the scene constraint representation vector through the adaptive gating unit; designs the emotion-scene fusion matrix, and performs tensor fusion of the emotion target representation vector and the scene constraint representation vector to form a joint feature space; establishes a dynamic weight allocation mechanism in the joint feature space, sets differentiated weight coefficients for different scene types, and generates a time-series interaction intention vector through weighted fusion; uses the time-series interaction intention vector to drive the interaction strategy generation network, which is initialized with the parameters of the children's cognitive development model and adjusts the strategy generation parameters according to the cognitive characteristics of children of different age groups; divides the output results of the interaction strategy generation network into three types of strategy templates: soothing, companionship, and guidance according to the puppet interaction mode, selects the most matching strategy template according to the current emotion guidance goal, fills in the strategy parameters, and generates a phased interaction strategy.
[0050] Furthermore, the phased interaction strategies are sequentially combined and hierarchically analyzed to extract interaction intention features and timing constraint relationships, build a priority mapping matrix and a modal collaboration diagram, and generate an instruction screening strategy. The instruction screening strategy includes instruction priority sequences and modal combination rules. The specific steps are as follows: the phased interaction strategies of multiple time windows are combined through a timing fusion algorithm to form a complete interaction strategy chain; the interaction strategy chain is hierarchically decomposed to divide the strategy content into a core intention layer, an interaction behavior layer, and a modal expression layer, and the feature vectors of each layer are extracted to generate a three-layer feature map; based on the three-layer feature map, interaction intention features are extracted, including emotion regulation intention, companion interaction intention, and cognitive guidance intention, and the intention strength value is calculated; the timing constraint relationship in the interaction strategy chain is analyzed to identify and quantify priority execution conditions, mutual exclusion, and other conditions. Execution conditions and synchronous execution conditions; establish an intention-instruction mapping relationship library, map the extracted interaction intention features to the instruction type space, and generate a priority mapping matrix based on the intention strength value; analyze the timing constraint relationship in the interaction strategy chain, identify and quantify priority execution conditions, mutually exclusive execution conditions and synchronous execution conditions; establish an intention-instruction mapping relationship library, map the extracted interaction intention features to the instruction type space, and generate a priority mapping matrix based on the intention strength value; combine the timing control parameters in the emotion guidance target sequence to construct a voice-expression-action trimodal association network and generate a modal collaboration diagram; according to the priority mapping matrix and the modal collaboration diagram, calculate the execution priority and modal combination parameters of each type of instruction to form an instruction screening strategy, which includes an instruction priority sequence and modal combination rules.
[0051] Furthermore, the scenario-based emotional interaction instruction library is retrieved according to the instruction priority sequence to obtain candidate instruction sets, and the long-term usage habit features and scenario preference features are integrated into an interaction evaluation vector. The interaction evaluation vector is used to evaluate and screen the candidate instruction sets to obtain the optimal instruction set; an instruction execution dependency graph is constructed based on the modal combination rules. The instruction execution dependency graph is used to determine the parallel execution units and sequential execution units of the instructions, and the time window constraints of each execution unit are marked; the optimal instruction set is reorganized according to the instruction execution dependency graph, and the timing relationship of instruction execution is controlled through a recursive gating network to generate a basic instruction sequence. The recursive gating network is used to ensure the coordination of multimodal instructions.
[0052] Through the design of the above-mentioned scene interaction unit, the present invention realizes the precise mapping and efficient execution of emotional guidance goals to specific interactive instructions. Based on the two-layer sub-goal mapping network, the adaptability of the interaction strategy to scene constraints is improved; through the emotion-scene fusion matrix and the dynamic weight distribution mechanism, the precise expression of the interaction intention is achieved; the instruction screening strategy constructed by the hierarchical analysis and temporal combination method ensures the orderly execution of multimodal interaction instructions; the interaction evaluation mechanism combining the use habit characteristics and scene preference characteristics enhances the personalization of the interaction process; the instruction execution control based on the recursive gating network ensures the coordinated consistency of multimodal interactions such as voice, expression, and action. This scene interaction unit is particularly suitable for the precise interaction needs of emotional interactive companion dolls in complex scenes. It can convert abstract emotional guidance goals into specific and executable multimodal interaction instruction sequences, while ensuring the naturalness of the interaction and improving the execution efficiency of emotional guidance.
[0053] Furthermore, the workflow of the physiological regulation unit is as follows: extract periodic characteristic parameters from the physiological rhythm characteristics, and divide the basic instruction sequence into corresponding execution cycle units based on the periodic characteristic parameters; generate an intensity regulation matrix based on the physiological rhythm characteristics, and perform intensity mapping on the execution parameters of the voice module, expression module, and action module in each execution cycle unit; adjust the execution parameters in the execution cycle unit according to the intensity regulation matrix to obtain an adaptive instruction group; reorganize the adaptive instruction group according to the execution cycle unit to generate an execution instruction sequence.
[0054] The execution control module is used to control the voice module, expression module, and action module of the doll to complete emotional interaction behaviors based on the execution instruction sequence.
[0055] The database module includes an emotional memory library and a scenario-based emotional interaction instruction library. The emotional memory library stores historical emotional feature vector sequences, emotional state sequences, and scenario type sequences. The scenario-based emotional interaction instruction library stores voice instruction sets, expression instruction sets, and action instruction sets for different scenarios.
[0056] Furthermore, the present embodiment also provides a scene-based emotional interactive companion doll method, comprising collecting interaction data groups and scene data groups through the doll's multimodal sensing system to generate an extended data matrix; performing feature extraction on the extended data matrix, generating an emotion feature vector and a scene feature vector using a multimodal feature network, determining the user's current emotion state based on the emotion feature vector, and identifying the current scene type based on the scene feature vector; extracting the user's long-term usage habit characteristics, emotion preference characteristics, physiological rhythm characteristics, and scene preference characteristics based on the user's basic information, the historical emotion feature vector sequence, the historical emotion state sequence, and the historical scene type sequence stored in the emotion memory library, and constructing a user portrait; constructing an interactive decision module based on the user portrait, the user's current emotion state, and the current scene type, the interactive decision module including an emotion evolution unit, a scene interaction unit, and a physiological regulation unit ; The emotion evolution unit analyzes the emotion change rules based on the emotion preference characteristics and historical emotion state sequence, and predicts the emotion change trend in combination with the current emotion state and the current scene type, and determines the emotion guidance target sequence; the scene interaction unit generates an instruction screening strategy containing an instruction priority sequence and a modal combination rule based on the emotion guidance target sequence and the current scene type, and screens the multimodal interaction instructions corresponding to the current scene from the scenario-based emotion interaction instruction library according to the priority sequence, and combines the screened instructions to form a basic instruction sequence according to the modal combination rule; the physiological regulation unit adjusts the execution parameters of each instruction in the basic instruction sequence according to the physiological rhythm characteristics, and generates an execution instruction sequence; based on the execution instruction sequence, the voice module, expression module, and action module of the puppet are controlled to complete the emotional interaction behavior, and the emotional feature vector, emotional state, and scene type in the interaction scene are stored in the emotional memory library.
[0057] In summary, the present invention realizes the personalized and precise response of the emotional interactive doll, and is able to accurately identify the user's emotional state and scene requirements based on the comprehensive analysis of multimodal sensing and scene perception. Through the design of a dual-channel emotion prediction framework and a scene modulation mechanism, the system can make more accurate predictions of emotion change trends and dynamically adjust the interaction strategy according to the scene characteristics. Based on a hierarchical emotional state map and a progressive emotion guidance scheme, the system can provide a natural and smooth interactive experience while ensuring the safety of emotional regulation. Through modular multimodal interaction design and collaborative control of recursive gating networks, the system realizes the precise coordination of multiple interaction methods such as voice, expression, and action, and improves the execution efficiency of emotional guidance.
[0058] Example 2, reference Figures 1 to 4 , which is the second embodiment of the present invention, provides a scene-based emotional interactive companion doll system. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0059] In order to verify the effectiveness of the technical solution of the present invention in practical applications, an interactive therapy scene for children aged 3 to 6 years old in a children's rehabilitation center was selected for experimental verification. The experimental subjects were 40 children (20 boys and 20 girls), and the experiment lasted for 12 weeks. The experimental environment included three typical scenes: rehabilitation training room, children's activity room and personal rest area. Each scene was equipped with standard lighting equipment (400-600 lux) and constant temperature air conditioning (24±1℃) to ensure environmental consistency.
[0060] During data collection, a multimodal sensing system is used to monitor children. Facial images are captured using a built-in high-definition camera (30fps, 1080P); heart rate data is collected via a wearable photoplethysmography sensor (sampling rate 100Hz); pressure data is acquired using a distributed thin-film pressure sensor array (16×16 dot matrix, response time <50ms); and voice data is recorded via an omnidirectional microphone array (48kHz sampling rate). For scene data collection, environmental parameters including temperature, humidity, light, and noise are monitored in real time via an environmental sensor network. Time characteristics are recorded based on the system clock; location characteristics are tracked via a UWB positioning system; and characteristics of surrounding objects are captured using an RFID tag recognition system.
[0061] During the experiment, the researchers first configured the parameters of the doll system. The emotion type classifier uses an improved ResNet architecture and is pre-trained using a public emotion dataset; the emotion state judgment threshold is set to 0.75; and the scene matching threshold is set to 0.8. The emotion intensity quantification standard sets different benchmark values according to the age group of children: 3-4 years old group (baseline value 0.6), 4-5 years old group (baseline value 0.7), 5-6 years old group (baseline value 0.8). The time window parameters are set as follows: basic analysis window 5 seconds, emotion state assessment window 30 seconds, scene duration judgment window 3 minutes. In terms of interaction strategy, the basic weight ratio of the three types of strategies, namely comforting, companionship, and guidance, is set to 3:4:3, and dynamically adjusted according to the scene type.
[0062] Over 12 weeks of experimental operation, the system collected approximately 3,600 hours of interaction data. Data analysis showed that the accuracy of emotional state recognition gradually increased from an initial 76.5% to 94.2%. The scenario adaptability score increased from an initial 72.3 points to 91.8 points (out of 100). In the child acceptance assessment, 90.5% of test subjects expressed a positive willingness to interact, and the average duration of a single interaction increased from an initial 8.3 minutes to 23.7 minutes. The evaluation of the emotional regulation effect showed that the average relief time of negative emotions was shortened by 46.3%, while the duration of positive emotions was extended by 52.8%. Heart rate variability analysis showed that the children's stress index decreased by an average of 37.2% when accompanied by the dolls. This effect was particularly evident in the rehabilitation training scenario.
[0063] Based on the above experimental data, the research team conducted a comparative analysis of the present invention and the prior art, as shown in Table 1.
[0064] Table 1 Performance comparison between the present invention and the prior art
[0065] Evaluation Metrics Traditional doll making method Simple emotional interaction system System of the present invention Emotion recognition accuracy 65.3% 82.7% 94.2% Scene adaptability score No scene perception 75.6 91.8 Average interaction time 5.8min 12.4min 23.7h Negative emotion relief efficiency Baseline value 23.5% increase 46.3% increase Improvement in stress index Baseline value Down 18.9% down 37.2% Degree of personalization none Low high Multimodal collaboration Simple feedback Partial collaboration Fully collaborative Number of scenes covered 1-2 types 3-4 types More than 8 types Continuous learning ability none weak powerful
[0066] As can be seen from the comparative data in Table 1, the present invention has achieved significant improvements in key indicators such as emotion recognition accuracy, scenario adaptability, and interaction continuity. In particular, in terms of personalization and continuous learning capabilities, the innovative multimodal feature network and portrait construction module achieve precise grasp and dynamic adjustment of children's emotional needs. Experimental results demonstrate that the present invention not only outperforms existing technologies in technical indicators, but also demonstrates significant advantages in practical application, providing a new technical solution for the field of children's emotional companionship.
[0067] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A scene-based emotional interactive companion doll system, characterized by: include, A data acquisition module is used to collect interaction data groups and scene data groups to generate an extended data matrix; The feature analysis module is used to extract features from the extended data matrix, generate emotion feature vectors and scene feature vectors using a multimodal feature network, determine the user's current emotional state, and identify the current scene type; The portrait construction module is used to extract the user's long-term usage habit characteristics, emotional preference characteristics, physiological rhythm characteristics, and scene preference characteristics based on the emotional memory library to construct a user portrait; Interactive decision-making module, including emotion evolution unit, scene interaction unit and physiological regulation unit; The emotion evolution unit predicts emotion change trends based on emotion preference characteristics and historical emotion state sequences, and determines the emotion guidance target sequence. The scene interaction unit generates an instruction screening strategy containing an instruction priority sequence and modal combination rules based on the emotion guidance target sequence and the current scene type. It screens the instructions corresponding to the current scene according to the instruction priority sequence, and combines the screened instructions into a basic instruction sequence according to the modal combination rules. The physiological regulation unit adjusts the basic instruction sequence according to physiological rhythm characteristics to generate an execution instruction sequence.
2. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: Also includes, The execution control module is used to control the voice module, expression module, and action module of the doll to complete emotional interaction behaviors based on the execution instruction sequence; A database module, comprising an emotional memory library and a scenario-based emotional interaction instruction library, wherein the emotional memory library stores a sequence of historical emotional feature vectors, an emotional state sequence, and a scene type sequence, and the scenario-based emotional interaction instruction library stores a voice instruction set, an expression instruction set, and an action instruction set for different scenes; The interaction data group includes facial images, heart rate data, pressure data, and voice data, and the scene data group includes environmental parameters, time features, location features, and surrounding object features.
3. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: The method of generating the emotion feature vector and the scene feature vector by using the multimodal feature network includes the following steps: Preprocessing the extended data matrix to obtain a standardized data matrix, and dividing the standardized data matrix into an interaction sub-matrix and a scene sub-matrix; Processing the interaction sub-matrix and the scene sub-matrix through a dual-branch feature extraction network, the dual-branch feature extraction network including an interaction feature branch and a scene feature branch; wherein the interaction feature branch includes an emotion modality adaptation layer and an emotion fusion network, and the scene feature branch includes a scene encoder and a scene semantic understanding network; The emotion modality adaptation layer in the interaction feature branch is used to extract and transform the interaction sub-matrix, and the emotion fusion network is used to integrate the transformed features to output the emotion representation features; Using the scene encoder in the scene feature branch to perform feature encoding on the scene sub-matrix, and combining and processing through the scene semantic understanding network to output scene semantic representation features; Establishing a spatiotemporal alignment module to connect the interaction feature branch and the scene feature branch, wherein the spatiotemporal alignment module adopts an adaptive time window and an emotion delay compensation mechanism to input the emotion representation features and the scene semantic representation features into the emotion mapping layer for feature mapping; The mapped emotion representation features and scene semantic representation features are respectively input into the emotion feature generator and the scene feature generator. The emotion feature generator and the scene feature generator use a multi-head attention network to perform deep feature processing to generate emotion feature vectors and scene feature vectors.
4. The scene-based emotional interactive companion doll system according to claim 1, wherein: Determining the user's current emotional state and identifying the current scene type includes the following steps: Establishing an emotional state recognition model, the emotional state recognition model includes an emotional type classifier and an emotional intensity calculation unit, and inputting the emotional feature vector into the emotional state recognition model; Performing a classification operation on the emotion feature vector by the emotion type classifier to obtain an emotion type probability vector; Based on the comparison result of the maximum probability value in the emotion type probability vector and the emotion judgment threshold, the basic emotion type is determined in combination with the historical emotion state sequence; Inputting the emotion feature vector into the emotion intensity calculation unit, and calculating the emotion intensity value based on a preset emotion intensity quantification standard; Inputting the basic emotion type and the emotion intensity value into an emotion state mapping network to generate an emotion state characteristic curve, and determining the user's current emotion state based on the peak position and waveform characteristics of the emotion state characteristic curve; Establishing a scene recognition model, inputting the scene feature vector into the scene recognition model, and calculating the matching degree between the scene feature vector and each scene type template based on a built-in scene type template library; Based on the comparison result of the maximum matching degree and the scene matching threshold, the current scene type is determined in combination with historical scene data.
5. The scene-based emotional interactive companion doll system according to claim 1, wherein: The workflow of the portrait construction module is as follows: Read the historical emotional feature vector sequence, historical emotional state sequence, and historical scene type sequence from the emotional memory library, and concatenate the three sequences with the user's basic information to construct the original feature matrix; Performing time series decomposition on the historical emotion feature vector sequence in the original feature matrix, calculating the emotion state transition probability matrix, and obtaining the emotion preference feature; Mapping the historical scene type sequence to a scene association matrix, calculating the scene transition probability and scene duration distribution, and obtaining scene preference characteristics; Performing frequency domain transformation on the historical emotion feature vector sequence, combining it with the physiological parameters in the user basic information, calculating the period feature parameters, and obtaining the physiological rhythm feature; Aligning the historical emotional state sequence with the historical scene type sequence in time, calculating the joint distribution matrix, and obtaining long-term usage habit features; A three-layer user portrait is constructed based on the emotional preference feature, the scene preference feature, the physiological rhythm feature, and the long-term usage habit feature, and an inter-layer connection relationship is established through feature indexing.
6. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: The workflow of the emotion evolution unit is as follows: Obtain the historical emotional state sequence and emotional preference characteristics in the emotional memory library and construct a hierarchical emotional state map; Performing a time series feature analysis on the historical emotional state sequence, combining the hierarchical emotional state map, extracting emotional periodic patterns and transition critical points, and forming an emotional change feature set; A dual-channel emotion prediction framework is established, in which the first channel predicts state transition trends based on the emotion state transition probability matrix, and the second channel predicts emotion change patterns based on the emotion change feature set. The output results of the two channels are weighted and fused to generate a basic emotion prediction vector. Design a scene modulation subunit to convert the current scene type into a scene context vector, modulate the basic emotion prediction vector through a cross-attention network, and generate a context-adaptive emotion prediction result; Design a steady-state evaluation subunit to construct an emotional steady-state interval based on the user's current emotional state, calculate the degree of deviation between the situation-adaptive emotional prediction result and the emotional steady-state interval, and generate an emotional deviation vector; Obtaining at least two candidate emotion guidance directions from a preset doll companion therapy knowledge base according to the emotion deviation vector, and selecting the optimal emotion guidance direction based on the child emotion development theory; The optimal emotion guidance direction is compared with the user's current emotion state, the emotion guidance gradient is calculated, and an emotion guidance target sequence with timing control parameters is output.
7. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: The workflow of the scene interaction unit is as follows: Converting the timing control parameters in the emotion guidance target sequence into an interaction timing matrix, building a scene constraint model based on the interaction timing matrix and the current scene type, and outputting an interaction constraint vector; The emotion guidance target sequence is segmented using the interaction constraint vector, a sub-target mapping network is constructed for each time window, the emotion guidance target is dynamically associated with the scene constraint conditions, and a phased interaction strategy is generated; Performing sequential combination and hierarchical analysis on the phased interaction strategies, extracting interaction intention features and timing constraint relationships, constructing a priority mapping matrix and a modal collaboration graph, and generating an instruction screening strategy, which includes an instruction priority sequence and modal combination rules; Searching the scenario-based emotional interaction instruction library according to the instruction priority sequence to obtain a candidate instruction set, fusing long-term usage habit features and scenario preference features into an interaction evaluation vector, and using the interaction evaluation vector to evaluate and screen the candidate instruction set to obtain an optimal instruction set; Constructing an instruction execution dependency graph based on the modal combination rules; The optimal instruction set is reorganized according to an instruction execution dependency graph, and the timing relationship of instruction execution is controlled through a recursive gating network to generate a basic instruction sequence.
8. The scene-based emotional interactive companion doll system according to claim 1, characterized in that: The workflow of the physiological regulation unit is as follows: Extracting period characteristic parameters from the physiological rhythm characteristics, and dividing the basic instruction sequence into corresponding execution period units based on the period characteristic parameters; Generate an intensity adjustment matrix based on the physiological rhythm characteristics, and perform intensity mapping on the execution parameters of the voice module, expression module, and action module in each execution cycle unit; Adjusting execution parameters in the execution cycle unit according to the intensity adjustment matrix to obtain an adaptive instruction group; The adaptive instruction group is reorganized according to the execution cycle unit to generate an execution instruction sequence.
9. A method for using a scene-based emotionally interactive companion doll system, based on the scene-based emotionally interactive companion doll system according to any one of claims 1 to 8, characterized in that: Also includes, The interactive data set and the scene data set are collected through the multimodal sensing system of the doll to generate an extended data matrix; Performing feature extraction on the extended data matrix, generating an emotion feature vector and a scene feature vector using a multimodal feature network, determining a user's current emotion state based on the emotion feature vector, and identifying a current scene type based on the scene feature vector; Based on user basic information, historical emotional feature vector sequences, historical emotional state sequences, and historical scene type sequences stored in the emotional memory library, the user's long-term usage habit features, emotional preference features, physiological rhythm features, and scene preference features are extracted to construct a user profile. Constructing an interactive decision module according to the user portrait, the user's current emotional state, and the current scene type, wherein the interactive decision module includes an emotional evolution unit, a scene interaction unit, and a physiological regulation unit; The emotion evolution unit analyzes the emotion change rules based on the emotion preference characteristics and the historical emotion state sequence, and predicts the emotion change trend in combination with the current emotion state and the current scene type to determine the emotion guidance target sequence; The scene interaction unit generates an instruction screening strategy including an instruction priority sequence and a modal combination rule based on the emotion guidance target sequence and the current scene type, screens the multimodal interaction instructions corresponding to the current scene from the scene-based emotion interaction instruction library according to the priority sequence, and combines the screened instructions into a basic instruction sequence according to the modal combination rule; The physiological regulation unit adjusts the execution parameters of each instruction in the basic instruction sequence according to the physiological rhythm characteristics to generate an execution instruction sequence; Based on the execution instruction sequence, the voice module, expression module and action module of the doll are controlled to complete the emotional interaction behavior, and the emotional feature vector, emotional state and scene type in the interactive scene are stored in the emotional memory library.
Citation Information
Patent Citations
Information processing method and mobile terminal
CN105681900A
Method and system for processing multimedia scene interaction data
CN118260380A
Cited By
Children emotion pacifying intelligent doll
CN121808466A