Children emotion intelligent guiding method and system based on emotion instance objectization and explanation
By using a data structure that objectifies and externalizes emotion instances and a three-stage state machine protocol, combined with a physiological closed-loop parameter tuning mechanism, the problem of cognitive overload and insufficient physiological feature assessment in existing AI emotion interaction systems for young children is solved, achieving a more efficient emotion guidance effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 深圳市象形字科技股份有限公司
- Filing Date
- 2026-03-11
- Publication Date
- 2026-05-12
AI Technical Summary
Existing AI emotion interaction systems suffer from cognitive overload due to emotion tag binding when dealing with negative emotional outbursts in young children, lack of effective closed-loop control over the output of large language models, and lack of quantitative repositioning assessment based on physiological characteristics.
By constructing a data structure that objectifies and externalizes emotion instances, and employing a three-stage state machine protocol and a physiological closed-loop parameter tuning mechanism, the system decouples emotion tags from user identities, forcibly limits the output of the large language model, and performs quantitative evaluation and adaptive control based on physiological characteristics such as heart rate variability.
It significantly shortened the emotional reversion response time, reduced the LLM output overshoot rate, improved the physiological indicator reversion rate, and effectively reduced cognitive load.
Smart Images

Figure CN122018699A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and human-computer interaction technology, and in particular to a method and system for guiding children's emotional intelligence based on the objectification and externalization of emotional instances. Background Technology
[0002] With the rapid development of Large Language Models (LLM) and smart terminal technologies, AI interactive systems with affective computing capabilities have become a research and industrialization hotspot. However, existing technologies suffer from three types of systemic defects with inherent technical flaws when dealing with scenarios involving negative emotional outbursts in young children, rather than merely application-level optimization issues.
[0003] (a) The problem of cognitive overload caused by emotional label binding.
[0004] Existing AI emotion interaction systems generally adopt the "direct semantic empathy" model. Its underlying data processing logic is to directly attach the emotion recognition result (such as "anger") as an attribute label to the target user's conversation context and output it in the second person (such as "You are angry now"). From the perspective of computer data structure, this processing method establishes a strong reference relationship between emotion features and user ontology identifiers at the semantic layer.
[0005] Existing research in cognitive neuroscience has shown that when emotional labels are directly linked to self-identity, the regulatory function of the prefrontal cortex is suppressed, while the amygdala remains continuously activated, leading to cognitive overload. For children under 5 years old, whose prefrontal cortex is not yet fully developed, these negative effects are even more pronounced.
[0006] At the engineering level, such strong reference bindings also cause large language models to tend to generate ruminative texts (sentiment analysis, causal discussions, didactic advice) around the emotion label in multi-turn dialogues, further increasing cognitive load and forming a technical emotion reinforcement loop rather than a guidance loop.
[0007] (ii) The lack of effective closed-loop control over the output of large language models.
[0008] Large language models are essentially probabilistic text generation systems, and their outputs are uncertain. In emotion intervention scenarios, once the model output deviates from the preset guidance path (such as generating content that fails to evoke empathy or producing didactic text that triggers cognitive rumination), existing systems lack a technical mechanism to forcibly stop and redirect it.
[0009] Existing research has shown that relying solely on reinforcement learning human feedback (RLHF) fine-tuning and general safety alignment is insufficient to guarantee the output boundaries of a model in specific scenarios. In the highly sensitive vertical scenario of guiding children's emotions, a control protocol with greater engineering determinism than prompt constraints is needed.
[0010] (iii) Lack of quantitative repositioning assessment based on physiological characteristics.
[0011] Current affective computing systems typically rely on voice emotion recognition or facial expression analysis for feedback evaluation. However, these signals are noisy and unreliable during emotional outbursts in children. Heart rate variability (HRV), as an objective indicator of autonomic nervous system activity, can more stably reflect physiological changes in emotional arousal. Although HRV has been used in adult stress assessment, current technologies have not yet systematically incorporated it into real-time closed-loop control functions in the field of children's emotional guidance. Existing children's affective computing systems largely rely on external signals and lack quantitative repositioning evaluation mechanisms based on objective autonomic nervous system indicators such as HRV. This results in an inability to dynamically adjust intervention strategies (such as regulating the rhythmic cycle of sensory stimulation) according to the child's real-time physiological calming. This "open-loop" or "weak feedback" model is ill-suited to addressing the individual differences and dynamic changes during children's emotional outbursts, representing a pressing technical challenge in this field.
[0012] In summary, existing technologies have gaps in three dimensions: "emotion tag unbinding mechanism", "LLM output mandatory constraint protocol" and "physiological feature closed-loop parameter tuning". This invention proposes solutions to the above three specific technical problems. Summary of the Invention
[0013] (a) The technical problem that the invention aims to solve
[0014] This invention aims to solve the following three specific technical problems:
[0015] Technical Issue 1: How to debind and isolate emotion instances from user ontology identifiers at the data structure level, thereby cutting off the strong reference binding between emotion labels and user identities in the output of large language models from a technical perspective, and reducing the cognitive load in children's emotion guidance scenarios.
[0016] Technical Problem 2: How to construct a large language model output constraint protocol with engineering determinism, so that the model's output is forcibly limited to a preset temporal structure and content boundaries in the scenario of guiding children's emotions, without relying on the probabilistic alignment effect of the model.
[0017] Technical problem 3: How to systematically incorporate physiological characteristic data (especially heart rate and heart rate variability) into the quantitative evaluation function of the guidance effect, and build an adaptive closed-loop control mechanism that combines software and hardware based on the evaluation results.
[0018] (II) Technical Solution
[0019] To address the aforementioned technical problems, the first aspect of this invention provides a method for guiding children's emotional intelligence based on the objectification and externalization of emotional instances, executed by a computer device, comprising the following steps:
[0020] The system acquires interaction data from target users and inputs it into an emotion recognition model, outputting emotion feature values that include valence and arousal components.
[0021] Based on the objectified data structure of emotion instances, a virtual object role object is constructed: a virtual object role object independent of the target user's ontology identifier is instantiated according to the emotion feature value. The object carries a three-element attribute group, including an independent object name identifier, an embodied action-driven feature code, and an externalized narrative logic label, so that the instance identifier of the virtual object role object and the ontology identifier of the target user have no reference relationship at the data level.
[0022] Execute an emotion repositioning state machine protocol with mandatory state transition constraints: Construct system-level prompts and input them into a natural language generation model. The system-level prompts are configured with a resident agent identifier and a ternary attribute of the virtual object, and include narrative perspective constraint instructions. This forces the natural language generation model to strictly output text in stages according to a preset temporal structure, prohibiting free jumps or skipping of any stage. The preset temporal structure includes at least:
[0023] Phase A – Naming Externalization: Based on the externalized narrative logic tags, generate text describing the independent object name identifier in third-person declarative sentences, and the output text does not contain second-person emotion tag sentences that bind emotional attributes to the target user ontology;
[0024] Phase B – Action Soothing: Based on the resident agent identifier, soothing text is generated, and a first control instruction set is output simultaneously. The first control instruction set is used to drive the terminal to perform soothing actions or generate virtual interactive feedback.
[0025] Phase C – State Reset: Generate guiding text containing a preset object identifier and simultaneously output a second control instruction set, which is used to drive the terminal to output sensory stimulation signals with a preset rhythmic cycle.
[0026] Output the text generated in each stage and the corresponding set of control instructions.
[0027] In a preferred embodiment, the objectified data structure of the emotion instance is stored in an emotion object role database, and a key-value mapping rule is used to map the valence component and arousal component of the emotion feature value to the ternary attribute group. Specifically, when the valence component of the emotion feature value is negative and the arousal component is greater than a first arousal threshold, the externalized narrative logic label is assigned to the externally triggered type, and the embodied action-driven feature code is assigned to the high-energy output type; when the valence component of the emotion feature value is negative and the arousal component is less than a second arousal threshold, the externalized narrative logic label is assigned to the internally consumed type, and the embodied action-driven feature code is assigned to the low-energy slow type.
[0028] In another preferred embodiment, the state machine protocol further includes content boundary constraints and subject forced substitution constraints. The content boundary constraints prohibit the natural language generation model from outputting text paragraphs containing emotion causal analysis, cognitive reconstruction suggestions, or moral evaluations. The subject forced substitution constraint mandates that the independent object name identifier be inserted as the initial subject in the output text of stage A, prohibiting the use of the target user's personal pronoun as the initial subject.
[0029] In another preferred embodiment, the method further includes a physiological closed-loop parameter tuning step: after executing stage C, acquiring the physiological characteristic data of the target user, and calculating the repositioning score based on the physiological characteristic data; if the repositioning score does not reach a preset safety threshold, dynamically updating the control parameters of the state machine protocol, and cyclically executing the emotion repositioning state machine protocol until the termination condition is met. The dynamically updated control parameters include at least one of adjusting the rhythmic period parameter of the sensory stimulation signal in stage C, adjusting the length limit parameter of the output text, or adjusting the audio playback rate parameter.
[0030] Furthermore, the repositioning score is calculated based on the weighted sum of heart rate data, heart rate variability data, voice energy change rate data, and limb movement amplitude change rate data; wherein, the sum of the weight coefficients of heart rate and heart rate variability related indicators is not less than 0.60 of the total weight, and is higher than the weight ratio of other indicators.
[0031] Furthermore, during the loop execution, the loop will terminate and a guardian notification instruction will be triggered when any of the following conditions are met: the number of loops reaches the maximum number of loops limit; preset high-risk semantic features (including semantic patterns involving self-harm, attacking others, or extreme rejection) are detected in the interaction data; or the data rate of physiological features continuously exceeds the preset proportion of the safety baseline and the duration exceeds the preset duration.
[0032] Furthermore, the method also includes output verification and fault tolerance steps: real-time verification of the single output of the natural language generation model; if the output content violates the temporal structure, subject constraint, or content boundary constraint of the state machine protocol, a regeneration instruction is triggered; when the number of retries reaches a preset threshold, a degradation response strategy is executed, and a preset fixed appeasement script is output; if a degradation response is triggered during loop execution, the current loop is terminated in advance.
[0033] A second aspect of the present invention provides a child emotional intelligence guidance system based on the objectification and externalization of emotional instances, comprising:
[0034] The sensing and acquisition module is used to acquire interaction data from the target user;
[0035] An emotion feature extraction module is used to map the interaction data into emotion feature vectors;
[0036] The object role instantiation module is used to generate virtual object role object instances carrying a three-element attribute group based on the emotion feature vector. The three-element attribute group includes an independent object name identifier, an embodied action-driven feature code, and an externalized narrative logic label. The instance identifier and the target user ontology identifier are independent of each other at the data level.
[0037] The resident agent module is used to generate and maintain resident agent identifiers;
[0038] A state machine protocol engine is used to construct system-level prompt words and constrain the natural language generation model to output text according to a preset temporal structure. The preset temporal structure includes at least a naming externalization stage, an action soothing stage, and a state return stage.
[0039] A multimodal output module is used to execute the text output and corresponding control commands;
[0040] The physiological closed-loop parameter tuning module is used to collect physiological characteristic data of the target user to calculate the positioning score. When the score is lower than the safety threshold, it feeds back the parameter tuning instruction to the state machine protocol engine to dynamically update the control parameters of the state machine protocol.
[0041] In a preferred embodiment, the system further includes an output verification and fault tolerance module and a risk monitoring and notification module: the output verification and fault tolerance module is used to monitor whether the output of the natural language generation model conforms to the state machine protocol constraints, and trigger a regeneration or degradation strategy when a violation occurs; the risk monitoring and notification module is used to trigger a guardian notification instruction when a high-risk semantic feature or abnormal physiological indicator is detected and timed out continuously.
[0042] A third aspect of the present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method described in any of the preceding claims, or to constitute the system described in any of the preceding claims.
[0043] (III) Beneficial Effects
[0044] The beneficial effects of this invention are derived from functional verification data conducted by the internal R&D team in a controlled environment. Compared with existing technical solutions, this invention has the following quantifiable and verifiable technical effects:
[0045] Effect 1: Significantly reduced emotion reversion response time. In internal comparative tests simulating emotional outbursts in children, the average emotion reversion time using the method of this invention was reduced by approximately 42% compared to traditional direct semantic empathy methods.
[0046] Effect 2: The LLM output out-of-bounds rate is significantly reduced. The three-stage state machine protocol of this invention, combined with the output verification mechanism, reduces the rate of out-of-bounds content such as emotion cause analysis and didacticism in the output of the natural language generation model from 23.7% to 2.8%, effectively avoiding cognitive overload.
[0047] Effect 3: Improved recovery rate of physiological indicators. After introducing a closed-loop physiological parameter tuning mechanism based on HRV weights, the proportion of children in the experimental group whose heart rate and HRV indicators returned to the resting baseline range increased by approximately 42% compared to the control group without physiological feedback. Attached Figure Description
[0048] Figure 1 is a schematic diagram of the overall structure of the child's emotional intelligence guidance system based on the objectification and externalization of emotional instances in an embodiment of the present invention.
[0049] Figure 2 is a flowchart illustrating the child's emotional intelligence guidance method based on the objectification and externalization of emotional instances in an embodiment of the present invention.
[0050] Figure 3 is a schematic diagram of the data structure and query timing architecture of the emotion object role database in an embodiment of the present invention.
[0051] Figure 4 is a schematic diagram of the state transition of the three-stage emotion repositioning state machine protocol in an embodiment of the present invention.
[0052] Figure 5 is a diagram of the joint architecture of the hardware output layer, system-level prompt word generation, and physiological closed-loop parameter tuning in an embodiment of the present invention. Detailed Implementation
[0053] The specific embodiments of the present invention will be further described in detail below with reference to Figures 1 to 5 and specific examples. The following examples are only used to illustrate the technical solutions of the present invention and do not limit the scope of protection of the present invention.
[0054] It should be stated beforehand that the "resident agent intelligent agent identifier" appearing in this document and accompanying figures can also be simply referred to as "resident agent" or manifested as a specific "soothing role" in specific application scenarios. All three refer to the same fixed identity identifier independent of the target user and the emotional object at the system data level. In addition, the "control instruction set" (including the first control instruction set and the second control instruction set) mentioned in the specification is a higher-level concept. In specific hardware implementations, it can manifest as "hardware control instructions" (such as driving a robotic arm, LED lights, vibration motors, etc.), and in pure software or virtual interaction implementations, it can manifest as "virtual interaction instructions" (such as screen animation rendering, virtual character actions, etc.). Both fall within the protection scope of this invention.
[0055] Example 1: System Architecture, Data Compliance, and Emotional Object Role Database Design
[0056] As shown in Figure 1, the children's emotional intelligence guidance system of the present invention includes a perception acquisition module, an emotion feature extraction module, an object role instantiation module, a resident subject intelligent agent module, a state machine protocol engine, an output verification and fault tolerance module, a multimodal output module, a physiological closed-loop parameter tuning module, and a risk monitoring and notification module.
[0057] Data Compliance and Privacy Protection Design: In the system architecture of this invention, the design of the sensing and data acquisition module and the risk monitoring module follows the principle of "data minimization." For sensitive data such as children's heart rate, HRV, voice, and images, the system performs real-time streaming processing in memory to calculate feature values or repositioning scores, and does not persistently store raw biometric data by default, unless a high-risk warning is triggered requiring evidence retention. Optionally, the system can also be expanded to include a guardian configuration module, allowing parents to customize the granularity of data collection, storage duration, and notification methods to ensure compliance with the Personal Information Protection Act and GDPR requirements regarding child data protection.
[0058] Referring to steps S10-S30 in Figure 2 and Figure 3, this embodiment details the data structure design of the emotion object role database. In the prior art, the system typically records emotion tags (such as user_session["emotion"]="anger") in the session context and generates "The user is now angry, please comfort him" in the prompt. This method establishes a strong reference relationship between emotion and user at the data level.
[0059] The emotional object role database of this invention adopts an independently instantiated data structure, as shown in Figure 3. Key fields include: an independently generated object instance identifier (instance_id) that has no reference relationship with the user identifier (the user_ref field value is null); and ternary attributes including an independent object name identifier, an embodied action-driven feature code, and an externalized narrative logic tag. The externalized narrative logic tag is injected as a constraint parameter into the prompt word generation function, forcing the system to use a third-person narrative perspective triggered by external events when generating text, rather than a second-person perspective.
[0060] In a preferred mapping rule example, a first arousal threshold is set to be greater than or equal to a second arousal threshold. Specifically: when the valence component of the emotional feature value is negative and the arousal component is greater than the first arousal threshold, the externalized narrative logic label is assigned an externally triggered type, and the embodied action-driven feature code is assigned a high-energy output type; when the valence component of the emotional feature value is negative and the arousal component is less than the second arousal threshold, the externalized narrative logic label is assigned an internally consumed type, and the embodied action-driven feature code is assigned a low-energy slow type. Furthermore, if a high-frequency vibrato component is detected in the acoustic feature, regardless of the arousal level, it is preferentially mapped to a "high-frequency micro-amplitude vibration type" action-driven feature code to match the child's trembling physiological manifestations.
[0061] Note: The specific role names in the independent object name identifier (such as "Snoring Monster" or "Little Cloud") belong to the content configuration of the application layer and do not constitute a limitation of the technical features of this invention. They can be freely configured by the implementer.
[0062] Example 2: Engineering Implementation of Three-Phase State Machine Protocol and System-Level Prompt Words
[0063] Referring to step S40 in Figure 2 and Figure 4, this embodiment details the engineering implementation of the state machine protocol engine.
[0064] As shown in Figure 4, the state machine protocol engine maintains a finite state machine with a state set including {IDLE, NAMING, SOOTHING, RETURNING, EVALUATING, ALERT}. The state transitions are deterministic functions, disallowing free jumps based on the output content of a large model.
[0065] For a specific emotion feature vector scenario, the system-generated system-level prompts include: a role setting module, which retrieves the identifier from the resident subject intelligent agent module and sets the role identity of the soothing agent; a parameter slot module, which injects the three-element attributes of the currently detected virtual object role; and an output format enforcement module, which stipulates that the model must and can only output in the order of stage A (name externalization), stage B (action soothing), and stage C (state return).
[0066] In a preferred embodiment, the maximum character limit parameters (L_A, L_B, L_C) set for stages A, B, and C can be dynamically configured according to the child's age. For example, exemplary values could be: L_A=10, L_B=12, L_C=15 for the 3-4 year old group; and L_A=15, L_B=18, L_C=20 for the 5-6 year old group. It should be noted that the above character limit is only an exemplary value of the preferred embodiment, and those skilled in the art can adjust it according to the speech synthesis speed of the actual device and user test feedback.
[0067] If the output of the natural language generation model violates any constraint (such as the subject at the beginning of a sentence not being an object name, or containing emotionally didactic vocabulary), the state machine is triggered to regenerate. After retries reach a preset number (e.g., 3 times), a degradation response strategy is executed (outputting a preset fixed reassurance script). This mechanism ensures that even with the probabilistic output of the LLM, the system can still maintain deterministic boundaries in engineering. In particular, in loop execution scenarios containing physiological closures, if the number of retries in the current loop reaches a preset threshold and triggers a degradation response, the system will terminate the execution of the current loop early, without waiting for stage C to complete, and will decide whether to directly end the process or proceed to the next loop judgment based on the current repositioning score, in order to avoid resource waste and user annoyance caused by invalid loops.
[0068] Example 3: Specific Implementation of the Closed-Loop Parameter Tuning Mechanism for Physiological Characteristics
[0069] Referring to step S60 in Figure 2 and Figure 5, this embodiment details the specific implementation of the closed-loop parameter tuning mechanism for physiological characteristics.
[0070] As shown in Figure 5, after the execution phase C outputs sensory stimulation signals with a preset rhythmic cycle (such as LEDs gradually brightening and dimming, and the vibration module rhythmically vibrating), the physiological closed-loop parameter tuning module evaluates based on the physiological characteristic data of the target user.
[0071] In a preferred embodiment with a specific algorithm formula, the calculation logic of the relocation score Q is as follows:
[0072] Q = w1·f1(HR) + w2·f2(HRV) + w3·f3(vocal) + w4·f4(motion)
[0073] Wherein, f1(HR) is the heart rate normalization function, f2(HRV) is the heart rate variability normalization function based on indicators such as the root mean square difference of continuous RR intervals (RMSSD), f3(vocal) is the voice energy change rate function, and f4(motion) is the limb movement amplitude change rate function. w1 to w4 are weighting coefficients and ∑wᵢ=1. Specifically, each normalization function is configured such that: when the monitored physiological indicators tend towards the baseline value of the user in a resting state, the function output value increases; otherwise, it decreases.
[0074] To ensure that physiological indicators accurately reflect true emotional arousal, this solution, tailored to the developmental characteristics of children's autonomic nervous system, assigns higher weights to heart rate and heart rate variability-related indicators than to other indicators (w1+w2 ≥ 0.60, e.g., w1=0.35, w2=0.35, w3=0.15, w4=0.15). This is a specific technical adaptation for children's incomplete prefrontal cortex development and their greater reliance on the autonomic nervous system for emotion regulation, distinguishing it from general adult stress assessment models. The safety threshold (Q_safe) can be set to 0.65. When Q is lower than Q_safe, a parameter tuning instruction is triggered, such as extending the initial values of the LED's gradual brightening and dimming durations (e.g., 4 seconds) by a step size ΔT (e.g., 1 second).
[0075] It is particularly important to note that the specific weighting coefficients, safety threshold values, initial rhythm duration, and parameter tuning step sizes in the above formulas are merely exemplary values or preferred settings for specific testing scenarios. Those skilled in the art should understand that, without departing from the core logic of this invention—"using a higher weighting percentage for heart rate and heart rate variability-related indicators than other indicators for multiple weighted evaluations" and "dynamically adjusting rhythm and text parameters when the safety threshold is not reached"—the values of the above parameters can be routinely replaced or dynamically calibrated based on sensor hardware accuracy, the specific age group of the target user, or individual baseline differences.
[0076] Example 4: Complete Scenario Execution Example (Children's Separation Anxiety When Starting Kindergarten)
[0077] This embodiment illustrates a complete execution flow in conjunction with Figures 1 to 5.
[0078] As shown in step S10 of Figures 1 and 2, the sensing and acquisition module acquires acoustic data, visual data, and heart rate data transmitted from the wristband.
[0079] In step S20, the emotion feature extraction module outputs an emotion feature vector with high negative valence and high arousal.
[0080] As shown in step S30 of Figure 2 and Figure 3, the system instantiates a virtual object role in the database (e.g., the name identifier is "Panic Beast", the driving feature code is "high-frequency micro-amplitude vibration", and the narrative tag is "internal fear type"). At this time, the instance identifier is completely independent, and user_ref=null.
[0081] As shown in step S40 of Figure 2 and Figure 4, the state machine enters the NAMING (name externalization) state, constructs system-level prompt words, and constrains the model output.
[0082] As shown in step S50 of Figure 2 and Figure 5, the multimodal output module executes synchronously. During speaker output phase A (text), after a wait period, phase B (text) is output, simultaneously triggering the first control instruction set (e.g., robotic arm closing drive command, with time synchronization error controlled within 200 milliseconds); during phase C (text), the second control instruction set (e.g., LED light array breathing rhythm and low-frequency white noise) is simultaneously triggered. Throughout this process, the output verification and fault tolerance module monitors the generated content in real time to ensure no violations occur during output.
[0083] As shown in step S60 of Figure 2 and Figure 5, the system collects evaluation data. If the initial repositioning score Q does not reach the safety threshold, the system dynamically updates parameters (such as extending the duration of the breathing light rhythm and reducing the voice playback rate) and executes S40 again in a loop. If the termination condition is met after multiple loops, a repositioning completion prompt is output. If the heart rate is detected to continuously exceed the preset proportion of the safety baseline, the risk monitoring and notification module sends a notification instruction to the guardian.
[0084] Example 5: Implementation of Electronic Devices
[0085] This embodiment provides an electronic device, including at least one of a processor (CPU / NPU), a memory, a microphone array, a camera, a speaker, a display module, an LED lighting module, a vibration module, a robotic arm drive module, and a communication module. When the processor executes program instructions in the memory, it implements the steps described in the foregoing method embodiments, or constitutes the system described in the foregoing system embodiments.
[0086] The electronic device can employ either cloud-based or local inference modes. In cloud-based collaborative mode, the local terminal is responsible for multimodal data acquisition and hardware instruction execution, while the cloud server is responsible for the state machine protocol engine and LLM inference. To meet the time synchronization error requirements of the first control instruction set, end-to-cloud communication and instruction delivery are configured as low-latency, high-priority channels.
[0087] Example 6: Optional Variations and Extensions
[0088] 1. The ternary attribute data structure of this invention is not limited to a specific character image. The character can be anthropomorphic, a natural phenomenon, or an abstract symbol, as long as it meets the constraints of data unbinding and objective narrative.
[0089] 2. In addition to the basic three-stage state machine protocol, the present invention can be extended to include a "pre-emotional confirmation stage" or a "post-emotional review stage". As long as the mandatory timing and format constraints of the core stage are retained, it will fall within the scope of protection.
[0090] 3. In addition to heart rate and heart rate variability, other physiological indicators such as skin conductance (EDA) or electroencephalography (EEG) can also be introduced for physiological assessment, as long as the multivariate weighting logic has not been fundamentally changed.
[0091] The above embodiments are merely preferred embodiments of the present invention. Any equivalent substitutions and improvements made based on the technical features disclosed in this invention, within the spirit and principles of the invention, should be included within the scope of protection of this invention.
Claims
1. A method for guiding children's emotional intelligence based on the objectification and externalization of emotional instances, executed by a computer device, characterized in that: Includes the following steps: The system acquires interaction data from target users and inputs it into an emotion recognition model, outputting emotion feature values that include valence and arousal components. Based on the objectified data structure of emotion instances, a virtual object role object is constructed: a virtual object role object independent of the target user's ontology identifier is instantiated according to the emotion feature value. The object carries a three-element attribute group, including an independent object name identifier, an embodied action-driven feature code, and an externalized narrative logic label, so that the instance identifier of the virtual object role object and the ontology identifier of the target user have no reference relationship at the data level. Execute the emotion return state machine protocol with mandatory state transition constraints: construct system-level prompt words and input them into the natural language generation model. The system-level prompt words are configured with a resident subject intelligent agent identifier and the three-element attribute of the virtual object role object, and contain narrative perspective constraint instructions, which force the natural language generation model to output text in stages according to the preset temporal structure, and prohibit free jumping or skipping any stage. The preset timing structure includes at least: Phase A – Naming Externalization: Based on the externalized narrative logic tags, generate text describing the independent object name identifier in third-person declarative sentences, and the output text does not contain second-person emotion tag sentences that bind emotional attributes to the target user ontology; Phase B – Action Soothing: Based on the resident agent identifier, soothing text is generated, and a first control instruction set is output simultaneously. The first control instruction set is used to drive the terminal to perform soothing actions or generate virtual interactive feedback. Phase C – State Reset: Generate guiding text containing a preset object identifier and simultaneously output a second control instruction set, which is used to drive the terminal to output sensory stimulation signals with a preset rhythmic cycle. Output the text generated in each stage and the corresponding set of control instructions.
2. The method according to claim 1, characterized in that, The objectified data structure of the emotion instance is stored in the emotion object role database, and the valence component and arousal component of the emotion feature value are mapped to the ternary attribute group using key-value mapping rules. The mapping rules include at least: When the valence component of the emotional feature value is negative and the arousal component is greater than the first arousal threshold, the externalized narrative logic label is assigned the external trigger type, and the embodied action-driven feature code is assigned the high-energy output type. When the valence component of the emotional feature value is negative and the arousal component is less than the second arousal threshold, the externalized narrative logic label is assigned the internal consumption type, and the embodied action-driven feature code is assigned the low-energy slow type.
3. The method according to claim 1, characterized in that, The state machine protocol also includes content boundary constraints and subject forced substitution constraints: The content boundary constraints prohibit the natural language generation model from outputting text paragraphs that contain emotional causal analysis, cognitive restructuring suggestions, or moral evaluations. The subject-forced substitution constraint mandates that the independent object name identifier be inserted as the sentence-initial subject in the output text of stage A, and prohibits the use of the target user's personal pronoun as the sentence-initial subject.
4. The method according to claim 1, characterized in that, It also includes physiological closed-loop parameter tuning steps: After performing stage C, the physiological characteristic data of the target user is obtained, and the relocation score is calculated based on the physiological characteristic data. If the repositioning score does not reach the preset safety threshold, the control parameters of the state machine protocol are dynamically updated, and the emotion repositioning state machine protocol is executed cyclically until the termination condition is met. The dynamically updated control parameters include at least one of the following: the rhythmic period parameter of the sensory stimulus signal in adjustment phase C, the length limit parameter of the output text, or the audio playback rate parameter.
5. The method according to claim 4, characterized in that, The repositioning score is calculated based on a weighted sum of heart rate data, heart rate variability data, voice energy change rate data, and limb movement amplitude change rate data. Among them, the sum of the weight coefficients of heart rate and heart rate variability-related indicators shall not be less than 0.60 of the total weight, and shall be higher than the weight ratio of other indicators.
6. The method according to claim 4, characterized in that, During the loop execution, the loop terminates and a guardian notification instruction is triggered when any of the following conditions are met: The loop count has reached the maximum loop count limit; Preset high-risk semantic features were detected in the interactive data. These preset high-risk semantic features include semantic patterns involving self-harm, aggression towards others, or extreme rejection. The physiological data rate continuously exceeds the preset proportion of the security baseline and the duration exceeds the preset duration.
7. The method according to claim 1, characterized in that, It also includes output verification and fault tolerance steps: The output of the natural language generation model is validated in real time. If the output violates the temporal structure, subject constraint, or content boundary constraint of the state machine protocol, a regeneration instruction is triggered. When the number of retries reaches a preset threshold, a downgraded response strategy is executed, and a preset fixed appeasement script is output. If a degradation response is triggered during the execution of the loop, the current loop iteration will be terminated early.
8. A child emotional intelligence guidance system based on the objectification and externalization of emotional instances, characterized in that, include: The sensing and acquisition module is used to acquire interaction data from the target user; An emotion feature extraction module is used to map the interaction data into emotion feature vectors; The object role instantiation module is used to generate virtual object role object instances carrying a three-element attribute group based on the emotion feature vector. The three-element attribute group includes an independent object name identifier, an embodied action-driven feature code, and an externalized narrative logic label. The instance identifier and the target user ontology identifier are independent of each other at the data level. The resident agent module is used to generate and maintain resident agent identifiers; A state machine protocol engine is used to construct system-level prompt words and constrain the natural language generation model to output text according to a preset temporal structure. The preset temporal structure includes at least a naming externalization stage, an action soothing stage, and a state return stage. A multimodal output module is used to execute the text output and corresponding control commands; The physiological closed-loop parameter tuning module is used to collect physiological characteristic data of the target user to calculate the positioning score. When the score is lower than the safety threshold, it feeds back the parameter tuning instruction to the state machine protocol engine to dynamically update the control parameters of the state machine protocol.
9. The system according to claim 8, characterized in that, It also includes an output verification and fault tolerance module and a risk monitoring and notification module: The output verification and fault tolerance module is used to monitor whether the output of the natural language generation model conforms to the state machine protocol constraints, and triggers a regeneration or degradation strategy when a violation occurs. The risk monitoring and notification module is used to trigger a notification instruction from the guardian when a high-risk semantic feature or abnormal physiological indicator is detected for an extended period of time.
10. An electronic device comprising a processor, a memory, and a computer program stored in the memory, characterized in that, When the processor executes the computer program, it implements the method described in any one of claims 1 to 7, or constitutes the system described in claim 8 or 9.