A multimodal emotion intervention data protocol and co-location evaluation method based on neural mapping parameters
By defining ESDS and a multimodal collaborative repositioning algorithm, the problems of inconsistent data and inaccurate evaluation in children's emotional intervention were solved, achieving device interoperability and high-precision physiological repositioning evaluation, thus ensuring the consistency and effectiveness of intervention methods.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 深圳市象形字科技股份有限公司
- Filing Date
- 2026-03-20
- Publication Date
- 2026-05-26
AI Technical Summary
Existing child emotion intervention technologies suffer from problems such as inconsistent data interfaces, failure to translate cultural archetypes into machine instructions, and a single evaluation dimension lacking physiological basis. These issues result in poor device interoperability, inconsistent intervention effects, and a high rate of misjudgment.
We define an Emotional State Data Structure (ESDS) and a multimodal collaborative attribution evaluation algorithm. By strongly binding emotional prototypes with machine instructions through Role_ID, Action_Code, and Narrative_Tag, and combining neural mapping parameters and multimodal signal analysis, we can achieve standardized and high-precision evaluation of emotion intervention.
It achieves interoperability between devices from different manufacturers, ensuring consistency of intervention methods and high-precision physiological repositioning assessment, and reducing the misjudgment rate in extreme emotional scenarios.
Smart Images

Figure CN122075873A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence, human-computer interaction and digital healthcare, specifically to an Emotion Token Protocol that standardizes emotional prototypes into machine-executable data protocols, and a quantitative evaluation method for emotion attribution based on multimodal biosignal collaborative feedback. Background Technology
[0002] Currently, smart hardware and software applications in the field of children's emotional intervention are emerging in large numbers. However, existing technologies suffer from the following serious problems of fragmentation and non-standardization:
[0003] First, data interfaces are not standardized. Existing emotion recognition systems mostly output unstructured natural language labels (such as "angry" or "sad") or simple numerical values. Downstream actuators (such as robotic arms, lights, and speech synthesizers) lack unified driving standards, resulting in incompatibility between devices from different manufacturers and making it difficult to form a standardized intervention strategy library.
[0004] Second, there is a disconnect between cultural archetypes and machine instructions. While many products incorporate IP characters (such as bears and rabbits), these characters remain merely visual displays or storytelling elements, failing to translate them into underlying machine control instructions. In other words, there is a lack of mandatory technical binding between "characters," "actions," and "narrative logic," resulting in intervention effects relying on randomly generated content and lacking scientific consistency.
[0005] Third, the evaluation of emotional attunement is based on a single dimension and lacks physiological basis. Existing systems typically rely solely on voice emotion recognition or simple facial expression analysis to determine whether a user is "calm." These indicators are highly noisy and unreliable in scenarios involving high arousal (crying) in children. More importantly, existing evaluations lack multi-dimensional collaborative algorithms that incorporate real physical bioelectrical signals (such as heart rate variability (HRV) and skin conductance (GSR), making it impossible to accurately quantify the true physiological degree of "emotional attunement."
[0006] Fourth, the physical output parameters lack scientific basis and psychological support. Existing light and sound feedback are mostly based on experience (e.g., "red represents anger"), lacking a mapping standard based on specific frequencies, waveform curves, and emotional archetypes verified by neuroscience. Furthermore, existing interactions often employ second-person empathy (e.g., "You are angry"), ignoring the psychological theory of "self-distancing." This theory states that in a state of high emotional arousal, individuals can effectively reduce amygdala activation by examining emotional objects from a third-person perspective. Current technologies lack mechanisms to enforce this theory (e.g., mandatory third-person narrative constraints), potentially exacerbating emotional involvement rather than alleviating it.
[0007] Therefore, there is an urgent need to establish a standardized emotional computing data protocol to transform emotional prototypes into machine-executable underlying instructions, and to construct a multimodal collaborative positioning evaluation system that integrates real bioelectrical signals, so as to promote the standardized and scientific development of the emotional intelligence industry. Summary of the Invention
[0008] This invention aims to address the problems of inconsistent data formats, failure to translate cultural prototypes into machine instructions, and inaccurate attribution assessments in existing technologies for emotion intervention. It establishes an industry-standard technical framework by defining an Emotional State Data Structure (ESDS) and a multimodal collaborative attribution assessment algorithm.
[0009] To achieve the above objectives, the present invention provides the following technical solution:
[0010] Firstly, this invention provides an emotion computing interface protocol based on emotion prototype roles, defining a standardized Emotional State Data Structure (ESDS). The ESDS includes the following strongly bound fields: Role_ID, a unique identifier selected from a pre-defined finite set containing at least six basic emotion prototypes corresponding to different combinations of valence and arousal; Action_Code, a uniquely mapped embodied action-driven feature code to the Role_ID, used to directly control the execution mechanism of the terminal device to execute a specific sequence of embodied actions; and Narrative_Tag, a uniquely mapped narrative logic constraint parameter to the Role_ID, used to enforce the person, tense, and objectified content boundaries of the natural language generation model, to achieve emotion objectification based on self-distancing theory. Any system based on this protocol, when generating emotion intervention instructions, must encapsulate the unstructured emotion recognition results into the ESDS format, and the combination of the three fields maintains indivisible atomicity within a single intervention cycle.
[0011] Secondly, this invention provides a physical layer output standard based on neural mapping parameters. For a specific Role_ID, the protocol defines a range of standard physical parameters in the device driver lookup table (LUT). For high-arousal negative emotion prototypes, the flicker frequency range of the visual output features is specified as 3-5Hz, the color temperature range as 3000K-4500K, and the dynamic waveform as an acute-angle geometric shape. For low-arousal negative emotion prototypes, the flicker frequency range of the visual output features is specified as 0.1-0.5Hz, the color temperature range as 5000K-6500K, and the dynamic waveform as a smooth sine wave. These physical parameters serve as the neural activation signature of the emotion prototype, forcing downstream actuators to call values within the corresponding parameter ranges for sensory feedback output.
[0012] Thirdly, this invention provides a multimodal collaborative emotion repositioning evaluation method, comprising the following steps: Step 1, real-time acquisition of multimodal signal streams of the user during the intervention process, wherein the signal streams include at least biosignal streams, acoustic feature streams, and visual interaction streams; Step 2, calculation of the repositioning score Q_return using a multimodal repositioning evaluation algorithm, the calculation formula being: Q_return = alpha × Norm(S_bio) + beta × Norm(S_audio) + gamma × Sync(S_visual) + delta × Decay(S_time). Wherein, Norm(S_bio) is the normalized smoothing index of the biosignal flow, derived from the actual physical electrical signals of the sensor, including heart rate variability (HRV) or skin conductance response (GSR) signals; Norm(S_audio) is the convergence index of the acoustic feature flow; Sync(S_visual) is the synchronization rate of the interaction between user actions and device feedback; Decay(S_time) is the time decay factor; alpha, beta, gamma, and delta are weighting coefficients, and satisfy alpha + beta + gamma + delta = 1. In step 3, when Q_return exceeds the preset safe return threshold Q_safe, the system determines that the emotion return is successful and terminates the intervention procedure. Preferably, the value of the biosignal flow weighting coefficient alpha is forcibly specified to be no less than 0.50.
[0013] The beneficial effects of this invention are as follows:
[0014] 1. Established an industrial-grade interface: standardized emotion intervention commands, supported interoperability across vendor devices, and solved the data silo problem.
[0015] 2. Certainty of scientific verification: Transforming cultural archetypes into physical parameters that can be verified by neuroscience, and combining them with the psychological theory of "self-distance" to force third-person narration, ensuring the consistency and effectiveness of intervention methods.
[0016] 3. High-precision closed-loop assessment: By introducing high-weight real bioelectrical signals (HRV / GSR) and visual synchronization rate, the misjudgment rate of a single sensor in extreme emotional scenarios is significantly reduced, and the degree of physiological repositioning is objectively reflected. Attached Figure Description
[0017] Figure 1 is a schematic diagram of the data structure of the Emotional Computing Interface Protocol (ESDS) in an embodiment of the present invention.
[0018] Figure 2 is a diagram of the lookup table (LUT) structure of emotion prototypes and neural mapping parameters (visual / auditory) in an embodiment of the present invention.
[0019] Figure 3 is a flowchart of the multimodal collaborative emotion repositioning evaluation algorithm in an embodiment of the present invention.
[0020] Figure 4 is a diagram of a distributed cloud / terminal collaborative network architecture based on this protocol in an embodiment of the present invention. Detailed Implementation
[0021] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0022] Example 1: Engineering Implementation of the ESDS Data Protocol
[0023] As shown in Figure 1, the system defines a standardized ESDS data packet containing three strongly bound fields: Role_ID (e.g., "PROTO_ANGRY_HIGH"), Action_Code (e.g., "ARM_PUSH_FAST"), and Narrative_Tag (forced third person).
[0024] When a child is detected to be in a hyperarousal, angry state, the upper-layer AI model generates this ESDS packet. The downstream hardware driver layer parses the data packet and directly calls the parameters in the local LUT: controlling the LED to flash orange-red light at a frequency of 4Hz (within the range of 3000K-4500K), and the robotic arm to perform a pushing action.
[0025] Meanwhile, based on the constraints of Narrative_Tag and the theory of "Self-distancing", the voice module forces the generation of third-person declarative sentences such as "Look, that red fire dragon has run out!" and strictly prohibits the generation of second-person sentences such as "You are very angry". This helps users objectify their emotions at the cognitive level and reduce psychological defenses.
[0026] Example 2: Standardized Application of Neural Mapping Parameters
[0027] As shown in Figure 2, for low-arousal fear prototypes (such as the "fearful mouse"), the protocol stipulates that their visual output must be: soft blue light with a color temperature of 6000K, brightness variation following a sine wave pattern, and frequency strictly controlled at 0.2Hz (i.e., one cycle every 5 seconds) to match the frequency of calm human breathing. Any device conforming to this protocol must use parameters within this frequency range when processing such emotions to ensure consistent neuro-soothing effects. For high-arousal negative prototypes, a sharp flashing of 3-5Hz is strictly limited to quickly interrupt emotional rumination using the neural entrainment effect.
[0028] Example 3: Multimodal Cooperative Relocation Evaluation Algorithm
[0029] As shown in Figure 3, the system calculates the repositioning score Q_return in real time. The calculation formula is as follows:
[0030] Q_return = alpha × Norm(S_bio) + beta × Norm(S_audio) + gamma ×Sync(S_visual) + delta × Decay(S_time)
[0031] Each parameter has a clearly defined physical signal source:
[0032] (1) Norm(S_bio): Normalized score of biosignal flow. This signal originates from the actual physical electrical signals of the wearable sensor, including the RMSSD index of heart rate variability (HRV) acquired by photoplethysmography (PPG) or the micro-Siemens (μS) conductance change value of the skin conductance response (GSR) sensor. This parameter directly reflects the physiological calming degree of the user's autonomic nervous system.
[0033] (2) Norm(S_audio): Normalized score of acoustic feature flow, which is derived from the sound wave signal collected by the microphone and calculated based on the slope of the speech fundamental frequency (F0) and the energy convergence.
[0034] (3) Sync(S_visual): Visual interaction synchronization rate, calculated by the phase matching degree between the rhythm of the user's body movements captured by the camera and the rhythm of the device's light feedback.
[0035] (4) Decay(S_time): Time decay factor, which represents the inverse normalized value of the time interval from the start of the intervention to the first occurrence of positive semantics.
[0036] (5) Weighting coefficients: alpha, beta, gamma, delta satisfy alpha + beta + gamma + delta = 1. The key constraint in this embodiment stipulates that the biosignal weight alpha ≥ 0.5. This constraint ensures that the evaluation result is mainly determined by the user's real physiological electrical signals, avoiding misjudgment caused by feigning calmness based solely on facial expressions or voice.
[0037] When Q_return exceeds the preset safety threshold Q_safe (e.g., 0.75) for N consecutive sampling periods (e.g., N=5, sampling frequency 1Hz), the system determines that the emotion has been successfully restored, automatically terminates the intervention, and switches to the daily mode.
[0038] Example 4: Building an Ecosystem for Standard Essential Patents (SEPs)
[0039] As shown in Figure 4, in a distributed emotion computing network, the cloud master node periodically sends updated Role_ID attribute parameter packets (such as fine-tuning flashing frequency or action trajectory). Slave nodes must follow the protocol parsing and synchronously update their local LUTs. Any data packets that do not conform to the ESDS format will be rejected at the network gateway layer, thereby establishing the uniqueness and exclusivity of the protocol at the system architecture level and ensuring the consistency of the intervention strategy across the entire network.
[0040] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An emotion computing interface protocol based on emotion archetypes, characterized in that, A standardized Emotional State Data Structure (ESDS) is defined, which includes the following strongly bound fields: Role_ID: A unique identifier selected from a pre-defined finite set containing at least six basic emotional archetypes that correspond to different combinations of valence and arousal. Action_Code: A uniquely mapped embodied action-driven feature code that is used to directly control the actuator of the terminal device to execute a specific embodied action sequence; Narrative_Tag: A narrative logic constraint parameter uniquely mapped to the Role_ID, used to enforce the person's perspective, tense, and objectified content boundaries of the natural language generation model, in order to achieve emotion objectification based on the theory of self-distancing; Specifically, any system based on this protocol must encapsulate the unstructured emotion recognition results into the ESDS format when generating emotion intervention instructions, and the combination of the three fields must maintain indivisible atomicity within a single intervention cycle.
2. The protocol according to claim 1, characterized in that, It also includes database annotation and model training standards: When constructing the emotion computing training dataset, it is mandatory to map and label the emotion descriptions in the original corpus as Role_ID in the ESDS protocol, and it is prohibited to use ambiguous natural language labels as the final target value for model training. The training dataset formed in this way is used to train or fine-tune a large language model, enabling the model to establish an end-to-end logical mapping path from "natural language input" directly to "structured ESDS sequence".
3. A physical layer output standard based on neural mapping parameters, applied to the protocol described in claim 1, characterized in that: For a specific Role_ID, the protocol defines a range of standard physical parameters in the device driver lookup table (LUT); For the high-arousal negative emotion prototype, the flicker frequency range of the visual output characteristics is specified to be 3-5Hz, the color temperature range is 3000K-4500K, and the dynamic waveform is an acute-angle geometric shape. For low-arousal negative emotion prototypes, the flicker frequency range of visual output features is specified to be 0.1-0.5Hz, the color temperature range is 5000K-6500K, and the dynamic waveform is a smooth sine wave. The physical parameters serve as the neural activation signature of the emotional prototype, forcing downstream actuators to call values within the corresponding parameter range to output sensory feedback.
4. A multimodal collaborative emotion attribution assessment method, characterized in that, Includes the following steps: Step 1: Real-time acquisition of multimodal signal streams from the user during the intervention process, wherein the signal streams include at least biosignal streams, acoustic feature streams, and visual interaction streams; Step 2: Calculate the relocation score Q_return using the multimodal relocation evaluation algorithm. The calculation formula is as follows: Q_return = alpha × Norm(S_bio) + beta × Norm(S_audio) + gamma × Sync(S_visual) + delta × Decay(S_time) in: Norm(S_bio) is the normalized flattening index of the biosignal flow, which originates from the actual physical electrical signals of the sensor, including heart rate variability (HRV) or skin conductance response (GSR) signals. Norm(S_audio) is the convergence index of the acoustic characteristic flow; Sync(S_visual) represents the synchronization rate between user actions and device feedback. Decay(S_time) is the time decay factor; alpha, beta, gamma, delta are weighting coefficients, and satisfy alpha + beta + gamma + delta = 1; Step 3: When Q_return exceeds the preset safe return threshold Q_safe, the system determines that the emotion has been successfully returned and terminates the intervention procedure.
5. The method according to claim 4, characterized in that: The value of the biosignal flow weighting coefficient alpha is mandated to be no less than 0.50, and it serves as the core weight for determining emotional calmness, ensuring that the evaluation results primarily rely on the user's actual physiological electrical signals.
6. A standard essential patent implementation method based on the protocol of claim 1, characterized in that: In a distributed emotion computing network, the master node periodically sends out updated Role_ID attribute parameter packages, including motion trajectory parameters for fine-tuning Action_Code or frequency range of physical layer output standards; The slave node follows the protocol to parse and synchronously update the local lookup table (LUT); Any data packet that does not conform to the ESDS format is logically rejected at the network gateway layer to ensure the consistency of the intervention strategy across the entire network.
7. A child emotional intelligence guidance system based on the objectification and externalization of emotional instances, characterized in that, include: A protocol parsing module is used to receive and parse data packets conforming to the ESDS format described in claim 1; The hardware driver mapping module is used to call the local lookup table (LUT) to drive the hardware actuator according to the fields in the data packet, so as to implement the physical layer output standard as described in claim 3; The repositioning evaluation engine is used to execute the multimodal collaborative emotion repositioning evaluation algorithm as described in claims 4 to 5, and to control the flow of the intervention process based on the evaluation results; The cloud-based collaboration module is used to execute the standard essential patent implementation method as described in claim 6.
8. An electronic device comprising a processor, a memory, and a computer program stored in the memory, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 6.
10. A cross-device interoperability method based on the protocol of any one of claims 1 to 3, characterized in that, Includes the following steps: The first device generates an emotion intervention instruction that conforms to the ESDS data structure; The second device receives the instruction and, based on the Action_Code field therein, calls the local lookup table (LUT) to map it into a control signal that conforms to the physical layer output standard; The third device collects multimodal signals and executes the multimodal collaborative emotion repositioning evaluation method, and feeds back the generated Q_return score as a unified acceptance index for the intervention effect to the first device; By standardizing the ESDS data structure, physical layer output standards, and homing evaluation methods, closed-loop interoperability of emotion intervention between devices from different manufacturers can be achieved.