Non-semantic micro-probing confirmation method and system based on separability of candidate emotions

By using separability judgment based on candidate emotion pairs and non-semantic micro-probing actions, the accuracy and resource consumption problems of companion terminals in multi-candidate emotion scenarios are solved, achieving efficient emotion confirmation with low disturbance and low resource consumption. It is applicable to desktop companion robots, children's companion devices and elderly companion terminals.

CN122493889APending Publication Date: 2026-07-31SUZHOU ZHIMENG ERA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU ZHIMENG ERA TECHNOLOGY CO LTD
Filing Date
2026-06-12
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

When faced with multiple candidate emotions that are similar and have unclear boundaries, existing companion terminals struggle to accurately identify emotions without significantly increasing disturbance, relying on high-cost heavy computing, and meeting privacy and power consumption constraints. This is especially true in low-resource independent terminals, where existing solutions often lead to increased power consumption or interactive disturbances, affecting the companion experience.

Method used

A separability judgment mechanism based on candidate emotion pairs is introduced. Non-semantic micro-probing actions such as light, slight displacement, and vibration are used for low-disturbance confirmation. Combined with differential response features and hierarchical confirmation strategy, the problem of distinguishing similar candidate emotions is prioritized, reducing resource consumption and improving confirmation accuracy.

Benefits of technology

In scenarios with uncertain emotions, this approach improves the accuracy and stability of emotion confirmation through low-intrusion and low-resource consumption. It is suitable for low-resource independent terminals, maintains the naturalness and privacy of interaction, and reduces resource consumption and interaction disturbance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493889A_ABST
    Figure CN122493889A_ABST
Patent Text Reader

Abstract

This invention discloses a non-semantic micro-probing confirmation method and system based on the separability of candidate emotion pairs. The method includes the following steps: collecting user input signals, performing lightweight emotion assessment, selecting candidate emotion pairs, calculating the separability of the candidate emotion pairs, determining the separability of the candidate emotion pairs and emotion confirmation, selecting and executing corresponding non-semantic micro-probing actions, collecting user responses within a short observation window after the micro-probing action is executed, matching response templates to obtain the matching degree of each candidate emotion, performing targeted correction on the weights of each candidate emotion within the candidate emotion pair, and confirming the emotion state based on the corrected candidate emotion values, or implementing escalation, delay, or rollback strategies. This invention, by first extracting candidate emotion pairs and then combining them with separability to determine whether further confirmation is needed, can more effectively handle the differentiation problem between similar candidate emotions, thereby improving the confirmation accuracy and result stability in uncertain emotion scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of affective computing, intelligent human-computer interaction, companion robots, and multimodal signal processing, and particularly to a non-semantic micro-probing confirmation method and system based on the separability of candidate emotion pairs. It is applicable to desktop companion robots, children's companion devices, elderly care companion terminals, and other intelligent devices that need to complete user emotion perception and companionship feedback under local conditions. Background Technology

[0002] With the increasing application of companion-type smart terminals, these devices increasingly need the ability to recognize and respond to users' current emotional states. Current technologies typically identify user emotions through voice, image, posture, touch, distance changes, or multimodal fusion, and then control responses, actions, lighting feedback, or companionship strategies based on the recognition results. These solutions achieve good recognition results when user emotional characteristics are relatively obvious. However, in actual use, user emotions are often not stable or clearly defined, especially between states such as depression, irritability, calmness, fatigue, tension, and surprise, where multiple candidate emotion scores are often close and difficult to distinguish directly.

[0003] In this context, if the system directly outputs an emotional state based solely on a single passive sampling result, misjudgments are likely to occur, affecting the relevance and naturalness of subsequent companionship feedback. If the system continues to collect more data over a longer period or increases the sampling intensity, it will incur additional power consumption, latency, and resource consumption, which is detrimental to the long-term operation of low-resource independent companionship terminals. If the system directly questions the user using natural language, it is likely to significantly disturb the user, reduce the interactive experience, and place higher demands on the terminal's semantic understanding and dialogue capabilities. Therefore, when the recognition result is not clear enough, how to further confirm similar candidate emotions without increasing obvious disturbance, without relying on high-cost heavy computation, and while satisfying privacy and power consumption constraints as much as possible, is a problem that existing companionship terminal emotion recognition technologies need to solve.

[0004] Furthermore, most existing solutions focus on improving the accuracy of emotion recognition or on executing subsequent interactions based on the recognition results, while few establish a complete technical loop for "low-intrusion confirmation when multiple candidate emotions are close together." Especially in scenarios such as nighttime, quiet, privacy-sensitive, or resource-constrained environments, terminals often cannot frequently utilize high-power sensors, nor are they suitable for clarification and confirmation through explicit question-and-answer methods. This limits the usability of existing technologies in real-world companionship scenarios.

[0005] Based on the above, a new technical solution is needed. After initially identifying multiple candidate emotions, the system should not perform a general re-identification of all emotion categories. Instead, it should prioritize candidate emotion pairs that are close to each other, determining whether further confirmation is necessary based on the separability of these pairs. When the separability is insufficient, it should guide the user to generate a short-term response through low-semantic, short-duration, and weak-stimulus non-semantic micro-probing actions. This response change should then be used to directionally correct the candidate emotions, thereby completing emotion confirmation with lower resource consumption and less interactive disruption. This approach is more suitable for engineering implementation of independent companion terminals in real-world usage scenarios and also helps improve the stability and reliability of emotion confirmation results.

[0006] The existing technologies related to this application can be broadly classified into two categories:

[0007] The first type involves directly executing interactive feedback after emotion recognition:

[0008] These solutions typically identify user emotions through information such as voice, images, facial expressions, and gestures, and then control the device to output corresponding responses, actions, facial expressions, or light feedback based on the recognition results. The key is "how to execute a response after recognizing an emotion," making it a fairly typical "emotion recognition + interactive control" model.

[0009] The second category involves using multimodal information or interaction data to help determine user emotions:

[0010] In addition to utilizing conventional inputs such as voice and images, this type of solution also incorporates data on user touch behavior, gaze direction, distance changes, and interaction frequency to comprehensively analyze user emotions and improve recognition accuracy. The core idea behind this type of solution is to enhance the ability to judge the user's emotional state by increasing the dimensions of observed information or integrating more interaction features.

[0011] Most existing solutions focus on "how to identify emotions" or "how to provide feedback after identification." When identification is uncertain, the common approach is to continue passively collecting more data, which leads to increased power consumption and resource usage, making it unsuitable for low-resource independent companion terminals. Another common approach is to clarify through natural language questions, but natural language clarification is highly disruptive and relies heavily on semantic understanding capabilities.

[0012] This invention targets independent companion devices such as Doraemon, desktop robots, companion ornaments, and interactive toys. These devices typically have limited main processor computing power, battery capacity, or power consumption budget. Prolonged use of cameras, microphones, or high-frequency inference increases resource consumption, and frequent questioning of the user via natural language can be disruptive and negatively impact the companionship experience. Furthermore, they are limited by privacy requirements and permissible stimulation intensity in scenarios such as homes, bedrooms, and offices. Therefore, this invention does not primarily rely on "continued passive identification" or "direct natural language clarification." Instead, it introduces a low-intrusion proactive confirmation mechanism for uncertain emotional scenarios. This mechanism uses candidate emotion pairs as the confirmation object, separability as the trigger condition, non-semantic micro-probing as the probing method, short-time differential response as the judgment criterion, and a tiered confirmation and fallback strategy as a fallback path. Summary of the Invention

[0013] The purpose of this invention is to provide a non-semantic micro-probing confirmation method and system based on the separability of candidate emotion pairs. It introduces a low-disturbance proactive confirmation mechanism for uncertain emotion scenarios, using candidate emotion pairs as the confirmation object, separability as the trigger condition, non-semantic micro-probing as the probing means, short-time differential response as the judgment basis, and hierarchical confirmation and backoff strategies as fallback paths.

[0014] The technical solution of this invention is:

[0015] A non-semantic micro-trial confirmation method based on the separability of candidate sentiment pairs includes the following steps:

[0016] S1. Acquire the user's multimodal input signals;

[0017] S2. Perform a lightweight emotion assessment on the multimodal input signal to obtain multiple candidate emotion values;

[0018] S3. Select at least two candidate emotions that rank highest to form a candidate emotion pair;

[0019] S4. Calculate the separability of the candidate emotion pairs, whereby the separability is used to characterize the degree of distinction between the candidate emotion pairs;

[0020] S5. Determine whether the separability has reached the preset confirmation threshold: if it has, directly output the confirmed emotional state; if it has not, enter the micro-probing confirmation mode.

[0021] S6. Based on the candidate emotion pairs and the current scene constraints, select and execute the corresponding non-semantic micro-probing action. The non-semantic micro-probing action is a low-semantic, short-term, weak-stimulus non-questioning action.

[0022] S7. Collect user response within a short observation window after the micro-probing action is executed, and extract differential response features before and after the micro-probing.

[0023] S8. Match the differential response features with the predefined response templates of the corresponding candidate emotion pairs to obtain the matching degree of each candidate emotion.

[0024] S9. Based on the matching degree, the weights of each candidate emotion in the candidate emotion pair are adjusted in a targeted manner to obtain the adjusted candidate emotion value.

[0025] S10. Based on the modified candidate emotion value, determine whether the confirmation condition is met. If it is met, output the confirmed emotion state. If it is not met, execute the upgrade, delay or rollback strategy.

[0026] Preferably, in S4, the separability of candidate emotion pairs is calculated using the following formula:

[0027] ;

[0028] in, For the current candidate sentiment pairs, , For the two emotion categories in the candidate emotion pair, , These are the initial candidate scores for both. The difference in scores between the two. For sorting stability within a continuous time window, To support consistency across multiple modalities, For the current sampling quality, The prior constraints of historical confirmation results on the current candidate pair are α, β, γ, δ, and η, which are preset weight coefficients.

[0029] Preferably, in S6, the selection process for the non-semantic micro-probing action includes:

[0030] Predefine a priority table of trial actions for different candidate emotions, as well as the power consumption cost, disturbance level, and privacy risk level of the actions;

[0031] Based on the current scenario's power consumption budget, privacy level, permissible stimulus intensity, remaining device battery power, resource usage level, and action cooldown time constraints, a set of candidate actions that meet the requirements is selected;

[0032] Each candidate action is scored based on the following action selection scoring formula, and the action with the highest score is selected for execution:

[0033] ;

[0034] in, For the j-th candidate action, The degree to which this action distinguishes and fits the current candidate emotion pairs. For scene permission, For the sake of power consumption, In exchange for the disturbance, For the cost of privacy risks, , , , , These are the preset weighting coefficients.

[0035] Preferably, in S7, the extraction of the differential response features before and after the micro-probe specifically includes:

[0036] Extracting feature vectors before micro-probe execution and the feature vector within a short observation window after microprobe execution The features include at least one or more of the following: changes in gaze direction, changes in device distance, changes in touch probability, changes in head rotation, changes in avoidance actions, and changes in voice activity.

[0037] Calculate the differential response characteristics :

[0038] .

[0039] Preferably, in S9, the weights of each candidate emotion within a candidate emotion pair are adjusted in a targeted manner using the following formula:

[0040] ;

[0041] in, To correct candidate sentiment The weight, Given its initial candidate score, denoted as , where is the matching degree between the differential response features and the corresponding response template of the candidate emotion, and μ is the correction strength coefficient of the matching result. This represents the current candidate sentiment pair.

[0042] Preferably, in S10, determining whether the confirmation condition is met based on the corrected candidate emotion value specifically includes:

[0043] Calculate the overall confirmation score for each candidate emotion:

[0044] ;

[0045] Wherein, ρ1, ρ2, and ρ3 are preset weight coefficients;

[0046] If the overall confirmation score of a candidate emotion is greater than the preset confirmation threshold, then the candidate emotion is output as a confirmed emotion state; otherwise, an escalation, delay, or rollback strategy is executed.

[0047] Preferably, the upgrade, delay, or rollback strategy specifically includes:

[0048] If the confirmation conditions are not met and the current scenario allows for continued probing, the micro-probing action will be upgraded to a higher level of stimulation and the confirmation process will be repeated.

[0049] If the current scenario does not allow for immediate continuation of the trial, the confirmation process will be delayed until the next interactive window.

[0050] If the confirmation threshold is not reached after multiple consecutive confirmations, or if the preset maximum number of attempts is reached, the output will revert to a neutral state, a conservative state, or a pending confirmation state, and no new micro-probing actions will be performed.

[0051] Set the maximum number of attempts within the same confirmation cycle, and the minimum cooldown interval between two adjacent micro-probe actions to avoid continuous stimulation and cumulative disturbance.

[0052] Preferably, the non-semantic micro-probing actions include at least one of the following: changes in light intensity / color, slight device displacement, orientation rotation, short vibrations, and low-brightness visual cues, but do not include natural language inquiry actions.

[0053] A non-semantic micro-trial confirmation system based on the separability of candidate emotion pairs includes:

[0054] The signal acquisition module is used to acquire the user's multimodal input signals;

[0055] A lightweight evaluation module is used to process the multimodal input signal and output multiple candidate emotions and their initial candidate scores;

[0056] The candidate pair construction module is used to select at least two top-ranked candidate emotions to form a candidate emotion pair.

[0057] A separability calculation module is used to calculate the separability of the candidate emotion pairs;

[0058] The trigger judgment module is used to compare the separability with a preset threshold to determine whether to enter the micro-probe confirmation mode;

[0059] The action selection and execution module is used to select and execute non-semantic micro-probing actions based on candidate emotion pairs and scene constraints.

[0060] The response acquisition and feature extraction module is used to acquire user responses within a short observation window after the micro-probing action is executed, and to extract differential response features.

[0061] The template matching module is used to match the differential response features with the corresponding response templates to obtain the matching degree of each candidate emotion;

[0062] The weight correction module is used to make targeted corrections to the candidate emotion weights within candidate emotion pairs based on the matching degree.

[0063] The confirmation decision module is used to output confirmation results based on the corrected candidate sentiment values, or to invoke upgrade, delay, or rollback strategies when the target is not met.

[0064] Preferably, the system is deployed in a desktop companion robot, a child companion device, an elderly companion terminal, or a low-power intelligent interactive device, and all processing flows are completed locally.

[0065] Compared with the prior art, the present invention has at least the following advantages:

[0066] 1. More suitable for handling uncertain emotional scenarios

[0067] This invention is not limited to conventional scenarios where emotional features are clear and directly identifiable, but rather focuses on complex scenarios where multiple candidate emotions are similar, have unclear boundaries, and are difficult to distinguish directly. By first extracting candidate emotion pairs and then combining this with separability to determine whether further confirmation is needed, the system can more effectively address the distinction between similar candidate emotions, thereby improving the accuracy and stability of confirmation in uncertain emotional scenarios.

[0068] 2. It can reduce resource consumption and is more suitable for low-resource independent companion terminals.

[0069] This invention triggers active confirmation only when the separability of candidate emotion pairs is below a preset threshold, avoiding continuous high-frequency sampling, repeated high-cost recognition processes, or long-term use of high-power sensors. Furthermore, this invention emphasizes local closed-loop processing, short observation windows, a limited library of micro-trial actions, and a tiered confirmation mechanism, making it more suitable for deployment in low-resource, independent devices such as desktop companion robots, children's companion devices, and elderly care companion terminals. Its engineering implementation path is clearer, and its practical applicability is stronger.

[0070] 3. It can reduce interactive disturbances.

[0071] When the recognition results are unclear, this invention does not rely on natural language to directly ask the user for clarification. Instead, it uses non-semantic micro-probing actions such as light, slight displacement, vibration, and changes in orientation to provide low-intrusion confirmation. This method can obtain further distinguishing information without significantly interrupting the user's current state, thus better maintaining the naturalness, continuity, and comfort of the companionship interaction.

[0072] 4. It has better privacy protection.

[0073] In privacy-sensitive scenarios, this invention can prioritize non-voice, non-camera-enhanced methods such as changes in lighting, subtle movements, and vibration alerts to complete confirmation, without frequently resorting to voice or continuous visual data collection. This reduces reliance on user voice content, image information, or continuous behavioral data while meeting confirmation requirements, better aligning with the privacy protection requirements of companion devices in bedrooms, rest areas, and private spaces.

[0074] 5. The logic is more robust.

[0075] When a single micro-probing confirmation is insufficient to reach the confirmation threshold, this invention does not rush to output a highly certain emotional conclusion. Instead, it continues to implement processing strategies such as escalating confirmation, delaying confirmation, or reverting to a conservative state based on the current scenario conditions. Through this tiered confirmation and revert mechanism, the system can avoid making premature and aggressive judgments when information is insufficient, thereby improving the stability, controllability, and fault tolerance of the overall confirmation process. Attached Figure Description

[0076] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0077] Figure 1 This is an overall flowchart of the non-semantic micro-probing confirmation method of the present invention;

[0078] Figure 2 A flowchart for upgrade, delay, and rollback mechanisms. Detailed Implementation

[0079] like Figure 1 As shown, the non-semantic micro-trial confirmation system method based on the separability of candidate emotion pairs of the present invention includes the following 12 main steps in its overall process.

[0080] 1. Acquire multimodal input signals

[0081] The system collects user input signals, including voice, image, distance, touch, or behavioral signals. These input signals may include one or more of the following: voice, facial image, head posture, distance changes, touch behavior, interaction frequency, and changes in device relative orientation. For ease of subsequent processing, the multimodal input collected at the current moment can be denoted as:

[0082] ;

[0083] in, Indicates visually relevant input, This indicates voice-related input. Indicates distance-related input, Indicates touch-related input. This indicates behaviorally relevant input. The purpose of this step is to provide a unified input basis for subsequent candidate emotion generation, rather than making final confirmation directly at the input layer.

[0084] 2. Light Emotional Assessment

[0085] Lightweight feature extraction is performed on the input signal to obtain multiple candidate emotion values. This step preferably outputs at least the top two candidate emotions along with their scores, ranking stability, and modality support information, rather than outputting only a single emotion category at once. A lightweight evaluation module evaluates the input... Output candidate sentiment score vectors:

[0086] ;

[0087] in, Indicates the first One emotion category, This represents the initial candidate score for the emotion at the current moment, and satisfies... First, multiple candidates are presented, and then it is determined whether confirmation is needed. This is a prerequisite for subsequent targeted confirmation of candidates.

[0088] 3. Forming candidate pairs

[0089] Select at least two top-ranked candidate emotions to form a candidate emotion pair. This step narrows the problem down to distinguishing between similar candidates, avoiding the costly re-determination of all emotion categories in uncertain scenarios. Alternatively, if ranked by score, select the top two candidate emotions... and The current candidate emotion pair can then be denoted as:

[0090] ;

[0091] The problem of re-identifying the entire category space is transformed into a local confirmation problem focusing only on candidate emotion pairs, thereby reducing the complexity of confirmation and enhancing the targeting of subsequent action selection.

[0092] 4. Determine the separability of candidate emotion pairs

[0093] Calculate the separability of candidate sentiment pairs. Separability calculation is based on at least one of the following: the difference in candidate sentiment scores, ranking stability across multiple consecutive time windows, multimodal support consistency, current sampling quality or confidence level, and prior constraints imposed by historical confirmation results on the current candidate pair. A separability threshold is used to determine whether a candidate sentiment pair is sufficiently explicit. Separability can be expressed as:

[0094] ;

[0095] in, Expressing emotions Candidate scores, Expressing emotions Candidate scores, Indicates the stability of the candidate ranking. This indicates that multimodal support supports consistency. Indicates the current sampling quality. This indicates a historical confirmation constraint. These are the weighting coefficients. The focus of this step is not simply comparing two scores, but rather determining whether the current candidate pair is sufficiently distinguishable.

[0096] 5. Determine if emotional confirmation is possible.

[0097] The system determines whether the separability reaches the confirmation threshold. If the separability is higher than the threshold, it means that the current candidate emotion pair can be distinguished, and the system directly outputs the confirmation result; if it is lower than the threshold, it enters the micro-probing confirmation mode. Active confirmation is only triggered when the separability is insufficient, thereby avoiding the system entering additional probing after each recognition, reducing resource consumption and interference with the user.

[0098] 6. Selection of micro-probing actions

[0099] The target non-semantic micro-probing action is selected based on candidate emotion pairs and scene constraints. This selection is simultaneously constrained by the current candidate emotion pair, power budget, privacy level, disturbance level, permitted stimulus types in the current scene, remaining device battery or resource usage level, and the cooldown time between the most recent probing action and the current moment. Micro-probing actions can be preset to multiple levels, such as Level 1, Level 2, and Level 3, corresponding to different stimulus intensities, power costs, and disturbance levels. Different candidate emotion pairs are associated with different probing action priority tables. For example, the "depressed-irritable" candidate pair prioritizes weak light pulses, slight displacement movements, or short vibrations, while the "tense-surprised" candidate pair prioritizes short-term orientation changes and low-brightness visual cues. To reflect the constrained selection nature of this step, the action selection score can be recorded as:

[0100] ;

[0101] in, Indicates action On candidate sentiment Distinguishing adaptability Indicates the scene permission level. Indicates the cost of power consumption. This indicates the cost of disturbing them. This indicates the cost of privacy risks. An action is only included in the candidate action set if it simultaneously satisfies the constraints of scene permission, cooldown time, and maximum number of attempts. The selected action is not a general cue action, but rather a targeted, non-semantic micro-probing action specific to the candidate emotion pair.

[0102] 7. Perform micro-probe actions

[0103] Perform the selected non-semantic micro-probing action. The action should be a low-semantic, short-duration, and weakly stimulating light movement to avoid interrupting the user. Micro-probing actions can be categorized into Level 1, Level 2, and Level 3 actions based on stimulus intensity. Level 1 actions are prioritized for initial confirmation, while Level 2 and Level 3 actions are used for escalation confirmation. Essentially, this step outputs a short external stimulus to the user, causing different emotional states to produce distinguishable micro-response differences within a short observation window. This step does not involve verbal questioning of the user, but rather elicits a response more suitable for subsequent judgment through non-semantic means.

[0104] 8. Collect short-time responses

[0105] User responses are captured within a short observation window following the micro-probing action, and short-term differential response features before and after the micro-probing are further extracted. The observation window length is selected to be between 0.5 and 3 seconds to capture the immediate incremental response after the probing under low-interference conditions; differential response features may include, but are not limited to, changes in gaze direction, changes in device distance, changes in touch probability, changes in head rotation, changes in avoidance actions, and changes in voice activity. Let the feature vector before the micro-probing be... The feature vector after micro-probing is Then the difference response characteristics can be expressed as:

[0106] ;

[0107] Instead of directly using the absolute response value, we use the short-time differential response before and after the micro-probe for confirmation, thereby better eliminating the interference caused by differences in individual baseline states.

[0108] 9. Response template matching

[0109] The differential response features are matched with corresponding response templates to determine which candidate emotion the response is closer to. Response templates are organized jointly according to candidate emotion pairs, trial action types, and scene constraints to improve the targeting of the matching. A matching degree threshold is used to determine whether the matching strength between the differential response and a candidate emotion template is sufficient to support weight adjustment. Let candidate emotions be... In action The response template below is The matching degree can then be expressed as:

[0110] ;

[0111] Here, sim() can be a relevance measure, distance similarity measure, or normalized matching function. The template is not statically built based on a single emotion, but rather organized in a targeted manner by combining candidate emotion pairs and trial actions.

[0112] 10. Adjust candidate weights

[0113] The weights of each candidate emotion in the current candidate emotion pair are adjusted based on the matching results to obtain the adjusted candidate emotion value. The adjustment only modifies the weights of the two candidate emotions within the current candidate emotion pair, rather than restarting a complete emotion recognition process. Candidate weight update:

[0114] ;

[0115] in, This represents the strength coefficient of the correction applied to the candidate weights by the matching result. For candidate emotions The matching degree. This step only performs targeted corrections within the candidate pairs, rather than recalculating across all categories, which makes the confirmation path shorter and more suitable for resource-constrained terminals.

[0116] 11. Emotional Confirmation

[0117] The system determines whether the corrected candidate emotion value reaches the confirmation threshold. If it does, it outputs a confirmed emotion state and sends this state to the subsequent companionship feedback control module, facial expression control module, lighting effect control module, or dialogue strategy module. The overall confirmation score is expressed as follows:

[0118] ;

[0119] when When the value exceeds the threshold, the corresponding confirmed emotional state is output. This step does not draw a conclusion based solely on a single template matching result, but rather incorporates the corrected candidate weights, matching strength, and original separability into the final confirmation.

[0120] 12. Upgrade, Delay, or Rollback Path

[0121] like Figure 2 As shown, if the confirmation threshold is still not reached after correction, at least one of the following strategies will be executed: upgrade to a higher level of micro-probing action; delay confirmation until the next interaction window; revert to a neutral or conservative output; or stop further stimulation after reaching the maximum number of probing attempts to avoid continuous interference to the user.

[0122] When the initial micro-probing attempt fails, if the current power budget, privacy level, and disturbance level allow, the attempt is escalated to a higher-level micro-probing action. If the current scenario does not allow for continued stimulation, the attempt is delayed until the next interaction window. If multiple consecutive attempts fail to reach the threshold, the system reverts to a neutral, pending confirmation, or conservative state. After reverting, the candidate emotion pairs, probing actions, and matching results of this round are preferably recorded for use in subsequent time windows. If confirmation is not completed, a conservative, pending confirmation, or delayed decision state is output to ensure stable and controllable system behavior. The maximum number of attempts parameter constrains the upper limit of the number of micro-probing attempts allowed within a confirmation cycle, and the cooling interval parameter constrains the minimum time interval between two micro-probing actions to prevent continuous stimulation from accumulating disturbance. Instead of immediately outputting an aggressive conclusion after a confirmation failure, the system maintains stability and controllability through escalation, delay, and reverting mechanisms.

[0123] This invention also proposes a non-semantic micro-trial confirmation system based on the separability of candidate emotion pairs, comprising:

[0124] The signal acquisition module is used to acquire the user's multimodal input signals;

[0125] A lightweight evaluation module is used to process the multimodal input signal and output multiple candidate emotions and their initial candidate scores;

[0126] The candidate pair construction module is used to select at least two top-ranked candidate emotions to form a candidate emotion pair.

[0127] A separability calculation module is used to calculate the separability of the candidate emotion pairs;

[0128] The trigger judgment module is used to compare the separability with a preset threshold to determine whether to enter the micro-probe confirmation mode;

[0129] The action selection and execution module is used to select and execute non-semantic micro-probing actions based on candidate emotion pairs and scene constraints.

[0130] The response acquisition and feature extraction module is used to acquire user responses within a short observation window after the micro-probing action is executed, and to extract differential response features.

[0131] The template matching module is used to match the differential response features with the corresponding response templates to obtain the matching degree of each candidate emotion;

[0132] The weight correction module is used to make targeted corrections to the candidate emotion weights within candidate emotion pairs based on the matching degree.

[0133] The confirmation decision module is used to output confirmation results based on the corrected candidate sentiment values, or to invoke upgrade, delay, or rollback strategies when the target is not met.

[0134] The system is deployed in desktop companion robots, children's companion devices, elderly companion terminals, or low-power intelligent interactive devices, and all processing flows are completed locally.

[0135] Example

[0136] A desktop companion robot is designed for use in a bedroom at night, where only low-light lighting and minimal movement are permitted; voice prompts and continuous video recording are not allowed. The system obtains candidate emotion pairs "depressed" and "irritable" within two consecutive time windows, with initial weights of [weights to be filled in]. and Stability is Modal consistency is The current sampling quality item is taken Historical confirmation constraints .

[0137] In taking , , , , When the separability is Because the value is below the threshold Instead of directly outputting the result, the system enters a micro-probing confirmation mode.

[0138] For the candidate emotion pair "depressed-irritable", the system calculates action selection scores for two optional primary actions, "weak light pulse" and "slight displacement action", in the current scene. Let the action selection weight be... , , , , For a "weak light pulse", let its candidate pair fit be... Scene Permission Power consumption cost The cost of disturbance Privacy risks and costs ,but For "weak displacement motion", let its candidate pair fit degree be... Scene Permission Power consumption cost The cost of disturbance Privacy risks and costs ,but .because The system selected "weak light pulse" as the first-level micro-probe action for this round and collected the response within a 1.2-second observation window.

[0139] Let the normalized response eigenvector before micro-trial be... The three components correspond to the device fixation probability, device approach tendency, and avoidance tendency, respectively; the normalized response feature vector after micro-probing is... Therefore, the short-time difference response characteristics are:

[0140] ;

[0141] The above results indicate that users exhibited increased fixation probability, shortened distance, and no obvious avoidance response characteristics after micro-probing. The system matched this differential response with the response template of the "dejected-irritable" candidate pair under the "weak light pulse" action, obtaining... , When taking the matching correction strength coefficient hour,

[0142] The corrected weight for "decline" is:

[0143] ;

[0144] The corrected weight for "irritability" is:

[0145] ;

[0146] As can be seen, after a micro-test confirmation, the candidate weight of "declined" changed from Upgraded to approximately The candidate weight for "irritability" is determined by... dropped to about Further, take the comprehensive confirmation coefficient. , , Confirm the threshold value The overall confirmation score for "depression" is:

[0147] ;

[0148] The overall confirmation score for "irritability" is:

[0149] ;

[0150] because And significantly higher than Therefore, the system outputs "depressed" as the confirmed emotional state during this confirmation cycle and sends this state to the subsequent companionship feedback control module. This implementation case shows that in nighttime scenarios with high privacy, low disturbance, and low resource constraints, the present invention can still complete emotion confirmation with low disturbance through a closed loop of candidate emotion pair limitation, separability determination, non-semantic micro-probing, short-time differential response, and template matching correction.

[0151] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All modifications made according to the spirit and essence of the main technical solution of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A non-semantic micro-probing confirmation method based on the separability of candidate emotions, characterized in that, Includes the following steps: S1. Acquire the user's multimodal input signals; S2. Perform a lightweight emotion assessment on the multimodal input signal to obtain multiple candidate emotion values; S3. Select at least two candidate emotions that rank highest to form a candidate emotion pair; S4. Calculate the separability of the candidate emotion pairs, whereby the separability is used to characterize the degree of distinction between the candidate emotion pairs; S5. Determine whether the separability reaches the preset confirmation threshold: if it does, directly output the confirmed emotional state. If the target is not met, proceed to the micro-probe confirmation mode; S6. Based on the candidate emotion pairs and the current scene constraints, select and execute the corresponding non-semantic micro-probing action. The non-semantic micro-probing action is a low-semantic, short-term, weak-stimulus non-questioning action. S7. Collect user response within a short observation window after the micro-probing action is executed, and extract differential response features before and after the micro-probing. S8. Match the differential response features with the predefined response templates of the corresponding candidate emotion pairs to obtain the matching degree of each candidate emotion. S9. Based on the matching degree, the weights of each candidate emotion in the candidate emotion pair are adjusted in a targeted manner to obtain the adjusted candidate emotion value. S10. Based on the modified candidate emotion value, determine whether the confirmation condition is met. If it is met, output the confirmed emotion state. If it is not met, execute the upgrade, delay or rollback strategy.

2. The non-semantic micro-probing confirmation method according to claim 1, characterized in that, In S4, the separability of candidate emotion pairs is calculated using the following formula: ; in, For the current candidate sentiment pairs, , For the two emotion categories in the candidate emotion pair, , These are the initial candidate scores for both. The difference in scores between the two. For sorting stability within a continuous time window, To support consistency across multiple modalities, For the current sampling quality, The prior constraints of historical confirmation results on the current candidate pair are α, β, γ, δ, and η, which are preset weight coefficients.

3. The non-semantic micro-probing confirmation method according to claim 2, characterized in that, In S6, the selection process for the non-semantic micro-probing action includes: Predefine a priority table of trial actions for different candidate emotions, as well as the power consumption cost, disturbance level, and privacy risk level of the actions; Based on the current scenario's power consumption budget, privacy level, permissible stimulus intensity, remaining device battery power, resource usage level, and action cooldown time constraints, a set of candidate actions that meet the requirements is selected; Each candidate action is scored based on the following action selection scoring formula, and the action with the highest score is selected for execution: ; in, For the j-th candidate action, The degree to which this action distinguishes and fits the current candidate emotion pairs. For scene permission, For the sake of power consumption, In exchange for the disturbance, For the cost of privacy risks, , , , , These are the preset weighting coefficients.

4. The non-semantic micro-trial confirmation method according to claim 1, characterized in that, In S7, the extraction of the differential response features before and after the micro-probe specifically includes: Extracting the feature vector before micro-probe execution and the feature vector within a short observation window after microprobe execution The features include at least one or more of the following: changes in gaze direction, changes in device distance, changes in touch probability, changes in head rotation, changes in avoidance actions, and changes in voice activity. Calculate the differential response characteristics : 。 5. The non-semantic micro-probing confirmation method according to claim 4, characterized in that, In S9, the weights of each candidate emotion within a candidate emotion pair are adjusted in a targeted manner using the following formula: ; in, To correct candidate sentiment The weight, Given its initial candidate score, denoted as , where is the matching degree between the differential response features and the corresponding response template of the candidate emotion, and μ is the correction strength coefficient of the matching result. This represents the current candidate sentiment pair.

6. The non-semantic micro-probing confirmation method according to claim 5, characterized in that, In S10, determining whether the confirmation condition is met based on the corrected candidate emotion value specifically includes: Calculate the overall confirmation score for each candidate emotion: ; Wherein, ρ1, ρ2, and ρ3 are preset weight coefficients; If the overall confirmation score of a candidate emotion is greater than the preset confirmation threshold, then the candidate emotion is output as a confirmed emotion state; otherwise, an escalation, delay, or rollback strategy is executed.

7. The non-semantic micro-trial confirmation method according to claim 1, characterized in that, The upgrade, delay, or rollback strategies specifically include: If the confirmation conditions are not met and the current scenario allows for continued probing, the micro-probing action will be upgraded to a higher level of stimulation and the confirmation process will be repeated. If the current scenario does not allow for immediate continuation of the trial, the confirmation process will be delayed until the next interactive window. If the confirmation threshold is not reached after multiple consecutive confirmations, or if the preset maximum number of attempts is reached, the output will revert to a neutral state, a conservative state, or a pending confirmation state, and no new micro-probing actions will be performed. Set the maximum number of attempts within the same confirmation cycle, and the minimum cooldown interval between two adjacent micro-probe actions to avoid continuous stimulation and cumulative disturbance.

8. The non-semantic micro-trial confirmation method according to claim 1, characterized in that, The non-semantic micro-probing actions include at least one of the following: changes in light intensity / color, slight device displacement, orientation rotation, short vibrations, and low-brightness visual cues, but do not include natural language questioning actions.

9. A non-semantic micro-trial confirmation system based on the separability of candidate emotion pairs, characterized in that, include: The signal acquisition module is used to acquire the user's multimodal input signals; A lightweight evaluation module is used to process the multimodal input signal and output multiple candidate emotions and their initial candidate scores; The candidate pair construction module is used to select at least two top-ranked candidate emotions to form a candidate emotion pair. A separability calculation module is used to calculate the separability of the candidate emotion pairs; The trigger judgment module is used to compare the separability with a preset threshold to determine whether to enter the micro-probe confirmation mode; The action selection and execution module is used to select and execute non-semantic micro-probing actions based on candidate emotion pairs and scene constraints. The response acquisition and feature extraction module is used to acquire user responses within a short observation window after the micro-probing action is executed, and to extract differential response features. The template matching module is used to match the differential response features with the corresponding response templates to obtain the matching degree of each candidate emotion; The weight correction module is used to make targeted corrections to the candidate emotion weights within candidate emotion pairs based on the matching degree. The confirmation decision module is used to output confirmation results based on the corrected candidate sentiment values, or to invoke upgrade, delay, or rollback strategies when the target is not met.

10. The non-semantic micro-probing confirmation system according to claim 9, characterized in that, The system is deployed in desktop companion robots, children's companion devices, elderly companion terminals, or low-power intelligent interactive devices, and all processing flows are completed locally.