Emotion state confirmation method based on micro-olfactory exploration and related device

CN122604375APending Publication Date: 2026-08-21SHANGHAI TIMAN BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610730900.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0003]然而,穿戴式生理信号识别需要用户长期佩戴电极或腕带等设备,舒适性差且长期佩戴接受度低,难以在日常场景中部署;非接触式视觉识别则在弱光、用户低头侧脸或人脸距镜头较远等弱特征工况下,由于表情几何信号、远程光体积描记心率信号与肤色变化信号同步退化,融合后输出的综合置信度低于可决策阈值,使得系统既无法输出可靠的情绪类别,又因持续等待自发表情而显著延长识别响应时延;同时,非接触式方案缺乏向用户主动施加可控刺激并观察响应的能力,识别过程完全依赖于自发表情或自发节律的偶发暴露,进一步加剧了弱特征工况下的识别失效

Benefits of technology

[0047] 1. By linking passive multimodal initial judgment with respiratory response analysis under active probing of trace odor vortex rings into a weighted confirmation link, non-contact emotion recognition under weak feature conditions such as low light, head down, and side profile is upgraded from passive waiting to active stimulation observation without relying on wearable devices, thereby improving recognition certainty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122604375A_ABST
    Figure CN122604375A_ABST
Patent Text Reader

Abstract

The application relates to an emotion state confirmation method based on micro odor exploration and a related device, which comprises the following steps: performing multi-modal feature fusion processing on collected face image frames to output a comprehensive emotion category and a comprehensive confidence; in response to the comprehensive confidence being lower than a preset dynamic trigger threshold, emitting a vortex ring carrying a micro stimulating odor to an identified user, and performing double-window comparison detection on breathing response audio and breathing baseline audio before the vortex ring is emitted, and outputting a breathing pause confidence; performing weighted fusion on the comprehensive confidence and the breathing pause confidence to output a confirmed emotion category; and in response to the dose level reaching an upper limit and the breathing pause confidence still being lower than an effective detection threshold, marking an identification blind area state. By connecting passive multi-modal visual preliminary judgment and breathing response analysis under active micro odor exploration into a weighted confirmation link, the confidence of emotion recognition under weak feature working conditions can be improved without relying on wearable devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of emotion state recognition, and in particular to a method and related apparatus for confirming emotion state based on trace odor probing. Background Technology

[0002] Emotion recognition has wide applications in human-computer interaction, security monitoring, and psychological assessment. Common emotion recognition technologies are divided into two main categories based on whether users need to wear devices: wearable physiological signal recognition and non-contact visual recognition. Wearable physiological signal recognition collects the user's autonomic nervous system response characteristics through devices such as heart rate, skin conductance, and electromyography, and has a high confidence level in judging emotion categories. Non-contact visual recognition captures facial image frames of the user through a camera and maps emotion categories based on facial expression geometric features, remote photoplethysmography of heart rate, skin color changes, and other channels, without requiring the user to wear any devices.

[0003] However, wearable physiological signal recognition requires users to wear electrodes or wristbands for extended periods, which is uncomfortable and has low acceptance for long-term use, making it difficult to deploy in everyday scenarios. Non-contact visual recognition, on the other hand, suffers from weak feature conditions such as low light, users looking down or with their faces turned to the side, or faces being far from the camera. This is because facial expression geometric signals, remote photoplethysmography heart rate signals, and skin color change signals degrade simultaneously, resulting in a combined output confidence level below the decision threshold. Consequently, the system cannot output a reliable emotion category and significantly prolongs the recognition response latency due to the continuous waiting for spontaneous empathy. Furthermore, non-contact solutions lack the ability to actively apply controllable stimuli to users and observe their responses. The recognition process relies entirely on occasional exposure to spontaneous empathy or rhythms, further exacerbating recognition failures under weak feature conditions.

[0004] Therefore, the art aims to propose a method for confirming emotional states that can proactively acquire analyzable responses, improve recognition confidence, and explicitly mark unrecognizable intervals without relying on wearable devices and under weak feature conditions. Summary of the Invention

[0005] In order to proactively acquire resolvable responses under weak feature conditions to improve recognition confidence without relying on wearable devices, and to explicitly mark unrecognizable intervals, this application provides a method and related device for confirming emotional state based on trace odor probing.

[0006] Firstly, this application provides a method for confirming emotional states based on trace odor probing, employing the following technical solution:

[0007] A method for confirming emotional state based on trace odor probing includes the following steps:

[0008] S1. Perform multimodal feature fusion processing on the acquired facial image frames and output the comprehensive emotion category and comprehensive confidence score;

[0009] S2. In response to the overall confidence level being lower than a preset dynamic trigger threshold, a vortex ring carrying a trace amount of irritant odor is emitted to the identified user, and within a predetermined time window when the vortex ring arrives at the expected time, a dual-window comparison detection is performed on the breathing response audio of the identified user and the breathing baseline audio before the vortex ring is emitted, and the breathing pause confidence level is output; wherein, the trace amount of irritant odor is an irritant gas with a concentration lower than a preset concentration threshold.

[0010] S3. Perform a weighted fusion of the overall confidence level and the confidence level of breathing pause, and output the confirmed emotion category;

[0011] S4. In response to the fact that the dose level of the vortex ring has reached the preset dose limit and the confidence level of respiratory arrest is still lower than the preset effective detection threshold, mark the current state as a recognition blind zone.

[0012] By adopting the above technical solution, the method uses the comprehensive confidence score output by multimodal feature fusion as the initial visual judgment channel. Only when the comprehensive confidence score is lower than the dynamic trigger threshold is a small amount of stimulating odor carried by the vortex ring actively applied. This transforms the original low-confidence observation process of passively waiting for natural expressions into a controllable process of active stimulation and observation of responses. Within a predetermined time window after the arrival of the vortex ring, a dual-window comparison detection is performed between the respiratory response audio and the respiratory baseline audio before the vortex ring is emitted. This compresses the scanning of the entire respiratory rhythm into a local locking of the stimulus response. Finally, a weighted fusion output is performed on the comprehensive confidence score and the confidence score of respiratory pause to confirm the emotion category. When there is no pause response even when the dose reaches the upper limit, the blind zone state is explicitly marked, thereby avoiding the forced output of incorrect emotions downstream based on low-quality criteria.

[0013] Optionally, the vortex ring in S2 is launched according to a dose-stepping strategy, which includes: setting the dose level of the vortex ring sequentially according to a preset dose sequence; performing a dual-window comparison detection after each vortex ring launch; and in response to the fact that the confidence level of respiratory arrest is still lower than the effective detection threshold and the current dose level of the vortex ring has not reached the upper dose limit, raising the dose level of the vortex ring to the next level in the dose sequence, and launching the vortex ring again according to the raised dose level.

[0014] By adopting the above technical solution, the dosage starts from the lowest level and is gradually increased as needed, avoiding user discomfort caused by administering a high dose at once, while ensuring that the stimulation intensity is continuously increased until a detectable respiratory arrest event is triggered when the response is insufficient.

[0015] Optionally, the method further includes: establishing an individual response baseline library for the identified user, the individual response baseline library storing multiple quadruple records, each quadruple record including odor type, dose level, response delay and pause amplitude; if the confidence level of the respiratory pause after the launch of the current vortex ring is higher than the effective detection threshold, the odor type, dose level and corresponding response delay and pause amplitude used in this launch are written into the individual response baseline library as new quadruple records; before the next execution of S2, based on the existing quadruple records in the individual response baseline library, the odor type and initial dose level of the trace stimulation odor carried by the next vortex ring are optimally determined according to the normalized ratio of pause amplitude and dose level.

[0016] By adopting the above technical solution, the method maps the stimulus response to the individual baseline after each successful trial and selects the odor type and starting dose accordingly before the next trial, forming a read-write closed loop, which continuously improves the trial efficiency of each user and avoids the use of odors with weak response.

[0017] Optionally, the method further includes: extracting respiratory rhythm harmonic components from the audio of the identified user's breathing response within a predetermined time window when the vortex ring in S2 reaches the expected time, and performing matching between the respiratory rhythm harmonic components and preset emotional state harmonic priors to output rhythm harmonic confidence; the input of the weighted fusion also includes the rhythm harmonic confidence, forming a three-way weighted fusion of comprehensive confidence, breathing pause confidence and rhythm harmonic confidence.

[0018] By adopting the above technical solution, the method extracts the respiratory rhythm harmonic components and matches them with the emotional prior within the trial window, extending the detection of a single respiratory pause event to a dual criterion coupling of pause and rhythm. Even if the respiratory pause is not significant, a reinforcing criterion can be obtained from the harmonic structure.

[0019] Optionally, the dynamic trigger threshold is dynamically adjusted based on the current ambient light intensity, the face orientation angle of the identified user, and the distance between the identified user and the image acquisition module that acquires the facial image frame; the execution of S2 also satisfies the following timing conditions: the face orientation angle of the identified user is within the preset effective orientation range; the noise intensity of the current environment is lower than the preset ambient noise threshold; the interval since the last vortex ring emission exceeds the preset cooling time; and the identified user is unique.

[0020] By adopting the above technical solution, the method separates the dynamic trigger threshold and timing conditions into a pre-gating layer, which suppresses unnecessary trial launches in adverse scenarios such as low light, strong noise, user deviation from the camera, or the cooling period immediately after launch, thus avoiding repeated or ineffective disturbances to the user.

[0021] Optionally, in response to the presence of multiple candidate users in the facial image frame, the candidate user whose face orientation angle is within the effective orientation range and is closest to the image acquisition module is identified as the recognized user, and S2 is performed only on the recognized user.

[0022] By adopting the above technical solution, the method can lock a single target in a multi-person scenario based on both orientation angle and distance criteria, and deliver the trace amount of irritating odor carried by the vortex ring to the identified user in a concentrated manner, thus avoiding unnecessary disturbance to other people in the scene.

[0023] Optionally, the method further includes: storing additional respiratory baseline fingerprint records for the identified user in an individual response baseline library, the respiratory baseline fingerprint records including resting respiratory frequency, respiratory amplitude distribution and respiratory rhythm harmonic structure; before performing dual-window comparison detection, calculating the respiratory activity variance of the respiratory baseline audio, and in response to the respiratory activity variance exceeding a preset variance threshold, using the respiratory baseline fingerprint records of the identified user instead of the respiratory baseline audio as the baseline reference for dual-window comparison detection.

[0024] By adopting the above technical solution, when the baseline audio window before the trial is contaminated by transient disturbances, the method uses the historical fingerprint baseline instead of the immediate baseline as a comparison reference to ensure that weak respiratory arrest events can still be effectively detected under atypical resting conditions.

[0025] Optionally, the method further includes: in response to the dose level of the vortex ring reaching the upper limit of the dose and the confidence of respiratory arrest still being lower than the effective detection threshold, reading the rhythm harmonic confidence, and in response to the rhythm harmonic confidence being higher than the preset effective threshold of the blind zone, outputting the emotion category corresponding to the rhythm harmonic confidence as a weakly affirmative emotion.

[0026] By adopting the above technical solution, the method uses rhythmic harmonic components as a fallback criterion in the blind zone state where the detection of pause events fails, so that it can still output weak affirmative emotions under the condition of limited dose upper limit, avoiding complete silence.

[0027] Optionally, the method also includes: in response to the dose level reaching the upper limit, the confidence level of respiratory arrest being lower than the effective detection threshold, and the confidence level of rhythm harmonics being lower than the effective threshold of the blind zone, entering an extended passive observation period, stopping active probing during the extended passive observation period and continuing to execute S1, waiting for environmental conditions to improve to the point that the overall confidence level is higher than the dynamic trigger threshold before resuming normal confirmation output.

[0028] By adopting the above technical solution, the method stops active probing and switches to long-term passive observation when rhythmic harmonic fallback is also unavailable, thus avoiding unnecessary repeated disturbances to users in adverse environmental conditions.

[0029] Optionally, the method further includes: during the identification blind zone, if the overall confidence level of passive observation continues to be higher than the dynamic trigger threshold for a preset cancellation time, the previously marked identification blind zone state is reversed and cancelled, and the overall emotion category of passive observation is output as a weak confirmation emotion.

[0030] By adopting the above technical solution, the method allows the blind zone state to be reverse-covered by high-quality passive observation after environmental improvement, upgrading the blind zone from a terminated state to a reversible state, and improving the coverage of the method in long-term use.

[0031] Optionally, the method further includes: calculating the equivalent flight velocity of the vortex ring based on the launch records of each vortex ring and the corresponding detection time of the breathing pause event, and dynamically correcting the center time of the current predetermined time window based on the calculated equivalent flight velocity of the vortex ring and the distance between the user being identified and the image acquisition module.

[0032] By adopting the above technical solution, the method uses the delay data of each trial to estimate the expected arrival time of the next vortex ring in self-calibration, so that the window position of the dual-window comparison detection is continuously aligned with the actual response time, thereby improving the detection rate of pause events in weak response scenarios.

[0033] Secondly, this application provides an emotional state confirmation system based on trace odor detection, which adopts the following technical solution:

[0034] A system for confirming emotional state based on trace odor probing includes: an image acquisition module configured to acquire facial image frames; an audio acquisition module configured to acquire respiratory response audio and respiratory baseline audio; an emotion recognition module configured to perform multimodal feature fusion processing on the facial image frames and output a comprehensive emotion category and a comprehensive confidence level; and a trace odor probing module comprising a trace odor chamber and a vortex ring forming mechanism, wherein the trace odor chamber is divided into multiple independent sub-cavities, each sub-cavity carrying a trace stimulating odor, and the vortex ring forming mechanism is configured to emit a trace odor to the identified user in response to a transmission command. The system includes: a vortex ring for detecting odor stimuli; a respiratory feature analysis module configured to perform a dual-window comparison detection of the respiratory response audio and the respiratory baseline audio within a predetermined time window when the vortex ring reaches its expected arrival time, and output the confidence level of respiratory pause; and a main control module configured to: issue a transmission command to the trace odor detection module in response to the overall confidence level being lower than a preset dynamic trigger threshold; perform weighted fusion of the overall confidence level and the confidence level of respiratory pause, and output the confirmed emotion category; and mark the current state as a blind zone in response to the vortex ring's dose level reaching a preset dose upper limit and the confidence level of respiratory pause still being lower than a preset effective detection threshold.

[0035] By adopting the above technical solution, the system integrates active excitation capability into the wearable-free architecture in a non-contact manner through the multi-odor structure of the independent sub-cavities in the micro-odor probing module and the remote delivery capability of the vortex ring forming mechanism; the emotion recognition module, breathing feature analysis module, and main control module form a serial link in the data flow of visual initial judgment, active probing, response analysis, and weighted confirmation, which corresponds one-to-one with the method steps in the first aspect, so that the method has complete support in the hardware entity.

[0036] Thirdly, the computer device provided in this application adopts the following technical solution:

[0037] A computer device comprising:

[0038] One or more processors;

[0039] Memory;

[0040] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to:

[0041] Perform the above-described method for confirming emotional state based on trace odor probing.

[0042] Fourthly, the computer-readable storage medium provided in this application adopts the following technical solution:

[0043] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above.

[0044] The storage medium stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the following:

[0045] Such as the above-mentioned method for confirming emotional state based on trace odor probing.

[0046] In summary, this application includes at least one of the following beneficial technical effects:

[0047] 1. By linking passive multimodal initial judgment with respiratory response analysis under active probing of trace odor vortex rings into a weighted confirmation link, non-contact emotion recognition under weak feature conditions such as low light, head down, and side profile is upgraded from passive waiting to active stimulation observation without relying on wearable devices, thereby improving recognition certainty.

[0048] 2. Through the synergy of weighted mechanisms such as dose stepwise escalation, individual response baseline quadrupole closure, rhythmic harmonic additional criteria, environmental adaptive threshold, and multi-user scene attention orientation, the method continuously optimizes the stimulus response efficiency of each user in long-term use and maintains accurate locking of a single target in multi-user and adverse environments.

[0049] 3. By supplementing mechanisms such as switching from the respiratory baseline fingerprint database to the historical reference when the instantaneous baseline is disturbed by transients, providing fallback output of rhythmic harmonics in the blind zone, allowing high-quality passive observation and reverse verification of the blind zone state, and self-calibrating the vortex ring arrival delay, the method maintains recognition coverage and window position accuracy under extreme conditions such as atypical resting states of users, blind zones with dose upper limits, and long-term continuous use. Attached Figure Description

[0050] Figure 1 This is a schematic diagram illustrating the application environment of an emotional state confirmation method based on trace odor probing in one embodiment of this application.

[0051] Figure 2 This is a flowchart illustrating the overall process of an emotional state confirmation method based on trace odor probing in one embodiment of this application.

[0052] Figure 3 This is a schematic diagram of an emotional state confirmation system based on trace odor probing in one embodiment of this application.

[0053] Figure 4 This is a schematic diagram of a computer device according to an embodiment of this application. Detailed Implementation

[0054] The present application will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative of the application and are not intended to limit the scope of the application.

[0055] This application provides a method for confirming emotional state based on trace odor probing. (Refer to...) Figure 1 and Figure 2 This method actively emits a vortex ring carrying a trace amount of stimulating odor as a probe source to the identified user when the overall confidence level given by multimodal passive perception is lower than the dynamic trigger threshold. By observing the respiratory pause response of the identified user within a very short time window after the arrival of the vortex ring, passive observation is transformed into controllable stimulus-response analysis. When the probe still fails to form an effective response, the identification blind zone state is explicitly marked to avoid forcibly outputting the wrong emotion category under weak feature conditions.

[0056] The term "trace amount of irritant odor" as used in this application refers to an odor dose sufficient to elicit a detectable respiratory response, but significantly lower than the dose required to cause subjective perception or discomfort to the identified user. In engineering implementation, the odor dose carried by a single vortex ring is typically at the lower end of the human olfactory perception threshold, and the identified user is usually unaware of a complete probe. The term "vortex ring" as used in this application refers to a ring-shaped airflow formed by a diaphragm or piston driven through a circular nozzle. Its flight trajectory is relatively concentrated, capable of directionally delivering the trace amount of odor carried within it to the identified user 1 to 2 meters away without diffusing it throughout the space. The term "comprehensive confidence level" as used in this application refers to the confidence value given by integrating multimodal passively perceived visual and physiological signals for the current emotion category determination. The term "identification blind zone state" as used in this application refers to a state where the method explicitly acknowledges that a credible emotion category cannot be determined using all employed methods, as opposed to forcibly outputting a low-confidence result.

[0057] In a typical home or office desktop application scenario, the method and corresponding system of this application are deployed at a distance of 1 to 2 meters from the user being identified, such as opposite the desktop or next to the monitor. The image acquisition module acquires facial image frames of the user being identified at a frame rate of approximately 30 frames per second, and the audio acquisition module acquires respiratory baseline audio and respiratory response audio at a sampling rate of approximately 48 kHz. In this scenario, the light intensity may vary significantly depending on the orientation of the window during the day and the position of the desk lamp at night. The user being identified may also temporarily deviate from the optimal acquisition posture due to looking down, turning their face to the side, or talking to others, thus constituting one of the weak feature conditions that the method must address.

[0058] In the above scenario, the method first completes the multimodal passive perception and preliminary confidence assessment of the identified user through S1, and then decides whether to trigger subsequent active probing.

[0059] S1. Perform multimodal feature fusion processing on the acquired facial image frames and output the comprehensive emotion category and comprehensive confidence level.

[0060] In some embodiments, the multimodal feature fusion processing consists of three passive perception channels: facial expression geometry modality, remote photoplethysmography modality, and skin color change modality. The facial expression geometry modality extracts the geometric positions and displacement trajectories of 68 facial key points from facial image frames, outputting an expression category probability distribution. The remote photoplethysmography modality performs bandpass filtering and spectral analysis on the temporal changes of green channel pixels within the facial skin region, outputting heart rate estimates and heart rate variability indices. The skin color change modality performs short-window statistics on the chroma channel components of the cheek and forehead regions, outputting skin color warm / cool offset. The outputs of the three modalities are concatenated and fed into a fusion layer, which then synthesizes and outputs a comprehensive emotion category. And the overall confidence level ρ, where ρ represents one of several discrete emotion categories supported by the method, such as calm, tension, relaxation, focus, etc., and ρ is a real number between 0 and 1.

[0061] Under favorable conditions, such as when the user's face is directly facing the image acquisition module, at a distance of 1.5 meters, and with uniform indoor lighting, the three modalities can each provide high confidence levels. The overall confidence level ρ output by the fusion layer is typically above 0.80, and the method can complete emotion confirmation solely based on the output of S1. However, when the user is in low light, looking down, or with their face turned to the side, the confidence level of the facial expression geometry modality decreases due to the detection bias of the 68 key points. The remote photoplethysmography modality and the skin color change modality degrade simultaneously due to the reduction in effective facial pixels and the distortion of chromaticity information. The overall confidence level ρ output by the fusion layer drops significantly to below 0.65. In this case, relying solely on S1 is insufficient to provide a reliable emotion category determination.

[0062] In some embodiments, the multimodal feature fusion processing may consist only of the facial expression geometry modality and the remote photoplethysmography modality, omitting the skin color change modality; in still other embodiments, the multimodal feature fusion processing may rely solely on the facial expression geometry modality. The aforementioned simplified embodiments can still operate under conditions of stable acquisition environment and limited computing resources, but their overall confidence level has lower tolerance for weak feature conditions than the three-modal embodiments.

[0063] The overall confidence level ρ output by S1 is the criterion for whether S2 triggers active probing. When ρ is higher than the preset dynamic trigger threshold, the method directly uses... As the confirmed emotion output for this time; when ρ is lower than the dynamic trigger threshold, the method continues to execute S2.

[0064] The term "trace amount of irritating odor" in the preceding claims is further explained here. The trace amount of irritating odor refers to an irritant gas with a concentration below a preset concentration threshold. The irritant gas is selected from volatile substances capable of activating the olfactory receptors or trigeminal nerve receptors of the identified user, including but not limited to menthol vapor, eucalyptol vapor, cinnamaldehyde vapor, citral vapor, and citronellol vapor.

[0065] The preset concentration threshold is set based on two principles: First, the concentration of the irritant gas is sufficient to allow the identified user to perceive the presence of the odor through olfactory receptors or trigeminal nerve receptors at the expected time when the vortex ring arrives, thereby triggering a brief involuntary breathing response, including inspiratory apnea, expiratory delay, and changes in the depth of a single exhalation; Second, the concentration of the irritant gas is below the lower limit of the odorant substance causing obvious discomfort to the identified user, including observable discomfort reactions such as sneezing, coughing, tearing, upper respiratory tract stinging, and skin itching.

[0066] In one embodiment, the preset concentration threshold is between 1 and 3 times the human olfactory detection threshold of the irritant gas. For example, menthol vapor has a human olfactory detection threshold of approximately 0.01 ppm to 0.04 ppm, so the preset concentration threshold is typically set between 0.01 ppm and 0.12 ppm; eucalyptol vapor has a human olfactory detection threshold of approximately 0.05 ppm to 0.10 ppm, so the preset concentration threshold is typically set between 0.05 ppm and 0.30 ppm. Simultaneously, the total amount of the trace irritant gas carried by a single vortex ring is typically between 0.1 ml and 5 ml, ensuring that even if the vortex ring breaks on the face of the identified user and the odorant is completely released, the instantaneous concentration increase in the indoor space remains far below the safety limit for the odorant.

[0067] The slightly irritating odor, as designed, ensures that when the vortex ring reaches the face of the user being identified, it triggers a respiratory response that can be captured by dual-window contrast detection. It also avoids causing obvious subjective feelings or health risks to the user being identified, thereby supporting the comfort and safety of the method described in this application embodiment in long-term daily use scenarios.

[0068] S2. In response to the overall confidence level being lower than the preset dynamic trigger threshold, a vortex ring carrying a trace amount of irritant odor is emitted to the identified user. Within a predetermined time window when the vortex ring arrives at the expected time, a dual-window comparison detection is performed on the breathing response audio of the identified user and the breathing baseline audio before the vortex ring is emitted, and the confidence level of breathing pause is output.

[0069] The vortex ring is generated by a vortex ring forming mechanism within the trace odor detection module. In one implementation, the vortex ring forming mechanism includes a circular nozzle and an elastic diaphragm located behind it. The diaphragm is driven by an electromagnetic drive unit to generate instantaneous displacement, rapidly propelling the gas within the trace odor chamber through the nozzle. As the airflow passes through the nozzle edge, it is swirled by edge shearing to form a ring-shaped vortex. The ring-shaped vortex maintains relative structural stability during flight, directionally delivering the trace amount of stimulating odor carried within it to a target location tens of centimeters to several meters away within hundreds of milliseconds, without diffusing throughout the room.

[0070] In the aforementioned home or office desktop scenarios, the distance *d* between the identified user and the image acquisition module is typically 1 to 2 meters. The equivalent flight speed of the vortex ring. At the device's calibrated dose, the time is typically between 2 and 4 meters per second, therefore the flight time required for the vortex ring to travel from the nozzle to the face of the identified user. Typically, this ranges from 0.25 seconds to 1 second. For example, when d is 1.5 meters... Under the condition of 3 meters per second, Approximately 0.5 seconds, the method is based on the vortex ring launch time. As zero point, The estimated arrival time is used as the center, and a predetermined time window is formed by extending 0.3 seconds forward and backward from the estimated arrival time. The setting of this time window covers both the small deviation of the actual arrival time of the vortex ring around the calibrated speed and the inherent delay of the identified user's respiratory response relative to the arrival time of the stimulus, so that the method will not miss the real respiratory response event due to small estimation errors in the predicted time.

[0071] Considering that the sound pressure level of the user's resting breathing at a distance of 1 to 2 meters is much lower than the ambient noise level, the audio acquisition module employs a directional microphone array facing the user at the hardware level. In one embodiment, the microphone array consists of four microelectromechanical system (MEMS) microphones arranged in a circle with a radius of 30 mm. Each microphone has a sensitivity of -38 dB and a full-scale reference sound pressure level of 94 dB, with a frequency response range covering 100 Hz to 8 kHz. The array performs delayed summation beamforming on the acquired multi-channel audio. The main lobe direction is determined in real-time by the angle of the user's face and the distance 'd' measured by the image acquisition module, ensuring the main lobe always points towards the user's mouth and nose. The output of the beamforming is the single-channel source signal for the subsequent respiratory baseline audio and respiratory response audio, typically improving the equivalent signal-to-noise ratio by 6 dB to 12 dB compared to an omnidirectional single microphone, thus preventing the user's breathing at a distance of 1 to 2 meters from being completely drowned out by ambient noise.

[0072] The dual-window comparison detection is performed within the aforementioned predetermined time window. Specifically, the method uses the vortex ring launch time... A fixed-duration audio segment (e.g., 2 seconds) is used as the baseline audio for breathing, and audio segments within a predetermined time window are used as the breathing response audio. Before feature extraction, both audio streams undergo a preprocessing step: first, a second-order Butterworth bandpass filter with cutoff frequencies of 80 Hz and 1 kHz is used to suppress low-frequency power supply interference, residual heartbeat sounds, and residual mid-to-high frequency ambient sounds, concentrating the filtered signal in the main energy frequency band of the respiratory airflow, namely the 200 Hz to 500 Hz range for inhalation sounds and the 100 Hz to 300 Hz range for exhalation sounds; then... Earlier periods (e.g.) The noise spectrum was estimated by using the previous 5 to 3 seconds period as a reference segment for environmental noise. Spectral subtraction noise reduction was performed on both audio streams to further suppress the influence of quasi-steady-state background noise such as air conditioner humming and remote keyboard typing.

[0073] For the two preprocessed audio streams, the method calculates the short-time energy window by window, with a step size of 10 milliseconds and an analysis window of 25 milliseconds. ,in This represents the number of samples within a single analysis window (1200 samples for 25 milliseconds at 48 kHz sampling). For the first analysis window Each sample value. The identified user exhibits "no airflow sound" segments during the resting expiratory interval and the end of inspiration. The method uses short-time energy of a single frame. Below the dynamic threshold And the duration should not be less than the reference duration. A continuous sequence of frames is defined as a continuous silence segment; where The average short-time energy of all analysis windows within the baseline audio segment according to Dynamically determined, The default value is 0.3 seconds. This dynamic threshold design enables the method to adapt to individual differences in the absolute intensity of the breathing sounds of different identified users, without the need to individually calibrate a fixed energy threshold for each user.

[0074] Within the baseline audio segment, the method identifies all consecutive silent segments and calculates their duration, obtaining the average duration of the baseline silent segments. With the standard deviation of duration Within the response audio segment, the method identifies consecutive silent segments using the same rules and takes the duration of the longest one. When the longest silent segment in the response segment is significantly prolonged relative to the average silent segment of the baseline segment, the method determines that the vortex ring has triggered a respiratory arrest response in the identified user. The specific criterion is as follows: ,in The extension coefficient is typically set to 2. This criterion, namely the statistical significance condition that "the longest silent segment of the response segment exceeds twice the average silent segment of the baseline segment", is used to exclude false triggers caused by random fluctuations in the baseline respiratory rhythm of the identified user.

[0075] Confidence of respiratory arrest It is obtained by normalizing the extension of the response segment relative to the baseline segment. Specifically, the method is as follows: Calculation, where For reference, a typical extension is 0.5 seconds, meaning the identified user experiences an additional respiratory pause of more than 0.5 seconds after receiving vortex stimulation. Reaching the upper limit of 1. After determining that the vortex ring triggered a respiratory arrest response, the method is based on the response delay (the offset of the start of the continuous silent segment relative to the expected arrival time of the vortex ring) and the pause amplitude ( Two features are used to perform nearest neighbor matching with the four-tuple records of the currently identified user in the individual response baseline library to determine the emotion category corresponding to this response; this emotion category is determined by... The data is output together to the weighted fusion of S3 as a comprehensive emotion category. Independent sources of evidence other than those mentioned above.

[0076] In some embodiments, the active excitation medium can also be replaced by a faint sound, a momentary flash of light, or a momentary air temperature disturbance instead of a vortex ring. However, faint sounds are easily interfered with by existing sound sources in the room, momentary flashes of light can be clearly noticed by the identified user, and momentary air temperature disturbances are difficult to maintain directionality and easily diffuse into non-target areas. Therefore, this application uses a trace amount of odor vortex ring as the preferred medium for active excitation to balance the directionality of the stimulus, the imperceptibility of the identified user, and the repeatability of stimulus-response coupling.

[0077] Through active stimulation by S2, the method transforms unresolvable passive observation under weak feature conditions into controllable stimulus-response coupling, so that the involuntary respiratory arrest response of the identified user is introduced into the subsequent fusion judgment as an additional emotional signal that can be clearly captured by the instrument.

[0078] In some embodiments, the dynamic trigger threshold is dynamically adjusted based on the current ambient light intensity, the face orientation angle of the identified user, and the distance between the identified user and the image acquisition module that acquires the facial image frame; the execution of S2 also satisfies the following timing conditions: the face orientation angle of the identified user is within a preset effective orientation range; the noise intensity of the current environment is lower than a preset ambient noise threshold; the interval between the last vortex ring emission and the last vortex ring emission exceeds a preset cooling time; and the identified user is unique.

[0079] Specifically, dynamic trigger threshold Based on the benchmark threshold Multiplying this by a correction term, the value of which increases with decreasing ambient light intensity, with the angle of the user's face deviating from a frontal view, and with increasing distance from the user to the image acquisition module. In the aforementioned home desktop scenario, the method uses a baseline threshold... Set to 0.65; when ambient light is sufficient, the face is nearly directly facing the subject, and the distance is approximately 1.5 meters, the method will... Lowering the threshold to approximately 0.55 makes the method more conservative in triggering S2, because at this point, multimodal passive sensing is usually sufficient to provide reliable results, and the method only engages in active probing when ρ degrades significantly; conversely, when the ambient light is dim, the face is slightly angled, and the distance is close to 2 meters, the method will... The value was increased to approximately 0.70, allowing the method to actively probe into S2 early on before S1 has severely degraded.

[0080] The four timing conditions of S2 are given the following default engineering values: the effective orientation range of the face orientation angle is assumed to be no more than 30 degrees in absolute value. If the angle exceeds this range, the user being identified is considered to have deviated from the effective observation domain of the image acquisition module; the ambient noise threshold is assumed to be 50 dB sound pressure level. If this threshold is exceeded, the continuous silent segment determination of the dual-window comparison detection may fail due to fluctuations in the ambient sound pressure level; the cooling time since the last vortex ring emission is assumed to be 30 seconds to avoid continuous vortex ring emission in a short period of time causing the user being identified to notice and feel uncomfortable, and also to avoid starting the next test before the weak airflow caused by the previous vortex ring has completely dissipated. If any of the above conditions are not met, the method will not trigger S2 this time, but will continue to maintain the current emotion category determination with the passive perception output of S1 until all conditions are met at the same time and the overall confidence level is still lower than the dynamic trigger threshold.

[0081] Furthermore, in response to the presence of multiple candidate users in the facial image frame, the candidate user whose face orientation angle is within the effective orientation range and is closest to the image acquisition module is identified as the recognized user, and S2 is performed only on the recognized user.

[0082] Specifically, the method first uses the aforementioned effective orientation range as a screening condition to exclude candidate users whose face orientation angle is greater than 30 degrees, because these candidate users are in a state that is off-site to the observation domain, and it is difficult to obtain a resolvable visual and audio response even if S2 is triggered on them; then, among the remaining candidate users, the distance from the image acquisition module is used as the final judgment criterion, and the one closest to the module is determined as the identified user.

[0083] For example, in a scenario where two people are in the same facial image frame, candidate user A has an orientation angle of 15 degrees and a distance of 1.2 meters, while candidate user B has an orientation angle of 25 degrees and a distance of 1.6 meters. Both are within the valid orientation range. The method identifies candidate user A, who is closer, as the user to be identified, and candidate user B is not included in S2 in this confirmation. Conversely, if candidate user A has an orientation angle of 40 degrees and a distance of 1.0 meter, while candidate user B has an orientation angle of 10 degrees and a distance of 1.8 meters, the method excludes candidate user A because its orientation angle exceeds the valid orientation range, and then identifies candidate user B as the user to be identified.

[0084] In some operating conditions, there may be situations where two or more candidate users are close to each other and the difference is less than the calibration ranging resolution of the image acquisition module. In this case, the "uniqueness of the identified user" condition in S2 is not met, so S2 is not triggered in this case to avoid sending the vortex ring to the wrong object due to ambiguity in the selection of the identified user.

[0085] In some embodiments, the emission of the vortex ring in S2 is performed according to a dose stepwise escalation strategy, which includes: setting the dose level of the vortex ring sequentially according to a preset dose sequence; performing a dual-window comparison detection after each emission of the vortex ring; and, in response to the fact that the confidence level of respiratory arrest is still lower than the effective detection threshold and the current dose level of the vortex ring has not reached the dose upper limit, raising the dose level of the vortex ring to the next level in the dose sequence, and re-electing the vortex ring according to the upgraded dose level.

[0086] Specifically, the dose sequence consists of n levels, where n is typically 3 to 5. Each dose level corresponds to a fixed value set in the pulse width or diaphragm displacement amplitude of the vortex ring forming mechanism during a single emission, thereby controlling the volumetric increase of the odor dose carried by the vortex ring in a single emission. For example, the method presets four dose levels in the dose sequence. to , This corresponds to the minimum diaphragm displacement amplitude and ensures that the gas volume loaded in a single vortex ring is approximately 0.3 times the human olfactory perception threshold dose. to The gas volume was successively increased to 0.5 times, 0.8 times, and 1.2 times the threshold dose in a proportional manner. This is the upper limit of dosage. .

[0087] In a complete trial process, the method first follows... The vortex ring is launched and a dual-window comparison detection is performed. The output confidence level of the respiratory arrest is then... Still below the effective detection threshold (typically 0.55) and the current dose level If the upper dose limit is not reached, the method increases the dose level after the aforementioned cooling time. Relaunch; if If the threshold is still not reached, continue to rise to... , until The dose is above the effective detection threshold or the current dose level has reached its upper limit. For example, for a user with a relatively poor sense of smell, the first two identification attempts... and The emitted vortex rings may, due to insufficient dosage, induce detectable respiratory arrest, causing... It remains below 0.55; when the method increases to back, It may jump to 0.78, at which point the method stops further upshifting and... This serves as a record of the dosage levels used in this successful trial.

[0088] Upper limit of dose The engineering parameters are set with the bottom line that "the odor dose from a single emission should not make the identified user clearly aware that they have been tested afterward." In the aforementioned scenario, The corresponding odor dose is about 1.2 times the human olfactory perception threshold dose, which is on the edge of the dose range where "the identified user can barely distinguish it when actively sniffing, but usually cannot distinguish it when passively breathing". Emissions above this dose will significantly increase the probability of the identified user's detection, so the method no longer increases to higher doses.

[0089] The core mechanism of the step-by-step escalation strategy lies in starting with a low dose, which allows most of the identified users who are sensitive to odors to... or The level can be effectively detected, thus keeping the total dose of a single trial sufficiently low; the level is only increased when the low dose fails to produce an effective response, ensuring that even users who are relatively insensitive to odors can be correctly detected within the upper limit of the dose, thereby achieving a balance between the user's unconscious comfort and the method's effective response.

[0090] In some other embodiments, the method may also employ a fixed-dose scheme or a step-dose scheme. A fixed-dose scheme performs all probes with a single, higher dose, which is excessive for odor-sensitive users and appropriate for odor-insensitive users, thus increasing the overall probability of detection. A step-dose scheme directly uses the upper limit of the dose. When starting, or when ineffective, it drops to In contrast to the incremental strategy of this application, the starting dose is already excessive for identified users who are sensitive to odors. Compared to the above schemes, the dose incremental strategy of this application can provide a relatively optimal trial total dose for identified users with various sensitivities.

[0091] In some embodiments, the method further includes: establishing an individual response baseline library for the identified user, the individual response baseline library storing multiple quadruple records, each quadruple record including odor type, dose level, response delay and pause amplitude; in response to the confidence level of the breathing pause after the launch of the current vortex ring being higher than the effective detection threshold, the odor type, dose level and corresponding response delay and pause amplitude used in this launch are written into the individual response baseline library as new quadruple records; before the next execution of S2, based on the existing quadruple records in the individual response baseline library, the odor type and initial dose level of the trace stimulating odor carried by the next vortex ring are preferably determined according to the normalized ratio of pause amplitude and dose level.

[0092] Specifically, each quadruple record consists of four fields: the odor type field records the identifier of the trace stimulus odor selected for this trial among several odors that the device can emit; the dose level field records the dose level used when the vortex ring successfully induced a respiratory arrest response (e.g., ...). to (A specific file within the file), the response delay field records the time from the vortex ring launch. The measured time difference (typically 0.4 to 0.8 seconds) to the start of the continuous silence segment within the response window is recorded in the pause amplitude field, which records the extension of the continuous silence segment relative to the baseline window (typically 0.2 to 1.5 seconds). Multiple four-tuple records are indexed by the identity of the identified user and stored in the device's memory or in the associated server-side database.

[0093] The trigger condition for writing is: the confidence level of the breathing pause output by the dual-window comparison detection after this vortex ring launch. Above the effective detection threshold, meaning the trial is considered to have successfully elicited a detectable respiratory response. Only in this case is the method appended as a new quaternion record to the individual response baseline database, including the odor type, dose level, and corresponding response delay and pause amplitude. Unsuccessful attempts below the threshold are not written into the baseline library to avoid diluting subsequent optimization with invalid data.

[0094] The triggering timing for the reading end is: before the method determines that S2 needs to be executed next, but before the current vortex ring is emitted. At this time, the method reads all existing four-tuple records for the currently identified user in the individual response baseline library, and calculates the normalized ratio of pause amplitude to dose level for each record. Specifically, the pause amplitude is normalized to the maximum pause amplitude in the user's historical records, and the dose level is normalized to the highest level in the dose sequence. A larger ratio indicates that a more significant respiratory pause was induced at a lower dose level, reflecting the higher relative efficacy of the odor type for the user. The method selects the odor type corresponding to the record with the highest ratio as the preferred odor type for the trace stimulating odor carried in this vortex ring, and uses the dose level in that record as the starting dose level for this trial.

[0095] For example, after a user has successfully attempted three tests, the individual response baseline database contains the following three four-tuple records: the first record is odor type A and dose level. The response delay is 0.55 seconds, and the pause duration is 0.6 seconds; the second record specifies the odor type (B) and dosage level. Response delay 0.50 seconds, pause 0.9 seconds; the third record includes odor type A and dosage level. The response delay is 0.60 seconds, and the pause duration is 0.5 seconds. Assume the user's historical maximum pause duration is 0.9 seconds, and the upper limit of the dose sequence is... The normalized ratios of the three records are approximately (0.6 / 0.9) / (2 / 4)≈1.33, (0.9 / 0.9) / (3 / 4)≈1.33, and (0.5 / 0.9) / (1 / 4)≈2.22, respectively. The third record has the highest normalized ratio, therefore, in the next S2 step, the method selects odor type A and initial dose level... .

[0096] Through the above-mentioned writing and reading mechanism, the method forms a self-learning closed loop for the effective odor types and initial dose levels of each identified user. The trial efficiency of each user steadily increases with the accumulation of successful trials, and the dose required for a single trial also shows a decreasing trend.

[0097] In other embodiments, the method may not establish an individual response baseline library, and each S2 may be executed with a default odor type and a default starting dose; or a single-dimensional historical record (e.g., only recording the odor type of the last successful attempt without recording the dose level) may be used instead of the quadruplet record. The aforementioned simplified embodiments do not require maintaining historical data for each user, but the starting dose level and odor type for each trial cannot be optimized for the current user, resulting in an average number of trials and an average dose that are higher than the quadruplet closed-loop embodiment of this application.

[0098] Furthermore, the method also includes: storing additional respiratory baseline fingerprint records for the identified user in the individual response baseline library, the respiratory baseline fingerprint records including resting respiratory frequency, respiratory amplitude distribution and respiratory rhythm harmonic structure; before performing dual-window comparison detection, calculating the respiratory activity variance of the respiratory baseline audio, and in response to the respiratory activity variance exceeding a preset variance threshold, using the respiratory baseline fingerprint record of the identified user instead of the respiratory baseline audio as the baseline reference for dual-window comparison detection.

[0099] Specifically, the respiratory baseline fingerprint record consists of three fields: the resting respiratory rate field records the average respiratory rate of the identified user at rest, typically 12 to 18 breaths per minute; the respiratory amplitude distribution field records the energy distribution characteristics of the user's respiratory airflow in the form of a discrete probability density function; and the respiratory rhythm harmonic structure field records the first three harmonic components in the user's respiratory airflow spectrum. , , The relative intensity ratio. The data for the aforementioned three fields were built offline by accumulating long audio segments of the user in a quiet state multiple times and performing spectral analysis, and were continuously updated with high-quality baseline audio collected with each subsequent successful trial.

[0100] Before performing dual-window comparison detection, the method focuses on the vortex ring launch time. The variance of respiratory activity was calculated from the previous baseline audio clips. This variance reflects the degree of fluctuation in respiratory airflow energy over time in the baseline audio. The variance was low when the identified user was at rest; it increased significantly when the identified user had just finished speaking, coughing, or taking a deep breath.

[0101] For example, in the aforementioned office desktop scenario, the identified user has just ended a phone call less than 5 seconds ago, and the method determines that S2 needs to be executed. At this time, the baseline audio segment still retains residual vocalizations from the end of the call and high-frequency respiratory recovery signals, with variance several times higher than the typical value in the resting state. If this instantaneous baseline audio is directly used as the baseline reference for dual-window contrast detection, the contrast of the respiratory pause event caused by the vortex ring within the response window relative to this high-variance baseline is compressed, and the pause amplitude is underestimated. It is easily mistakenly judged as being below the effective detection threshold.

[0102] To address the aforementioned conditions, the method adds a variance determination step before the dual-window comparison detection: when the respiratory activity variance exceeds a preset variance threshold, the method replaces the immediate respiratory baseline audio with the respiratory baseline fingerprint record of the currently identified user from the individual response baseline library, serving as the baseline reference for the dual-window comparison detection. Specifically, a virtual baseline audio segment of the user in a quiet state is synthesized from the resting respiratory frequency and amplitude distribution in the fingerprint record. This virtual segment is then fed into the baseline window of the dual-window comparison detection, allowing the continuous silent segment determination in the response window to be compared with a stable baseline that represents the user's long-term respiratory characteristics, thereby restoring the detection sensitivity for weak respiratory arrest events.

[0103] The engineering default value for the variance threshold is typically set to three times the mean variance of the user's resting state. Below this value, the instantaneous baseline is considered sufficient to reflect the current respiratory state; above this value, it is considered to have deviated from the resting baseline and should be switched to the fingerprint baseline. When the individual response baseline library has not yet established a fingerprint record for the currently identified user, such as when the user first comes into contact with the device, the method still performs dual-window contrast detection using the instantaneous baseline audio, but the blind zone determination threshold of S4 is correspondingly relaxed to avoid outputting too many blind zone states to the user in the initial stage when a fingerprint baseline is lacking.

[0104] In some embodiments, the method further includes: extracting respiratory rhythm harmonic components from the respiratory response audio of the identified user within a predetermined time window when the vortex ring reaches the expected time in S2, and performing matching between the respiratory rhythm harmonic components and preset emotional state harmonic priors, and outputting rhythm harmonic confidence.

[0105] Specifically, the method performs a short-time Fourier transform on the audio segment of the response window to extract the fundamental frequency component of the breathing airflow from the spectrum. With several harmonic components , The relative intensity ratio. Physiologically, the spectral structure of spontaneous breathing airflow shows statistically significant differences when the identified user is in different emotional states: for example, the respiratory rate increases in states of tension or anxiety. High frequency and The proportion is relatively increased; the breathing rate decreases in a relaxed or pleasant state. Low frequency and Almost invisible. The method pre-stores the above statistical differences in the device in the form of emotional state harmonic priors, with each prior corresponding to an emotional category. , , Typical relative intensity ratios and typical fundamental frequency ranges.

[0106] The measured harmonic components extracted within the response window are matched with the preset prior harmonics of emotional states using cosine similarity to obtain the user's current most matching emotional category and its matching confidence score, which is denoted as the rhythmic harmonic confidence score. For example, the actual measurement obtained within the response window. , , If the relative intensity ratio is 1:0.45:0.15 and the fundamental frequency is around 0.30 Hz, and the matching degree with the "stress" prior (1:0.50:0.10, fundamental frequency 0.30 Hz to 0.40 Hz) reaches 0.82, then the method output... The value was 0.82, and the matched emotion category was labeled "tension".

[0107] The extraction of rhythmic harmonic components is limited to a predetermined time window when the vortex ring reaches its expected moment. This is because the minute stimulation of the vortex ring itself disrupts the original spontaneous breathing rhythm of the identified user within a short period, causing a "probe-like" perturbation in the harmonic structure observed within this window relative to the baseline without stimulation. The perturbed harmonic structure is more discriminative of emotional states compared to the harmonic structure of calm spontaneous breathing. Extraction only within the window after the initial stimulation ensures that the extracted harmonic components have been "activated" by this stimulus while avoiding contamination with the spectrum observed passively over long periods of time.

[0108] Rhythmic Harmonic Confidence The method has two uses in the subsequent process: the first output is fed into the weighted fusion of S3, along with the overall confidence ρ and the confidence of the respiratory arrest. Together they constitute the weighted input of the three parties; the second path serves as a fallback criterion in S4 under blind zone conditions, used to provide weak affirmative emotion when an effective respiratory arrest response cannot be obtained even after exhausting the upper limit of the dose in S2.

[0109] In other embodiments, the method further includes: back-calculating the equivalent flight velocity of the vortex ring based on the launch records of each vortex ring and the corresponding detection time of the breathing pause event, and dynamically correcting the center time of the current predetermined time window based on the back-calculated equivalent flight velocity of the vortex ring and the distance between the currently identified user and the image acquisition module.

[0110] Specifically, if the dual-window comparison detection successfully outputs a breathing pause event after each vortex ring launch, the method will record the launch time of this vortex ring. With pause event detection time difference As an estimate of the actual flight time, the ratio of this actual flight time to the distance d recorded at the time of launch is used. This serves as the measured value of the equivalent flight velocity of the vortex ring. The method performs a moving average of the measured velocities from the most recent period (e.g., the last 10 times) to obtain an estimate of the equivalent flight velocity of the current device under the current trace odor chamber gas formulation. .

[0111] Because the actual flight speed of the vortex ring may deviate from the factory calibration value due to factors such as diaphragm fatigue, changes in gas concentration within the chamber, or drift in ambient temperature and humidity during long-term use, if the factory calibration value is still used as the basis for calculating the estimated arrival time, the center time of the predetermined time window will gradually deviate from the actual arrival time over time, causing the response window to be misaligned with the actual respiratory arrest event. The false positive rate has increased. The method uses the aforementioned moving average... Replace the factory calibration value with the estimated arrival time for this test. As the center moment of the predetermined time window, the position of the predetermined time window adaptively follows the drift of the device's operating conditions, maintaining long-term detection accuracy without user intervention or manufacturer calibration.

[0112] S3. Perform a weighted fusion of the overall confidence level and the confidence level of breathing pause, and output the confirmed emotion category.

[0113] Specifically, the weighted fusion adopts a linear weighting form, that is, the confidence level after fusion. To combine the confidence level ρ with the confidence level of respiratory arrest In their respective weights and The weighted sum under, Among them, the weights and Satisfy normalization constraints And each takes a value between 0 and 1. In the aforementioned scenario, and Typical values ​​for ρ are 0.4 and 0.6, respectively. The reason for giving higher weight to the confidence level of respiratory arrest is that when S2 enters active probing, it means that the ρ given by S1 is already below the dynamic trigger threshold, and the effective information density carried by the respiratory arrest response is higher than that of the degraded visual signal.

[0114] Regarding the determination of emotion category: when ρ corresponds to the comprehensive emotion category... and When the corresponding breathing pauses correspond to the same emotion category, the method directly uses that category as the confirmed emotion category output; when the two do not correspond, the method uses the source with the greater weighted contribution (i.e., and The emotion category corresponding to the larger median value is used as the confirmed emotion category.

[0115] For example, under the aforementioned low-light, head-down operating condition, the output ρ of S1 is 0.55. For "neutral"; S2 in Output after vortex ring launch The value was 0.85, corresponding to the identified emotion category of "tension". The method uses... 0.4 The weighted fusion was calculated to 0.6, resulting in... Further comparison and The latter value is larger, so "tension" is used as the confirmed emotion category in this study.

[0116] Furthermore, when the method introduces rhythmic harmonic confidence, the input of the weighted fusion of S3 also includes rhythmic harmonic confidence, forming a three-way weighted fusion of comprehensive confidence, respiratory arrest confidence, and rhythmic harmonic confidence.

[0117] Specifically, the weighted fusion of the three parties is as follows: Three of the weights satisfy The relative magnitude of each weight is determined by the effective information density it carries: the confidence level of respiratory arrest directly reflects the discrete response events of the identified user to stimuli, and its discriminative power is the strongest relative to the other two inputs; therefore, its weight is the highest. Still set to maximum; the rhythmic harmonic confidence reflects the respiratory spectral structure after stimulation and perturbation, and is a continuous statistical characteristic, its weight... Assuming it's secondary; the overall confidence level ρ being triggered in S2 means it has degenerated below the dynamic trigger threshold, and its weight... Set to the minimum. In the aforementioned scenario, the typical values ​​for the three weights are as follows: 0.25 0.45 It is 0.30.

[0118] For example, under the aforementioned low-light, head-down operating condition, the output ρ of S1 is 0.55. For "neutral"; S2 in Output after vortex ring launch The value is 0.85, corresponding to the identified emotion category of "tension"; rhythm harmonic analysis output. The value is 0.82, and the corresponding identified emotion category is also "tension". The method calculates using the aforementioned weights. And compare the weighted contributions of the three paths: , , ;in The value was the highest, therefore "tension" was selected as the emotion category for this assessment. It is worth noting that the inputs from breathing pauses and rhythmic harmonics mutually corroborated the "tension" assessment, resulting in a higher weighted average. Higher than the weighted average The method further improves the robustness of this confirmation.

[0119] S4. In response to the fact that the dose level of the vortex ring has reached the preset dose limit and the confidence level of respiratory arrest is still lower than the preset effective detection threshold, mark the current state as a recognition blind zone.

[0120] Specifically, the method identifies blind zone states using a separate status flag, which is connected to an external application system via the device's output interface. When the external application system detects that this flag is set, it should be considered that the method cannot provide a reliable emotion category in this confirmation. Based on this, the external application system can choose not to update the user's emotion record, not to trigger downstream decisions that depend on the emotion category, or to explicitly prompt the user that "the current emotion cannot be confirmed," rather than using a low-confidence result of the method as a reliable result.

[0121] The determination path for entering the recognition blind zone state is as follows: In S2, the method sequentially fires according to a dose gradient increasing strategy. to Each vortex ring undergoes a dual-window comparison test after each emission; when the current dose level of the vortex ring has reached the dose limit... Furthermore, the confidence level of the breathing pause output after this file is launched. If the reading is still below the preset effective detection threshold (typically 0.55), the method determines that the active probe has exhausted all dose levels without obtaining an effective response, and sets the status flag to the recognition blind zone state. In the aforementioned low-light, head-down operating conditions, for example, if a user being identified provides different dose levels at four different dose levels... The values ​​are 0.30, 0.42, 0.48 and 0.51 respectively. Since all of them are below the effective detection threshold of 0.55, the system enters the recognition blind zone after the fourth transmission.

[0122] In some embodiments, the method further includes: in response to the dose level of the vortex ring reaching the upper dose limit and the confidence level of respiratory arrest still being lower than the effective detection threshold, reading the rhythm harmonic confidence level, and in response to the rhythm harmonic confidence level being higher than a preset blind zone effective threshold, outputting the emotion category corresponding to the rhythm harmonic confidence level as a weakly confirmed emotion. Specifically, after entering the recognition blind zone state, the method first attempts to use rhythm harmonic fallback: reading the rhythm harmonic confidence level extracted from the response window at the highest dose level during this trial. ,when When the effective blind zone threshold is higher than the preset threshold (typically 0.70), the method is as follows: The corresponding emotion category is used as the weakly confirmed emotion output. For example, under the aforementioned operating condition, if the emotion extracted after the fourth-level vortex ring is emitted... If the value is 0.75 and corresponds to "relaxed", the method outputs "relaxed" as the weakly confirmed emotion category and adds a "weakly confirmed" sub-label to the status flag, indicating to the external application system that the confidence level of the result is lower than that of the normal weighted fusion result.

[0123] In other embodiments, the method further includes: responding to the dose level reaching the upper limit, the confidence level of respiratory arrest being lower than the effective detection threshold, and the confidence level of rhythm harmonics being lower than the effective threshold of the blind zone, entering an extended passive observation period, during which active probing is stopped and S1 continues to be executed, waiting for environmental conditions to improve to the point that the overall confidence level is higher than the dynamic trigger threshold before resuming normal confirmation output. Specifically, when even the rhythm harmonics fallback fails to provide a weak confirmation sentiment, i.e. The value is also below the effective threshold of the blind zone. The method does not immediately provide any emotion category, but instead enters an extended passive observation period. The default duration of this extended passive observation period is typically 5 minutes. During this period, the method continuously monitors the identified user using S1's multimodal passive perception, but no longer emits new vortex rings to avoid ineffective attempts if the user has already failed to respond to the aforementioned four vortex ring levels. When the output ρ of S1 rises above the dynamic trigger threshold during the observation period—for example, when the identified user changes posture, looks up directly at the image acquisition module, or ambient lighting recovers—the method terminates the extended passive observation period and uses the output of this S1 signal to confirm that the emotion category has returned to normal output.

[0124] In some embodiments, the method further includes: during the identification blind zone period, if the overall confidence level of passive observation continues to be higher than the dynamic trigger threshold for a preset cancellation time, the previously marked identification blind zone state is reversed and cancelled, and the overall emotion category of passive observation is used as a weak confirmation emotion output. Specifically, the reverse cancellation mechanism is a further reinforcement to extend the passive observation period. When the method is in the identification blind zone state and the cumulative time for the output ρ of S1 to be higher than the dynamic trigger threshold reaches the cancellation time (typically 30 seconds), the method determines that the state of the identified user has recovered to the condition where S1 can independently give a credible result, clears the previous identification blind zone state flag, and uses the overall emotion category output by S1 within the cancellation time as a weak confirmation emotion output.

[0125] The three-layer fallback mechanism is activated sequentially according to priority: the first layer is rhythmic harmonic fallback, which is attempted immediately after S2 has exhausted the upper dose limit, and can directly provide weak confirmatory emotions before the vortex stimulation has subsided; the second layer is extended passive observation, which is entered when rhythmic harmonics also fail to provide effective results, waiting for environmental conditions to improve naturally; the third layer is reverse checkpointing, which reverses the blind zone state when a significant recovery of ρ does occur during the passive observation period. The introduction of the three-layer mechanism transforms the identification blind zone state from a terminating "no solution" marker to a reversible intermediate state, improving the overall coverage of the method in providing credible confirmatory emotions across different operating conditions.

[0126] Reference Figure 3 This application also provides an emotional state confirmation system based on trace odor probing, which includes:

[0127] An image acquisition module, configured to acquire facial image frames;

[0128] An audio acquisition module, configured to acquire respiratory response audio and respiratory baseline audio;

[0129] The emotion recognition module is configured to perform multimodal feature fusion processing on facial image frames and output a comprehensive emotion category and a comprehensive confidence score.

[0130] The trace odor detection module includes a trace odor chamber and a vortex ring forming mechanism. The trace odor chamber is divided into multiple independent sub-cavities, each of which carries a trace amount of stimulating odor. The vortex ring forming mechanism is configured to emit a vortex ring carrying a trace amount of stimulating odor to the identified user in response to a transmission command.

[0131] The respiratory feature analysis module is configured to perform a dual-window comparison detection of the respiratory response audio and the respiratory baseline audio within a predetermined time window when the vortex ring arrives at the expected time, and output the confidence level of respiratory pause.

[0132] The main control module is configured to: send a transmission command to the trace odor detection module in response to the overall confidence level being lower than a preset dynamic trigger threshold; perform weighted fusion of the overall confidence level and the confidence level of respiratory arrest to output the confirmed emotion category; and mark the current state as a recognition blind zone in response to the dose level of the vortex ring reaching the preset dose upper limit and the confidence level of respiratory arrest still being lower than the preset effective detection threshold.

[0133] The micro-odor chamber inside the device is a cylindrical or rectangular cavity structure, divided into multiple independent sub-cavities along the axial direction by several rigid partitions. Each independent sub-cavity carries a trace amount of a stimulating odor, such as several mild odors that are perceived as neutral by the user and do not easily evoke strong emotional associations. Adjacent sub-cavities are physically isolated by partitions to prevent cross-contamination between different odors. Each independent sub-cavity is connected to the mixing chamber at the rear of the vortex ring forming mechanism via an independent gas path pipe. This connection can be made by controlling the opening and closing of the solenoid valve on the corresponding gas path pipe to select the type of odor carried by the vortex ring in this instance.

[0134] The vortex ring forming mechanism is located at the nozzle on the side of the device facing the identified user, and consists of a circular nozzle and a drive unit located behind it. In one engineering implementation, the drive unit is an elastic diaphragm in conjunction with an electromagnetic coil: after receiving a transmission command from the main control module, the electromagnetic coil generates an instantaneous magnetic field, attracting the elastic diaphragm and causing instantaneous displacement, rapidly pushing the gas in the mixing chamber out through the nozzle to form a ring vortex. In another engineering implementation, the drive unit is a piston in conjunction with a servo motor: the piston, driven by the servo motor, completes rapid feeding, pushing the gas in the mixing chamber out. In yet another engineering implementation, the drive unit is a piezoelectric ceramic oscillator: the piezoelectric ceramic oscillator undergoes rapid deformation after receiving a driving voltage pulse, with a similar mechanism to the diaphragm drive but a faster response speed. All three engineering implementations can form a relatively stable ring vortex; the main differences lie in cost, energy consumption, and pulse response time.

[0135] The signal connection relationships between the modules are as follows: Facial image frames output by the image acquisition module are sent to the emotion recognition module in the form of a video stream; the comprehensive emotion category and comprehensive confidence level output by the emotion recognition module are sent to the main control module in the form of a data stream; the respiratory baseline audio and respiratory response audio output by the audio acquisition module are sent to the respiratory feature analysis module in the form of an audio stream; the respiratory pause confidence level output by the respiratory feature analysis module is sent to the main control module in the form of a data stream; when the main control module determines that a vortex ring needs to be emitted, it sends an emission command to the trace odor detection module to trigger the vortex ring forming mechanism to act in the form of a control stream; after integrating the data streams from all sources, the main control module outputs a confirmation of the emotion category or a blind zone status flag to the device's external output interface.

[0136] In some embodiments, the logical functions of the aforementioned image acquisition module, audio acquisition module, emotion recognition module, respiratory feature analysis module, and main control module can be undertaken by the same embedded processing unit, with only the camera corresponding to the image acquisition module, the microphone corresponding to the audio acquisition module, and the trace odor detection module existing as independent hardware. In other embodiments, the logical functions of the emotion recognition module and the respiratory feature analysis module are undertaken by a supporting server, and the local device only undertakes data acquisition and vortex ring forming actions.

[0137] Reference Figure 4 This application also provides a computer device comprising: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to: perform the above-described method for confirming emotional state based on trace odor probing.

[0138] The processor can be a general-purpose central processing unit, an embedded microcontroller unit, a digital signal processor, a graphics processing unit, or a field-programmable gate array (FPGA), or other processing units that can execute machine instructions. The memory can be a storage medium that can store machine instructions and data, such as random access memory, read-only memory, flash memory, or a solid-state drive. One or more application programs can be stored in the memory in any form, such as executable binary code, bytecode, or interpreted scripts, and are loaded and executed by the processor to implement all the steps of the above method.

[0139] This application also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the above-described method for confirming emotional states based on trace odor probing.

[0140] The computer-readable storage medium can be any specific implementation of the aforementioned memory, or it can be other storage media that can be read by a processor, such as optical discs, magnetic tapes, or disks. Those skilled in the art will understand that all or part of the steps of the above method can be implemented by corresponding program instructions, which can be stored in one or more computer-readable storage media.

[0141] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0142] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for confirming emotional state based on trace odor probing, characterized in that, Includes the following steps: S1. Perform multimodal feature fusion processing on the acquired facial image frames and output the comprehensive emotion category and comprehensive confidence score; S2. In response to the overall confidence level being lower than a preset dynamic trigger threshold, a vortex ring carrying a trace amount of irritant odor is emitted to the identified user, and within a predetermined time window when the vortex ring arrives at the expected time, a dual-window comparison detection is performed on the breathing response audio of the identified user and the breathing baseline audio before the vortex ring is emitted, and the breathing pause confidence level is output; wherein, the trace amount of irritant odor is an irritant gas with a concentration lower than a preset concentration threshold; S3. Perform a weighted fusion of the overall confidence level and the breathing pause confidence level, and output the confirmed emotion category; S4. In response to the fact that the dose level of the vortex ring has reached the preset dose upper limit and the confidence level of the respiratory arrest is still lower than the preset effective detection threshold, the current state is marked as a blind zone.

2. The method for confirming emotional state based on trace odor probing according to claim 1, characterized in that, The emission of the vortex ring in S2 is performed according to a dose-incremental strategy, which includes: The dose levels of the vortex ring are set sequentially according to a preset dose sequence; After each launch of the vortex ring, the dual-window comparison detection is performed. In response to the fact that the confidence level of the respiratory arrest is still lower than the effective detection threshold and the current dose level of the vortex ring has not reached the dose upper limit, the dose level of the vortex ring is increased to the next level in the dose sequence, and the launch of the vortex ring is re-executed according to the increased dose level.

3. The method for confirming emotional state based on trace odor probing according to claim 1, characterized in that, The method further includes: An individual response baseline database is established for the identified user. The individual response baseline database stores multiple quadruple records. Each quadruple record includes odor type, dose level, response delay and pause amplitude. In response to the fact that the confidence level of the respiratory pause after the launch of the vortex ring is higher than the effective detection threshold, the odor type, the dose level, and the corresponding response delay and pause amplitude used in this instance are written into the individual response baseline library as new quadruplet records; Before the next execution of S2, based on the existing quadruple records in the individual response baseline library, the odor type and starting dose level of the trace stimulating odor carried by the vortex ring in the next execution are preferably determined according to the normalized ratio of the pause amplitude to the dose level.

4. The method for confirming emotional state based on trace odor probing according to claim 1, characterized in that, The method further includes: In the predetermined time window when the vortex ring reaches the expected time in S2, the respiratory rhythm harmonic component is extracted from the respiratory response audio of the identified user, and the respiratory rhythm harmonic component is matched with the preset emotional state harmonic prior to output the rhythm harmonic confidence. The input to the weighted fusion also includes the rhythm harmonic confidence level, which constitutes a three-way weighted fusion of the overall confidence level, the respiratory arrest confidence level, and the rhythm harmonic confidence level.

5. The method for confirming emotional state based on trace odor probing according to claim 1, characterized in that: The dynamic trigger threshold is dynamically adjusted based on the current ambient light intensity, the facial orientation angle of the identified user, and the distance between the identified user and the image acquisition module that acquires the facial image frame. The execution of S2 also satisfies the following timing conditions: the facial orientation angle of the identified user is within a preset effective orientation range; the noise intensity of the current environment is lower than a preset environmental noise threshold; the interval since the last emission of the vortex ring exceeds a preset cooling time; and the identified user is unique.

6. The method for confirming emotional state based on trace odor probing according to claim 5, characterized in that, In response to the presence of multiple candidate users in the facial image frame, the candidate user whose face orientation angle is within the effective orientation range and is closest to the image acquisition module is identified as the identified user, and S2 is performed only on the identified user.

7. The method for confirming emotional state based on trace odor probing according to claim 3, characterized in that, The method further includes: In the individual response baseline database, additional respiratory baseline fingerprint records are stored for the identified user. The respiratory baseline fingerprint records include resting respiratory rate, respiratory amplitude distribution, and respiratory rhythm harmonic structure. Before performing the dual-window comparison detection, the respiratory activity variance is calculated on the respiratory baseline audio. In response to the respiratory activity variance exceeding a preset variance threshold, the respiratory baseline fingerprint record of the identified user is used instead of the respiratory baseline audio as the baseline reference for the dual-window comparison detection.

8. An emotional state confirmation system based on trace odor probing, characterized in that, It includes: An image acquisition module, configured to acquire facial image frames; An audio acquisition module, configured to acquire respiratory response audio and respiratory baseline audio; An emotion recognition module is configured to perform multimodal feature fusion processing on the facial image frames and output a comprehensive emotion category and a comprehensive confidence level. A trace odor detection module includes a trace odor chamber and a vortex ring forming mechanism. The trace odor chamber is divided into multiple independent sub-cavities, each of which carries a trace amount of stimulating odor. The vortex ring forming mechanism is configured to emit a vortex ring carrying the trace amount of stimulating odor to the identified user in response to a transmission command. The respiratory feature analysis module is configured to perform a dual-window comparison detection of the respiratory response audio and the respiratory baseline audio within a predetermined time window when the vortex ring reaches the expected time, and output the confidence level of respiratory pause. The main control module is configured to: in response to the overall confidence level being lower than a preset dynamic trigger threshold, send the emission command to the trace odor detection module; Perform a weighted fusion of the overall confidence level and the breathing pause confidence level to output the confirmed emotion category; And in response to the fact that the dose level of the vortex ring has reached the preset dose upper limit and the confidence level of the respiratory arrest is still lower than the preset effective detection threshold, the current state is marked as a blind zone.

9. A computer device, characterized in that, It includes: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to: perform the emotional state confirmation method based on trace odor probing according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or instruction set is loaded and executed by a processor to implement: the emotional state confirmation method based on trace odor probing according to any one of claims 1 to 7.