Method of fine-grained respiratory interaction based on visual and auditory detection

By combining visual and auditory breathing detection technology to identify breathing direction, intensity, and type, the BreathUI breathing interaction system was designed. This solves the problem that existing breathing interaction technologies fail to fully utilize multiple parameters, and achieves efficient interactive control in VR scenarios.

CN116048270BActive Publication Date: 2026-03-31NANJING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-29
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing respiratory interaction technologies mainly focus on auxiliary input channels, failing to fully utilize various respiratory parameters for complex interactive operations, and lack fine-grained respiratory direction detection based on facial expressions and audio information.

Method used

By combining visual and auditory detection, utilizing the HTC Vive face tracker and lavalier microphone, the direction, intensity, and type of breathing are identified. Breath vectors are designed to represent breathing interaction actions, and the BreathUI breathing interaction system is constructed to achieve dual-channel detection of breathing as an independent interactive input channel.

Benefits of technology

It enables fine-grained breathing interaction in VR scenarios, enhances the controllability of the interactive interface, verifies the feasibility of using breathing as the main interactive input channel, and improves the efficiency of user interface operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116048270B_ABST
    Figure CN116048270B_ABST
Patent Text Reader

Abstract

The application discloses a method for fine-grained respiratory interaction based on visual and auditory detection, comprising the following steps: acquiring respiratory parameters based on respiratory types, and constructing a respiratory input method of an interactive vocabulary; based on the respiratory input method of the interactive vocabulary, designing several characteristic parameters of a respiratory vector representing respiratory interaction actions, and acquiring respiratory intensity, respiratory direction, whether inhaling, respiratory time, respiratory frequency, expression information and volume information from the several characteristic parameters; constructing a visual and auditory double-channel detection model, inputting the expression information and the volume information, and completing respiratory detection. The application realizes the recognition of multiple respiratory action characteristic parameters including respiratory action airflow intensity, respiratory airflow direction, respiratory airflow type and the like based on the combination of images and voices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of respiratory detection technology, and particularly relates to a method for fine-grained respiratory interaction based on visual and auditory detection. Background Technology

[0002] Breathing is one of the most natural physiological activities in humans, and it is considered an alternative control mechanism influencing the physical world and virtual environments. However, current interaction methods in VR scenarios mainly focus on eye contact, gestures, and voice interaction. Existing limited research on breathing often treats it as a physiological state indicator, a factor influencing user experience during VR game interaction, or uses specific breathing movements as supplementary input channels to gestures and other input methods, with blowing air as a "select and confirm" input action. Its potential as a standalone input channel has not been fully explored. With the development of sensor technology, it has become possible to further identify parameters such as the type (blowing or inhaling), strength, and direction of active breathing movements. The introduction of more breathing parameters may further enhance the interactive input capabilities of the breathing interaction channel, enabling it to meet the needs of complex user interface operations. Furthermore, for specific scenarios or special groups who find gesture interaction inconvenient, breathing interaction may become a more effective interactive input method.

[0003] By using the muscles around the mouth, humans can control parameters such as the direction, speed, and duration of airflow during exhalation or inhalation to some extent. Current research on interactive breathing applications primarily relies on sensors such as microphones, vibration sensors, and cameras to obtain information like breathing sound intensity, breathing frequency, and facial orientation. Parameters such as the direction of airflow are difficult to obtain directly using a single type of optical or vibration sensor, and therefore, relevant research remains lacking.

[0004] By combining facial expression recognition technology with audio recognition technology, it is possible to monitor the fine movements of a user's mouth and the acoustic characteristics of breathing airflow. This allows for the analysis of more precise breathing airflow direction and breathing state parameters, further increasing the input vocabulary space for breathing movements and thus enhancing the user interface's control capabilities.

[0005] This paper proposes a dual-channel (visual and auditory) breathing detection technology that combines facial and audio information during the breathing process. A facial tracker is used to capture facial expressions during breathing to determine the user's breathing direction and state. Simultaneously, a microphone is used to capture audio information during breathing to detect breathing intensity, and this information is combined with facial recognition to determine breathing direction and state. Based on these detected parameters, two breathing interaction application systems were designed, using breathing as an independent interactive input channel, demonstrating the feasibility of breathing as an independent interactive input method. Summary of the Invention

[0006] The purpose of this invention is to propose a fine-grained breathing interaction method based on visual and auditory detection. By combining images and voice, it realizes the recognition of various breathing action feature parameters, including breathing airflow intensity, breathing airflow direction, and breathing airflow type. By comprehensively utilizing various breathing action feature parameters, it realizes breathing interaction control of typical interactive interfaces in VR scenes.

[0007] To achieve the above objectives, the present invention provides a method for fine-grained respiratory interaction based on visual and auditory detection, comprising the following steps:

[0008] A breathing input method for interactive vocabulary is constructed by obtaining breathing parameters based on breathing type.

[0009] Based on the breathing input method of the interactive vocabulary, several feature parameters of breathing vectors are designed to represent breathing interactive actions, and the breathing intensity, breathing direction, whether inhalation occurs, breathing time, breathing frequency, facial expression information and volume information are obtained from the several feature parameters.

[0010] A dual-channel detection model combining visual and auditory inputs is constructed to perform breathing detection.

[0011] Optionally, the breathing type includes exhalation and inhalation. When a person exhales, the intensity of the exhalation varies, the duration of the breath varies, and the direction of the exhalation is diverse. The exhalation frequency is calculated by the number of exhalations over a period of time. When a person inhales, the intensity of the inhalation, the duration of the breath, and the frequency of the breath also change.

[0012] Optionally, the design of several feature parameters representing respiratory interaction actions using a respiratory vector, and the method for obtaining respiratory intensity, respiratory direction, inhalation status, respiratory time, and respiratory frequency from these feature parameters, includes: defining information during the respiratory process, with respiratory intensity as variable m, where m is a continuous analog quantity; respiratory direction as variable d, where d is a discrete directional variable; inhalation status as variable i, where i is a boolean variable; respiratory time as variable t; and respiratory frequency as variable f; the input variables for the dual-channel detection method are facial expression information e and volume information v, respectively.

[0013] Optionally, the breathing vector transforms the user interface interaction design process into a process of mapping breathing vectors to user interface operation action word vectors.

[0014] Optionally, the breathing vector includes a breathing direction component, a breathing intensity component, a breathing time, and an inspiratory state; the breathing direction component is mapped to the directional operation component of the user interface and is used for directional selection of the pie menu; the breathing intensity component is mapped to the speed change control component of the user interface action word vector and is used for the scrolling speed of the scroll menu; the breathing time is used for zooming in or out of interface elements; and the inspiratory state is used as a trigger switch for a specific interface input state.

[0015] Optionally, a dual-channel detection model combining visual and auditory signals can be constructed. The method for completing breath detection, by inputting the facial expression information and the volume information, includes:

[0016] First, design an interactive system called BreathUI that uses breathing gestures as the interactive input method;

[0017] A dual-channel detection model for vision and hearing is constructed. The facial expression information and the volume information are input. A lavalier microphone is used to receive audio information. The HTC Vive face tracker is used to detect changes in facial expressions. Facial movements are captured through 38 tracking points on the lips, chin, teeth, tongue and cheeks.

[0018] Combining the aforementioned situational information and volume information, the BreathUI interactive system detects the user's breathing type and determines whether the user is exhaling or inhaling.

[0019] When inhalation is detected, the system also identifies the inhalation intensity. When exhalation is detected, the system also needs to determine the user's current breathing direction and detect the exhalation intensity.

[0020] Optionally, the breathing type determination includes: using the indentation of the cheeks during breathing as the basis for determining the breathing type, and combining the microphone volume to detect the occurrence of breathing actions. When the volume reaches a preset threshold, it will be determined that the user is breathing at this time; after adjusting the volume threshold for breathing determination in each direction, the breathing direction is mapped to the interactive interface to replace the traditional directional operation component.

[0021] Optionally, breathing direction recognition specifically includes: when the user makes blowing motions in different directions, facialTracker will return the values ​​of the corresponding tracking points on the face, classify them based on the relevant values, and determine the user's blowing direction.

[0022] Optionally, breathing intensity detection specifically includes: using a microphone to receive volume information when the user blows or inhales. The volume information not only determines whether the user is blowing or inhaling at this time, but also maps it to breathing intensity. BreathUI uses changes in breathing intensity values ​​to control menu scrolling speed and interface scaling.

[0023] Technical Effects of this Invention: This invention discloses a fine-grained breathing interaction method based on visual and auditory detection. It proposes using breathing as an independent interaction channel, realizing a dual-channel breathing detection technology that integrates auditory and visual perception. Based on a combination of image and voice, it achieves the recognition of various breathing action feature parameters, including breathing airflow intensity, breathing airflow direction, and breathing airflow type. This invention designs a user interface system that uses breathing actions as the interactive input means, comprehensively utilizing various feature parameters of breathing actions to realize breathing interaction control in typical interactive interfaces in VR scenarios. Based on the user interface system, user experiments were conducted to study the effectiveness of using breathing airflow intensity, breathing airflow direction, and breathing airflow type for user interface operation, verifying the feasibility of using breathing actions as the main interactive input channel, and exploring possible future applications of breathing interaction. Attached Figure Description

[0024] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0025] Figure 1 This is a schematic diagram illustrating the vector description of respiratory characteristic parameters in an embodiment of the present invention:

[0026] Figure 2 This is a schematic diagram of the detection process of BreathUI according to an embodiment of the present invention;

[0027] Figure 3 This is a schematic diagram illustrating user breathing direction recognition in an embodiment of the present invention;

[0028] Figure 4 The system provides a schematic diagram of the target ball on the right (b) and the user blowing the ball on the right (a);

[0029] Figure 5 This is a schematic diagram showing the accuracy (left) and number of breaths (right) for each direction in the preliminary experiment;

[0030] Figure 6 The image shows the "1 / 2" option for the pie menu.

[0031] Figure 7 A diagram illustrating the use of the breathing control sliding menu;

[0032] Figure 8 The quantitative results of User Experiment 1 are illustrated with the average task completion time (left) and average error rate (right).

[0033] Figure 9 Quantitative results for User Experiment 2: Schematic diagram of average task completion time (left) and average number of menu scrolls (right);

[0034] Figure 10 The results of the questionnaire interviews for User Experiment 1;

[0035] Figure 11 The results of the questionnaire interviews for User Experiment 2;

[0036] Figure 12 This is a flowchart illustrating a method for fine-grained respiratory interaction based on visual and auditory detection according to an embodiment of the present invention. Detailed Implementation

[0037] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0038] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0039] like Figure 1-12 As shown, this embodiment provides a method for fine-grained breathing interaction based on visual and auditory detection, including the following steps:

[0040] A breathing input method for interactive vocabulary is constructed by obtaining breathing parameters based on breathing type.

[0041] Based on the breathing input method of the interactive vocabulary, several feature parameters of breathing vectors are designed to represent breathing interactive actions, and the breathing intensity, breathing direction, whether inhalation occurs, breathing time, breathing frequency, facial expression information and volume information are obtained from the several feature parameters.

[0042] A dual-channel detection model combining visual and auditory inputs is constructed to perform breathing detection.

[0043] Research on breathing in VR interaction scenarios can be divided into two main categories: treating it as a physiological state indicator parameter or an environmental influencing factor during the interaction process.

[0044] Most researchers use breathing as a physiological state indicator, identifying VR users' breathing parameters through sensors, breathing monitoring machines, or microphones. These parameters are then used to study the impact of other user interactions in VR on their physiological state. For example, Marieke van Rooij et al. monitored children's breathing rate and other physiological information to help assess treatment effectiveness when using VR games to treat childhood anxiety. Adler et al. studied visual-breathing conflict in VR, demonstrating that body ownership illusion (BOI) affects respiratory perception.

[0045] A smaller number of researchers treat breathing as an environmental influencing factor, simulating its impact on the real environment within VR scenarios to enhance the game's fun or realism. For example, Joan Sol Roo et al. developed the Inner Garden VR-assisted mindfulness system, where the sea level in the scene rises or falls with the user's inhalation and exhalation. Lorian Soyka et al. developed a VR underwater world where users adjust their breathing rhythm to follow the rise and fall of jellyfish. Martijn JLKors et al. designed a mixed reality game called ABTJ, where players take on the role of a refugee fleeing a war-torn country, equipped with a mask containing an odor diffuser and breathing sensors to provide an olfactory sensory experience and determine the user's success in stealth.

[0046] In non-VR scenarios, a few studies use single breathing parameters (such as whether or not to exhale) as additional input channels for other input methods such as gestures. For example, some literature uses "whether or not to exhale" as the input action for "click to select". Currently, there is a lack of research on interface interaction technology that comprehensively utilizes various parameters such as breathing frequency, intensity, and direction to complete complex interactive tasks, and that focuses on breathing interaction.

[0047] Research on respiratory interaction parameters has largely focused on respiratory intensity, duration, and frequency. Sra et al. differentiated five types of exhalation based on duration, intensity, and frequency, mapping these patterns into game controls. Joe Marshall et al. controlled the rotation of a riding device through breathing; inhalation caused clockwise rotation, and exhalation caused counter-clockwise rotation, with the rotation speed determined by the breathing rate. Harris J et al. designed a system that promotes slow breathing through auditory feedback; the closer the user's breathing rate is to the target rate, the higher the quality of the music played. Bingham et al. developed a video game to aid in the treatment of cystic fibrosis, where users control a green circle through breathing; inhalation lowers it, and exhalation raises it, with the breathing flow affecting the distance it moves.

[0048] Regarding the parameter of breathing direction, current research is not yet sufficient. Kim JH and Lee J used the position and angle between the user and the mobile device to calculate the breathing direction. Shwetak N. Patel and Gregory D. Abowd used the microphone built into the laptop to identify the position of the user blowing towards the computer screen; a short, forceful breath towards the target could act as a "click to select." However, this blowing direction is based on changes in head rotation, and there is currently no research on blowing direction based on mouth movements and facial expressions.

[0049] Currently, traditional breathing recognition methods mostly rely on vibration sensors to identify parameters such as the intensity and duration of breathing. Koichi Kuzume et al. used a piezo film sensor array to detect exhalation signals and a bone conduction microphone to detect sound signals from tooth contact, creating a hands-free interface for a disabled input device using exhalation and tooth contact sound signals. John Desnoyers-Stewart et al. used a Biosignalsplux piezoelectric breathing sensor around the diaphragm to detect chest and abdominal bulges, thereby controlling the size of a virtual jellyfish. Misha Sra et al. detected breathing parameters by wearing a Zephyr biostrap breathing detection band. Markus Tatzgern et al. enhanced the virtual experience with their self-developed AirRes mask, potentially creating more immersive scenarios for applications by enhancing hazard perception or improving situational awareness in training simulations, or providing additional physical stimulation for psychological therapy. Corbishley et al. proposed a miniature breathing detection system that uses a microphone placed in the neck to listen to the sound signals of airflow. Georgia Institute of Technology in the United States has developed an RVSM (Rader sensing of heartbeat and respiration) system that can monitor human heartbeat and respiration signals from a distance of more than 10 meters.

[0050] In addition to traditional vibration sensor-based respiration detection, D. Shao et al. explored an optical camera-based detection method. Since the shoulders undergo slight vertical movement during respiration, they selected a small area on the shoulder image and analyzed its movement over time. This allows them to monitor important physiological signals such as heart rate, pulse transit time, and respiratory patterns for disease diagnosis and treatment. However, current optical respiration detection methods lack research on detecting airflow direction during respiration based on facial recognition.

[0051] APPROACH

[0052] Design Space

[0053] Previous studies on respiratory interaction mainly used a single parameter of respiratory action as an auxiliary input channel, without fully exploring the ability of breathing as an independent interactive input channel.

[0054] By comprehensively utilizing parameters such as the ray direction, planar direction, breathing intensity, and duration of breathing, a breathing input method with rich interactive vocabulary can be constructed.

[0055] Firstly, breathing can be divided into two types: exhalation and inhalation. When a person exhales, the intensity of the exhalation varies, the duration of the breath varies, and the direction of the exhalation can also be diverse. Furthermore, the exhalation frequency can be calculated by counting the number of exhalations over a period of time. Conversely, when a person inhales, the intensity, duration, and frequency also vary, just like during exhalation.

[0056] To more clearly describe the various information detection methods and applications in the breathing interaction process, a breathing vector was designed to characterize various feature parameters of the breathing interaction action. First, the information in the breathing process is defined as follows: breathing magnitude is a variable *m*, which is a continuous analog quantity; breathing direction is a variable *d*, which is a discrete directional variable; inhalation is a variable *i*, which is a boolean variable; breathing time is a variable *t*; and breathing frequency is a variable *f*. Second, the dual-channel detection method requires input variables of facial expression information *e* and volume information *v*.

[0057] like Figure 1As shown, a dual-channel detection method using facial expression and volume information can detect information such as the intensity and direction of breathing. This information is integrated into a one-dimensional vector, forming a breathing vocabulary space, which can be used to execute corresponding interactive operations. Each breathing process can thus be described in vector form, facilitating subsequent feature description and data processing of breathing interaction actions.

[0058] BreathUI

[0059] Based on the description method of breathing vectors, the interaction design process of user interface can be transformed into a process of mapping breathing vectors to user interface action word vectors. The breathing direction component can be mapped to the directional operation component of the user interface, used for directional selection in a pie menu. The breathing intensity component can be mapped to the velocity change control component of the user interface action word vector, used for the scrolling speed of a scroll menu. Breathing time can be used for zooming in or out of interface elements. Inhalation state can be used as a trigger switch for specific interface input states, and so on.

[0060] Based on this approach, an interactive system called BreathUI was designed, using breathing movements as the primary input method. BreathUI is a user interface based on a breathing interaction system, allowing users to perform traditional interface operations such as menu selection simply by breathing. Its main breathing interaction recognition and processing process is as follows: Figure 2 As shown. After a user takes a breath, the breathing interaction system first detects whether the user's breathing state is exhalation or inhalation. When the system detects inhalation, it also identifies the inhalation intensity at that time. When it detects exhalation, the system also needs to determine the user's current breathing direction and detect the exhalation intensity at that time.

[0061] Respiratory Detection Based on Vision and Speech

[0062] This section will detail VisaudiBreath, a dual-channel breathing detection method used to identify the direction, intensity, and type of breathing. It uses a lavalier microphone to receive audio information and an HTC Vive facial tracker to detect changes in facial expressions, capturing facial movements through 38 tracking points on the lips, chin, teeth, tongue, and cheeks.

[0063] Respiratory type determination

[0064] In this invention, breathing is classified into two types: exhalation and inhalation, based on the overall gas flow direction of the active breathing action. Existing works mostly rely on audio information to determine the breathing type, but the sound during inhalation is relatively quieter than that during exhalation, increasing the difficulty of detection. BreathUI primarily uses the indentation of the cheeks during breathing as the criterion for determining the breathing type. This allows for the differentiation of the two types of breathing even in noisy environments, thus achieving better robustness.

[0065] However, facial expressions may not necessarily be accompanied by actual breathing. Considering that inhalation and exhalation usually produce sound, the system uses microphone volume to detect breathing movements. When the volume reaches a certain threshold, the system determines that the user is breathing. By adjusting the volume thresholds for breathing detection in different directions, the breathing direction can be mapped to components in the user interface, replacing traditional directional controls.

[0066] Recognition of breathing direction

[0067] During inhalation or exhalation, the airflow primarily occurs within a cone-shaped region in front of the mouth. The direction of airflow within this cone can be controlled to some extent by shaping the mouth. Current research on respiratory airflow direction largely uses the area directly in front of the mouth as a fixed direction for blowing or inhaling; in reality, the head orientation is used to represent the airflow direction of the breathing action, a method known as head-based breathing. Based on this method, interactive operations require multiple head rotations within the VR environment. Furthermore, it is only suitable for scenarios where the interactive object is fixed in the VR environment and does not move with the head.

[0068] This invention identifies breathing direction based on visual and auditory features, enabling finer-grained recognition of breathing direction and significantly increasing the lexical space for breathing interactions. This is termed facial expression-based breathing direction recognition.

[0069] like Figure 3 As shown, when a user makes blowing motions in different directions, the facial tracker returns the values ​​of the corresponding tracking points on the face. Based on the relevant values, the direction of the user's blowing can be determined.

[0070] Breathing intensity test

[0071] Breath intensity can naturally be used as a continuous, interactive analog input. A microphone is used to receive the volume information of the user's exhalation or inhalation; this volume information not only determines whether the user is exhaling or inhaling, but also maps it to breathing intensity. BreathUI uses changes in breathing intensity values ​​to control menu scrolling speed and interface scaling.

[0072] User Experiment

[0073] This chapter validates the breathing detection technology based on facial and auditory detection and explores the ability to utilize breathUI as an input mode in VR applications. Specifically, it examines the ability to apply breathing direction, intensity, and type to menu controls, as well as the interaction capabilities with different types of menus.

[0074] This experiment used the HTC Vive Pro Eye HMD as the VR experimental device. Facial expression recognition was based on the HTC Vive Facial Tracker. A Sudotack lavalier microphone was used to capture breathing sounds.

[0075] The participants were 12 college students on campus. They had an average of 0.8 years of experience using VR devices, ranging from 0 to 4 years, but none of them had experience with breathing interactions in VR.

[0076] Two commonly used menu formats, pie menus and slider menus, were set up to explore the controllability of breathing interaction on different menu types. Breathing control on a pie menu was compared with control using a joystick; this constitutes User Experiment 1. Prior to User Experiment 1, a simple pre-experiment was conducted to verify the performance of the proposed detection method in determining breathing direction.

[0077] The 12 participants were randomly divided into four groups, and the experimental order was determined using a Latin square experimental design, with random balance among the four groups.

[0078] In the preliminary experiment, participants were instructed to hit eight targets as quickly and accurately as possible in different directions (i.e., up, down, left, right, upper left, upper right, lower left, and lower right). The system would randomly display any target in the eight directions, shown in red (b). Within a given time (6 seconds), participants could blow several times. When a target was hit, the ball turned green, and a new target appeared immediately (a). Otherwise, a new target would appear after 6 seconds. Each participant had eight tasks in eight different directions. Targets in each direction would appear randomly and only once. The number of hits and the hit rate for each task would be recorded.

[0079] Figure 5 (Left) shows the probability of participants hitting the target by blowing air. It can be seen that the hit rate in the up, down, left, and right directions is much higher than in the other four directions. Figure 5 (Right) shows the average number of breaths required for each direction. It can be seen that the left and right directions require the fewest breaths, averaging only about 1.5 breaths. The up and down directions almost always require a second breath. The remaining four directions almost always require more than a second breath.

[0080] This may be because people don't usually blow air at an angle in daily life, making it difficult to perform the action, leading to inaccurate facial expression recognition. Therefore, in the pie chart design for Experiment 1, only four directions were selected: up, down, left, and right.

[0081] This experiment aims to verify the controllability of breathing interaction over a pie chart. Menu selection is a fundamental input task in VR applications. Menus come in various types, and due to the unique directional characteristics of breathing interaction, a pie chart is well-suited as an interface for breathing interaction.

[0082] Based on the results of the preliminary experiment, the pie chart menu was designed with three layers, each with four sub-menus. The menu content was designed to resemble number classification tasks familiar to science and engineering students, such as... Figure 6 As shown.

[0083] When using the breathing interaction system, participants trigger the next menu level by blowing in different directions. The target may appear in any of the three levels. If the target menu is selected by blowing, the task is considered successful. If the selection is incorrect, it can be canceled by inhaling, returning to the previous menu. Each participant must complete eight menu selection tasks using the breathing interaction and joystick.

[0084] When using the joystick, participants aim at the target menu with a ray emitted from the top of the joystick, then pull the trigger button to select a menu item and cancel using a side button on the joystick. The time it takes for participants to complete the task and the number of invalid interactions are recorded.

[0085] In user experiment 2, two breathing-based sliding menu interaction methods were designed compared to traditional sliding menus: Gale and Gust. Gale refers to a strong, continuous blowing pattern, similar to blowing air when blowing a cake. Gust refers to a short but strong airflow from the mouth, similar to blowing dust off an object.

[0086] In GALE mode, the menu scrolling speed is proportional to the intensity of the breath. The stronger the breath, the faster the menu scrolls. Once the breathing stops, the menu stops scrolling. In GUST mode, the initial movement speed of the menu depends on the intensity of the breath, then slows down until it stops. Here, the sliding menu is designed in two layers. The first layer has 11 menu items, each corresponding to 10 menu items in the second layer. Figure 7 Similar to Experiment 1, users could scroll up and down the menu by blowing air in four directions (up, down, left, and right) to return to the previous or next level. Each participant was randomly assigned a task with five menu options.

[0087] Because the breathing interaction menu task was too unfamiliar to most participants, and to analyze how breathing interaction performance improved with proficiency, each experiment was conducted twice, with a two-day interval. In both experiments, the average time to complete the entire task was approximately 20 minutes (M1 = 20.67 minutes, σ1 = 9.60 minutes). After becoming more proficient, the average time was approximately 17.5 minutes (M2 = 17.75 minutes, σ2 = 8.08 minutes). It can be considered that the first experiment reflected performance during the initial experience with the breathing interaction, while the second experiment reflected performance after achieving relative proficiency.

[0088] Paired t-tests (P < .05) were used to analyze the mean completion time in the two experiments and the number of menu scrolls in Experiment 2. Wilcoxon rank-sum tests (P < .05) were used to analyze the error rate in Experiment 1.

[0089] Figure 8 Data from User Experiment 1 are shown. In Experiment 1, there was a significant difference in average task completion time between breathing interaction (m = 105.21 s, σ = 30.74 s) and joystick-based interaction (m = 72.67 s, σ = 14.61 s) (T(11) = 4.74, P < 0.005). Similarly, the error rate using breathing (m = 12.38%) was significantly higher than that using the joystick (m = 4.52%) (P < .05). In Experiment 2, breathing interaction significantly reduced task completion time (m = 68.11 s, σ = 9.70 s) and error rate (m = 5.36%). The completion time (m = 64.29 s, σ = 10.09 s) and error rate (m = 3.90%) using the joystick were almost the same, with little difference in time (T(11) = 2.10, P > .05) and error rate (P > .05).

[0090] It can be seen that, once mastered, the breathing interaction in this invention is comparable to traditional operating methods in both efficiency and accuracy.

[0091] Figure 9The average completion time and the average number of scrolls up and down the menu bar are shown in User Study 2. It can be observed that there are significant changes in both experiments under the two different scrolling modes, similar to User Study 1. Regarding task completion time, the Gale mode takes significantly less time than the Gust mode (m = 168.17s, σ = 60.96s) (T(11) = 5.012, P < 0.005) (m = 119.41s, σ = 52.07). The number of scrolls using the Gale mode (m = 31.42, σ = 10.66s) is also significantly less than the number of scrolls using the Gust mode (m = 44.83, σ = 11.25s) (T(11) = 5.90, P < 0.005).

[0092] The task completion time (m_gale = 70.49s, σ = 16.70s) and scroll count (m_gust = 124.09s, σ = 26.60) and scroll count (m_gale = 22.83, σ = 8.71) and (m_gust = 39.75, σ = 7.92) of both modes decreased as they became more familiar with the breathing interaction. However, there were significant differences between them (T_gale(11) = 6.99, P < .005) and (T_gust(11) = 6.58, P < .005). In summary, Gale performed better than Gust when interacting with the sliding menu. This is likely because in Gale mode, the stopping of menu scrolling is easier to control, rather than the unpredictable final cursor position as in Gust.

[0093] Following the second experiment, participants' subjective feelings about the breathing interaction design were investigated using a questionnaire. These questions are shown on the left. Each question has two bars, corresponding to the breathing interaction and the joystick interaction, respectively. Each bar indicates the Likert scale selection made by each participant for each question. The uncorrected p-value for each question is shown on the right. An asterisk (*) before p indicates statistical significance (p < 0.05). The Wilcoxon rank-sum test (P < .05) was used to analyze the questionnaire results; detailed data can be found in [link to data]. Figure 10 .

[0094] Regarding user experiment 1, participants rated the learnability of the breathing interaction (Q1) as similar to that of a joystick. However, in terms of ease of use (Q2), participants clearly found the breathing interaction difficult to use. This may be because the joystick operation is more similar to the WIMP method, while the breathing interaction is relatively unfamiliar to them (Q3).

[0095] Regarding satisfaction (Q5, Q7), although objective data showed no significant difference in efficiency and error rate between the breathing interaction system and the joystick after becoming proficient, most participants found the breathing interaction faster and more natural. This may be because pie menus and breathing interaction are better suited for tasks involving directional selection.

[0096] It was also noted that due to repeated or prolonged irregular blowing, users experienced significant dizziness during the first experiment (Q6).

[0097] Regarding user experiment 2, participants considered the two scrolling modes to be similar in terms of learnability and ease of use (Q9, Q10). However, both objective data and subjective satisfaction (Q12) indicated that GALE was better than GUST. In particular, most people found the GUST mode more tiring and dizzying to use (Q14). This also indirectly proves that the GALE mode is superior to the GUST mode when using the breathing scrolling menu.

[0098] Regarding the dizziness, it is believed to be caused by the high-frequency use of the respiratory interaction system to complete interactive tasks within a certain period of time. In everyday application scenarios, such high-density interaction is rare, so the dizziness can be effectively controlled.

[0099] Another noteworthy phenomenon is that participants generally experienced severe dizziness in the first experiment. However, the dizziness was not significant in the second experiment, indicating that the user experience was greatly improved once participants became accustomed to the breathing interaction.

[0100] This invention proposes treating breathing as an independent interaction channel and constructs a parameter vector related to the breathing channel. Simultaneously, a dual-channel breathing detection technology integrating auditory and visual perception—VisaudiBreath—is implemented. To verify the feasibility of VisaudiBreath, a VR prototype, BreathUI, is designed and implemented to explore the feasibility and efficiency of using breathing direction, intensity, and type to control pie menus and sliding menus.

[0101] Experimental results show that breathing interaction can be used as an independent input channel to control regular interface menus, achieving efficiency comparable to a joystick in pie-shaped menu scenarios. In scrolling menu scenarios, the Gale mode, where the menu scrolling speed is matched to the breathing intensity in real time, performs better. Questionnaire interview results indicate that using breathing for interface interaction has high novelty and user acceptance.

[0102] Currently, the breathing detection method in this invention is relatively simple. In fact, everyone has different mouth movements when breathing. Furthermore, facial trackers are not sensitive enough to subtle changes in mouth movements. It is believed that with the development of hardware and detection algorithms, breathing interaction will have a broader application prospect in VR.

[0103] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method of fine-grained respiratory interaction based on visual and auditory detection, characterized in that, It comprises the following steps: Obtaining breathing parameters based on breathing types, and constructing a breathing input method of interactive vocabulary; Designing a breathing vector to represent several characteristic parameters of breathing interactive actions based on the breathing input method of interactive vocabulary, and obtaining breathing intensity, breathing direction, whether inhaling, breathing time, breathing frequency, expression information and volume information from the several characteristic parameters; Constructing a visual and auditory dual-channel detection model, inputting the expression information and the volume information, and completing breathing detection, including: Firstly, designing an interactive system BreathUI with breathing actions as interactive input methods; Constructing a visual and auditory dual-channel detection model, inputting the expression information and the volume information, receiving audio information by using a lapel microphone, detecting changes in facial expressions by using an HTC Vive face tracker, and capturing facial actions through 38 tracking points on the lips, chin, teeth, tongue and cheeks; Combining the expression information and the volume information, detecting the breathing type of the user based on the interactive system BreathUI, and determining whether the user is exhaling or inhaling; When inhaling is detected, the inhaling intensity at this time is identified, and when exhaling is detected, the system also needs to determine the current breathing direction and detect the blowing intensity at this time.

2. The method for fine-grained respiratory interaction based on visual and audio detection as claimed in claim 1, wherein, The breathing type includes exhaling and inhaling. When a person exhales, the intensity of exhalation is large or small, the time of breathing is long or short, the direction of exhalation is various, and the exhalation frequency is calculated through the number of exhalations in a period of time. When a person inhales, the intensity of inhalation, the time of breathing and the frequency of breathing also change.

3. The method for fine-grained respiratory interaction based on visual and audio detection as claimed in claim 1, wherein, The method for designing a breathing vector to represent several characteristic parameters of breathing interactive actions and obtaining breathing intensity, breathing direction, whether inhaling, breathing time and breathing frequency from the several characteristic parameters comprises: defining information in the breathing process, breathing intensity as a variable m, m is a continuous analog quantity; breathing direction as a variable d, d is a discrete direction variable; whether inhaling as a variable i, i is a bool type variable; breathing time as a variable t; breathing frequency as a variable f; and input variables of the dual-channel detection method used are expression information e and volume information v.

4. The method for fine-grained respiratory interaction based on visual and audio detection as claimed in claim 1, wherein, The breathing vector converts the interactive design process of the user interface into a mapping process of the breathing vector to the operation action vocabulary vector of the user interface.

5. The method for fine-grained respiratory interaction based on visual and audio detection as claimed in claim 4, wherein, The breathing vector includes a breathing direction component, a breathing intensity component, a breathing time and an inhaling state; the breathing direction component is mapped to a directional operation component of the user interface, and is used for directional selection of a pie menu; the breathing intensity component is mapped to a speed change control component of the action vocabulary vector of the user interface, and is used for scrolling speed of a scrolling menu; the breathing time is used for zooming in or out operation of interface elements; The inhaling state is used as a trigger switch of a specific interface input state.

6. The method for fine-grained respiratory interaction based on visual and audio detection as claimed in claim 1, wherein, The breath type judgment specifically comprises: taking the concave condition of the two cheeks during breathing as the basis for judging the breath category, combining the volume of the microphone to detect the occurrence of the breath action, and when the volume reaches a preset threshold, it is determined that the user is making a breath action at this time; after adjusting the volume threshold of each direction breath judgment, the breath direction is mapped to the interactive interface instead of the components of the traditional direction operation.

7. The method for fine-grained respiratory interaction based on visual and audio detection as claimed in claim 1 wherein, The breath direction recognition specifically comprises: when the user makes different direction blowing actions, the facial Tracker returns the numerical value of the corresponding tracking point of the face, and based on the relevant numerical value, the blowing direction of the user is judged.

8. The method for fine-grained respiratory interaction based on visual and audio detection as claimed in claim 1, wherein, The breath intensity detection specifically comprises: using a microphone to accept the volume information when the user blows or inhales, and the volume information not only judges whether the user is blowing or inhaling at this time, but also maps to the breath intensity; in the BreathUI, the change of the breath intensity value is used to control the menu scrolling speed and the interface zooming degree.

Citation Information

Patent Citations

  • Virtual experiment system and method based on multi-modal interaction

    CN111651035A