Voice interaction control method and device, safety helmet, storage medium and program product
By integrating a throat vibration sensor and a lip movement detection camera into a smart helmet, and combining machine learning models and threshold rules, the problem of misjudgment in voice activity detection has been solved, achieving more accurate voice activity detection and robustness.
Patent Information
- Application Number
- CN202511294059.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-09-11
AI Technical Summary
The voice activity detection function on smart helmets is prone to misjudgment in low voice command levels and special scenarios, causing wearers to be unaware of whether their voice has been picked up, and may mistakenly believe it is due to network latency or other malfunctions.
A throat vibration sensor and a lip movement detection camera are integrated into the smart safety helmet. By collecting throat vibration signals and lip movement signals, and combining machine learning models and threshold rules, it is determined whether the wearer is emitting throat and lip vibrations, and voice activity is detected when the confidence level meets the threshold.
It improves the accuracy of speech activity detection, reduces the probability of semantic misjudgment, and enhances the robustness of speech detection, especially in noisy environments.
Smart Images

Figure CN120853570B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of voice interaction, and in particular to a voice interaction control method, a voice interaction control device, a smart safety helmet, a storage medium, and a computer program product. Background Technology
[0002] Currently, smart helmets often integrate VAD (Voice Activity Detection) functionality to meet the needs of wearers who require voice communication or device control, ultimately improving the efficiency and safety of voice interaction in smart helmet wearing scenarios. Voice activity detection often relies on energy thresholds to distinguish between human voices and noise. However, in specific scenarios such as whispered commands (e.g., quiet warnings in dangerous situations) or loud breathing sounds (e.g., panting while climbing), whispered voices may be misinterpreted as non-human voices and discarded, while panting sounds may be mistaken for continuous speech, leading to microphone snatching. These factors contribute to the inherent vulnerability of voice activity detection; wearers may be unaware of whether their voice has been picked up and may mistakenly believe there is network latency or other malfunctions.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this application is to provide a voice interaction control method, a voice interaction control device, a smart helmet, a storage medium, and a computer program product, aiming to solve the technical problem that the voice activity detection function currently integrated into smart helmets is not accurate enough and has semantic misjudgment.
[0005] To achieve the above objectives, this application proposes a voice interaction control method applied to a smart safety helmet. The smart safety helmet integrates a laryngeal vibration sensor and a lip movement detection camera. The voice interaction control method includes:
[0006] Acquire the laryngeal vibration signal collected by the laryngeal vibration sensor and the lip movement signal collected by the lip movement detection camera;
[0007] The system determines whether the wearer of the smart helmet emits a throat vibration based on the throat vibration signal, and determines whether the wearer of the smart helmet emits a lip vibration based on the lip movement signal.
[0008] The wearer's speech activity is detected when the wearer emits throat and lip vibrations.
[0009] In one embodiment, after the steps of determining whether the wearer of the smart helmet emits throat vibrations based on the throat vibration signal and determining whether the wearer of the smart helmet emits lip vibrations based on the lip movement signal, the method further includes:
[0010] When the wearer emits throat vibrations and lip vibrations, a first score for the throat vibration and a second score for the lip vibration are determined, respectively.
[0011] The confidence level of speech activity detection is determined based on the first score and preset first weight of laryngeal vibration, and the second score and preset second weight of lip vibration.
[0012] When the confidence level is greater than a preset confidence threshold, voice activity detection is performed on the wearer;
[0013] When the confidence level is less than or equal to a preset confidence threshold, a reminder is issued to the wearer to speak again.
[0014] In one embodiment, the step of determining the confidence level of speech activity detection based on a first score and a preset first weight for laryngeal vibration, and a second score and a preset second weight for lip vibration includes the following prior steps:
[0015] Obtain the scene in which the wearer is located;
[0016] Based on the given scenario, a preset first weight for throat vibration and a preset second weight for lip vibration are determined.
[0017] In one embodiment, the step of determining whether the wearer of the smart helmet emits a throat vibration based on the throat vibration signal includes:
[0018] The vibration characteristics of the laryngeal vibration signal are extracted, wherein the vibration characteristics include vibration time-domain characteristics, vibration frequency-domain characteristics, and vibration physiological characteristics;
[0019] Based on the vibration characteristics and the pre-trained classification model, the throat vibration signal is determined to be either semantic vibration or non-semantic vibration.
[0020] In one embodiment, the step of determining whether the throat vibration signal is a semantic vibration or a non-semantic vibration based on the vibration features and a pre-trained classification model includes:
[0021] Acquire historical throat vibration signals;
[0022] The historical laryngeal vibration signal is subjected to semantic or non-semantic vibration pattern recognition to obtain training data, wherein the training data includes the historical laryngeal vibration signal and the labels of the vibration patterns of the historical laryngeal vibration signal.
[0023] The classifier model is trained based on the training data to obtain a classification model.
[0024] In one embodiment, the step of acquiring the laryngeal vibration signal collected by the laryngeal vibration sensor includes:
[0025] The throat vibration sensor emits a first infrared light of a first wavelength and a second infrared light of a second wavelength, and collects the first reflected light corresponding to the first infrared light and the second reflected light corresponding to the second infrared light, wherein the wavelength difference between the first wavelength and the second wavelength is less than a preset wavelength difference value.
[0026] Differential detection is performed based on the first photoelectric signal corresponding to the first reflected light and the second photoelectric signal corresponding to the second reflected light to obtain the throat vibration signal.
[0027] Furthermore, to achieve the above objectives, this application also proposes a voice interaction control device applied to a smart safety helmet. The smart safety helmet integrates a throat vibration sensor and a lip movement detection camera. The voice interaction control device includes:
[0028] The acquisition module is used to acquire the throat vibration signal collected by the throat vibration sensor and the lip movement signal collected by the lip movement detection camera;
[0029] The judgment module is used to determine whether the wearer of the smart safety helmet emits throat vibration based on the throat vibration signal, and to determine whether the wearer of the smart safety helmet emits lip vibration based on the lip movement signal;
[0030] The detection module is used to detect the wearer's speech activity when the wearer emits throat and lip vibrations.
[0031] In addition, to achieve the above objectives, this application also proposes a smart safety helmet, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the voice interaction control method described above.
[0032] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the voice interaction control method described above.
[0033] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the voice interaction control method described above.
[0034] One or more technical solutions proposed in this application have at least the following technical effects:
[0035] In this application, a laryngeal vibration sensor and a lip movement detection camera integrated into a smart helmet are used to acquire laryngeal vibration signals collected by the laryngeal vibration sensor and lip movement signals collected by the lip movement detection camera, respectively. Based on the laryngeal vibration signals, it is determined whether the wearer is emitting laryngeal vibrations, and based on the lip movement signals, it is determined whether the wearer is emitting lip vibrations. When the wearer emits both laryngeal and lip vibrations, speech activity detection is performed. Therefore, by incorporating the detection of laryngeal and lip vibrations, the accuracy of the speech activity detection function integrated into the smart helmet is improved, and the probability of semantic misjudgment is reduced. Attached Figure Description
[0036] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 This is a flowchart illustrating the first embodiment of the voice interaction control method of this application;
[0039] Figure 2 This is a detailed schematic diagram illustrating the steps of the first embodiment of the voice interaction control method of this application;
[0040] Figure 3 This is an application diagram provided for the first embodiment of the voice interaction control method of this application;
[0041] Figure 4 This is a schematic diagram of the module structure of the voice interaction control device according to an embodiment of this application;
[0042] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the voice interaction control method in the embodiments of this application.
[0043] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0044] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0045] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0046] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or smart helmet capable of performing the above functions. The following description uses a smart helmet as an example to illustrate this embodiment and the subsequent embodiments.
[0047] Based on this, embodiments of this application provide a voice interaction control method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the voice interaction control method of this application.
[0048] In this embodiment, the voice interaction control method is applied to a smart safety helmet, which integrates a throat vibration sensor and a lip movement detection camera. The voice interaction control method includes steps S10 to S30:
[0049] Step S10: Acquire the laryngeal vibration signal collected by the laryngeal vibration sensor and the lip movement signal collected by the lip movement detection camera;
[0050] In this embodiment, a laryngeal vibration sensor and a lip movement detection camera are integrated into the smart helmet to collect laryngeal vibration signals and lip movement signals of the wearer, respectively. The laryngeal vibration sensor can be an infrared laryngeal vibration sensor, which directly detects vocal cord vibrations via infrared signals, eliminating the need for sound energy detection. This allows non-semantic vibrations such as coughing and wheezing to be misinterpreted as speech, while still capturing soft commands. This improves the detection rate of low-volume speech below the target decibel level and reduces the false positive rate for such speech. Additionally, a near-infrared miniature camera for detecting lip movements can be installed under the brim of the smart helmet. This camera uses a CNN convolutional neural network to recognize lip movements in real time, including both vocal and non-vocal lip movements.
[0051] Step S20: Determine whether the wearer of the smart helmet emits throat vibration based on the throat vibration signal, and determine whether the wearer of the smart helmet emits lip vibration based on the lip movement signal;
[0052] Infrared laryngeal vibration sensors, used as laryngeal vibration sensors, utilize infrared technology to detect minute vibrations in the larynx, primarily sensing the physical vibrations of the vocal cords and laryngeal tissues during actions such as speaking and swallowing, without relying on ambient airborne sound signals. When acquiring lip movement signals using a lip movement detection camera, an image or video signal of the lips can first be obtained, followed by motion capture image processing to obtain the corresponding lip movement signal.
[0053] The following explanation uses an infrared laryngeal vibration sensor as an example to illustrate its principle. An infrared laryngeal vibration sensor contains an infrared LED that emits infrared light (usually in the near-infrared band). This infrared light shines on the target surface, namely the skin of the throat. When the wearer of the smart helmet speaks or swallows, the throat tissue vibrates, causing minute, rhythmic displacement and vibration of the skin surface. This vibration changes the intensity or pattern (such as the shape of the light spot or the intensity distribution) of the infrared light reflected or scattered back to the infrared laryngeal vibration sensor. For example, the closer the light is to the sensor, the stronger it is usually reflected; as the skin shifts, the angle and intensity of the reflected light also change accordingly. The photodetector inside the infrared laryngeal vibration sensor, such as a photodiode, is responsible for receiving the returned infrared light and converting the received light signal, which is modulated by the laryngeal vibration, into a corresponding weak electrical signal. Then, this electrical signal is processed. That is, this weak electrical signal is amplified by the preamplifier inside the sensor and then filtered (to remove DC components and suppress noise), shaped, etc., and finally outputs an analog or digital electrical signal that is basically synchronized with the laryngeal vibration waveform and can characterize speaking or swallowing activities.
[0054] In one feasible implementation, refer to Figure 2 Step S20 may be followed by steps S201-S202:
[0055] Step S201: Extract the vibration features of the laryngeal vibration signal, wherein the vibration features include vibration time domain features, vibration frequency domain features, and vibration physiological features;
[0056] Step S202: Based on the vibration characteristics and the pre-trained classification model, determine whether the throat vibration signal is semantic vibration or non-semantic vibration.
[0057] The key to how infrared laryngeal vibration sensors can distinguish between non-semantic vibrations such as coughing and wheezing and speech vibrations such as soft speech lies in the differences in the signal characteristics they capture and the intelligent pattern recognition of subsequent signal processing algorithms.
[0058] Although non-semantic vibrations such as coughing and wheezing, as well as speech vibrations such as soft speech, all belong to laryngeal vibrations, they exhibit significantly distinguishable characteristics in their vibration patterns and structures. Firstly, speech vibrations correspond to quasi-periodic speech signals. The regular opening and closing of the vocal cords generates a fundamental frequency and its harmonics, forming a clearly structured spectrum (energy concentrated at specific harmonic frequencies). Even in soft whispers, although the fundamental frequency may weaken or disappear, the vocal tract (oral cavity, tongue position) still modulates airflow, producing specific formants and forming a recognizable speech spectral envelope. For non-semantic vibrations such as coughing and wheezing, coughing is essentially a sudden, strong impact pulse. The glottis opens rapidly to release high-pressure airflow, producing a wide-bandwidth (wide spectral energy distribution), non-periodic, transient burst vibration. It has no stable fundamental frequency or harmonic structure, and its duration is short (usually 100-500 milliseconds). The peak amplitude is much higher than that of ordinary speech. Wheezing, on the other hand, is closer to noise-like or non-stationary signals. It may be a continuous turbulent sound, such as asthma wheezing, or a sound with irregular fluctuations, such as wheezing. It lacks a clear periodic fundamental frequency or harmonic structure, and its spectral energy is concentrated in the mid-low frequency or specific pathological frequencies.
[0059] Therefore, after the infrared laryngeal vibration sensor outputs the raw electrical signal, feature extraction and vibration classification can be performed using digital signal processing (DSP) or machine learning (ML) algorithms. First, the vibration features of the laryngeal vibration signal are extracted, including: 1) vibration time-domain features such as signal energy envelope, zero-crossing rate, kurtosis or sharpness, and short-term energy used to detect cough bursts; 2) vibration frequency-domain features such as spectral centroid, spectral bandwidth, Mel frequency cepstral coefficients, fundamental frequency estimate, and harmonic noise ratio; and 3) vibration physiological features such as glottal closure duration and the ratio between this duration and the vibration duration. Then, based on a pre-trained classification model such as Support Vector Machine (SVM) or neural network such as RNN / DNN / CNN, the extracted feature vectors are input into a machine learning classifier constructed based on the classification model. This machine learning classifier, based on the extracted vibration features of the laryngeal vibration signal, determines whether the laryngeal vibration signal is semantic or non-semantic.
[0060] It should be noted that during the training phase of the pre-trained classification model, it is provided with a large amount of labeled data, such as speech, cough, and wheezing samples, so that it learns to distinguish these three or more types of signals. Therefore, even if the speech energy is very weak, such as soft speech, as long as its related vibration features are successfully extracted, it can be correctly identified as speech.
[0061] In one embodiment, the throat vibration signal can also be judged based on threshold rules to determine whether the wearer of the smart helmet is emitting throat vibrations. For example: if a brief signal peak of less than 500ms and extremely high intensity is detected in the energy envelope, it is marked as a suspected cough; if a stable signal with a clear fundamental frequency (F0) and its harmonic structure (even if the energy is weak) is detected, along with dynamically changing formant structures (F1, F2), it is marked as speech; if continuous low-frequency energy is detected, lacking fundamental frequency and high-frequency harmonic structure (low spectral entropy), it is marked as wheezing.
[0062] In one possible implementation, the following steps are included prior to step S202:
[0063] Acquire historical throat vibration signals;
[0064] The vibration pattern of historical laryngeal vibration signals is identified by semantic or non-semantic vibration to obtain training data, which includes historical laryngeal vibration signals and labels of the vibration patterns of historical laryngeal vibration signals.
[0065] The classifier model is trained based on the training data to obtain the classification model.
[0066] In this embodiment, a training method for a classification model to distinguish whether a throat vibration signal is a semantic vibration or a non-semantic vibration is proposed.
[0067] The acquired historical throat vibration signals are subjected to semantic or non-semantic vibration pattern recognition, and the vibration patterns are labeled to obtain training data. This training data includes historical throat vibration signals and labels for their vibration patterns. It should be noted that before training the classification model, a threshold rule can be used to judge whether the wearer of the smart helmet is emitting throat vibrations, i.e., to determine the vibration pattern of the historical throat vibration signals. Then, the classifier model is trained based on the training data to obtain the classification model.
[0068] Step S30: When the wearer emits throat and lip vibrations, the wearer's speech activity is detected.
[0069] In this embodiment, refer to Figure 3 The trigger logic for detecting the wearer's voice activity is as follows: when the wearer emits throat and lip vibrations, the detection of the wearer's voice activity is activated.
[0070] Therefore, by incorporating the detection of throat and lip vibrations, the accuracy of the voice activity detection function integrated into the smart helmet is improved, and the probability of semantic misjudgment is reduced.
[0071] In one possible implementation, step S20 may be followed by:
[0072] When the wearer emits throat vibrations and lip vibrations, a first score for the throat vibration and a second score for the lip vibration are determined respectively.
[0073] The confidence level of speech activity detection is determined based on the first score and preset first weight of laryngeal vibration, and the second score and preset second weight of lip vibration.
[0074] When the confidence level is greater than a preset confidence threshold, the wearer's voice activity is detected;
[0075] When the confidence level is less than or equal to a preset confidence threshold, a reminder is issued to the wearer to speak again.
[0076] In this embodiment, a method is proposed to more accurately trigger speech activity detection of the wearer based on the throat and lip vibrations emitted by the wearer, aiming to improve the robustness of speech detection, especially in noisy environments.
[0077] First, when the wearer speaks, the smart helmet's throat sensor captures the throat vibration signal. Using preset rules or machine learning models such as support vector machines or neural networks, it outputs a first score (Score_throat). This score represents the reliability of the throat vibration and can range from 0 to 1, where 0 indicates unreliable and 1 indicates highly reliable. For example, the first score might be determined based on the signal-to-noise ratio (SNR) of the signal; for instance, a score of 0.8 is given if the SNR exceeds a threshold, and 0.3 is given if the SNR does not exceed the threshold.
[0078] Secondly, when the wearer speaks, the smart helmet's lip movement detection camera captures lip vibration signals. Similarly, using preset rules or machine learning models such as support vector machines or neural networks, it outputs a second score (Score_lip). This score represents the reliability of the lip vibration, and its value can also range from 0 to 1, where 0 indicates unreliable and 1 indicates highly reliable. For example, the second score might be determined based on the lip movement distance corresponding to the lip movement. For instance, if the lip movement distance exceeds a threshold, the score is 0.9; otherwise, if the lip movement distance does not exceed the threshold, the score is 0.4.
[0079] Then, a first weight (Weight_throat) and a second weight (Weight_lip) are preset. These weights can be set during initialization, and their sum can be set to 1 to reflect the importance of different vibration sources. For example, in a noisy environment, throat vibrations may be more stable, so the first weight (Weight_throat) is preset to 0.6; lip vibrations are sensitive to vocal clarity, so the second weight (Weight_lip) is preset to 0.4. These weights can be experimentally calibrated, such as by optimizing based on historical data, to ensure they are suitable for different wearers and wearing environments.
[0080] Next, the weighted average calculation formula for confidence is adopted: Confidence = (Weight_throat × Score_throat) + (Weight_lip × Score_lip), where confidence ranges from 0 to 1, and the higher the value, the more reliable the speech activity.
[0081] Finally, voice activity detection or alerts are performed based on the confidence level. A preset confidence threshold (typically between 0.5 and 0.7) is used as a decision boundary. Specifically, a higher confidence threshold reduces false detections but may result in missed detections, while a lower threshold has the opposite effect. If the confidence level is greater than the preset threshold, voice activity detection (VAD) is performed, and the results can be used for subsequent applications such as voice command recognition or security alarms. If the confidence level is less than or equal to the preset threshold, a reminder to speak again is sent to the wearer. That is, if there is insufficient vibration signal from the throat or lips (possibly due to environmental noise, unclear pronunciation, or sensor problems), the wearer is reminded to speak clearly again. The reminder can be implemented through the smart helmet's feedback module, such as flashing LEDs, beeping, or slight vibration.
[0082] In one feasible implementation, the step of determining the confidence level of speech activity detection based on a first score and a preset first weight for laryngeal vibration, and a second score and a preset second weight for lip vibration includes the following prior steps:
[0083] To obtain the wearer's current environment;
[0084] Based on the context, a preset first weight for throat vibration and a preset second weight for lip vibration are determined.
[0085] Before determining the confidence level of speech activity detection based on the individual scores of laryngeal and lip vibrations, the corresponding weights are pre-determined according to the wearer's context, namely, a preset first weight for laryngeal vibration and a preset second weight for lip vibration. In one embodiment, a mapping table is stored that contains context and corresponding weights for laryngeal and lip vibrations. This mapping table can be pre-defined with specific contexts and weights. After identifying the current context, the weights for laryngeal and lip vibrations are determined according to the context in this mapping table. For example, in a high-altitude work context, the first and second weights are both 0.5, as determined by consulting the ejection table; in a noisy equipment environment context, the first weight is 0.8 and the second weight is 0.2, indicating that laryngeal vibration is dominant; in a nighttime emergency communication context, the first weight is 0.3 and the second weight is 0.7, indicating that lip vibration is dominant. The above methods determine the weights of laryngeal and lip vibrations based on the specific context, thereby calculating a higher confidence level for speech activity detection that better reflects the actual situation and improves the accuracy of speech activity detection.
[0086] In one feasible implementation, step S10 may include:
[0087] The system emits a first infrared light of a first wavelength and a second infrared light of a second wavelength using a throat vibration sensor, and collects the first reflected light corresponding to the first infrared light and the second reflected light corresponding to the second infrared light. The wavelength difference between the first wavelength and the second wavelength is less than a preset wavelength difference value.
[0088] Differential detection is performed based on the first photoelectric signal corresponding to the first reflected light and the second photoelectric signal corresponding to the second reflected light to obtain the throat vibration signal.
[0089] In this embodiment, considering the weak signal and environmental interference issues—namely, the extremely small displacement of the throat vibration (micrometer level), the slight change in reflected light intensity, and its susceptibility to ambient light noise such as sunlight and infrared heat sources, circuit noise, and motion artifacts such as breathing and limb movements—a dual-wavelength differential detection method is proposed. This method emits two different wavelengths of near-infrared light, such as 850nm / 940nm, to amplify only the vibration-sensitive differential signal and suppress common-mode interference from ambient light.
[0090] First, dual-wavelength light source emission is achieved. The laryngeal vibration sensor has two built-in independent near-infrared LED light sources, which emit near-infrared light of wavelength A and wavelength B respectively. These two wavelengths are relatively close, for example, 850nm and 940nm. The two wavelengths of near-infrared light are simultaneously or rapidly alternately irradiated onto the skin at the same location in the throat.
[0091] Next, the reflected light is received. The skin of the throat reflects both types of infrared light, and the throat vibration sensor's built-in photodetector, such as a photodiode, receives this reflected light. The total received light intensity for wavelength A is the original A plus environmental interference A, and the total received light intensity for wavelength B is the original B plus environmental interference B. Here, original A / original B is the basic reflectance determined by the reflective characteristics of the throat skin itself, such as skin color and roughness, as well as the inherent absorption / reflection characteristics of the corresponding wavelength. Environmental interference A / environmental interference B are the components of stray light in the environment (mainly white light or other infrared sources such as the sun or light bulbs) near wavelengths A and B. Ambient light intensity is usually slowly changing or relatively stable, but it simultaneously affects the total received light intensity for all wavelengths; that is, it is a type of "common-mode" interference.
[0092] When speaking or vibrating the throat, the wearer's throat skin undergoes minute displacement and angular changes. These minute physical vibrations alter the skin's light reflection efficiency, angle, and scattering characteristics. For two near-infrared lights with very similar wavelengths, the changes in reflected light intensity (signal component A and signal component B) caused by this minute physical vibration of the throat skin are highly similar or even almost equal. This is because, in this case, the change in reflected light intensity is caused by the same physical motion (skin vibration) and is not significantly related to the minute differences in wavelength itself. Based on this, the useful signal caused by throat vibration mainly manifests as a small-amplitude, rapidly changing modulated signal (signal component A and signal component B) superimposed on the original base reflected light.
[0093] Then, ambient light (stray light) contains a broad spectrum of components. When the intensity of ambient light changes (such as clouds blocking the sun, lights being turned on or off, or someone walking by blocking the light), it affects light of all wavelengths, similar to a uniform background noise suddenly becoming stronger or weaker. This change in ambient light intensity produces almost the same amount of change in amplitude and direction at wavelengths A and B (change A ≈ change B), which is a highly correlated (synchronously changing) interference signal (i.e., common-mode interference).
[0094] Next, the total photoelectric signal corresponding to wavelength A (original A + environmental interference A + signal component A) is input to one input of the differential amplifier, and the total photoelectric signal corresponding to wavelength B (including original B + environmental interference B + signal component B) is input to the other input of the differential amplifier. The differential amplifier calculates and amplifies the difference between the two input signals, while suppressing the common-mode component. Specifically, when calculating the difference, the output signal = Gain * [(original A + environmental interference A + signal component A) - (original B + environmental interference B + signal component B)], which expands to: Output signal = Gain * [(original A - original B) + (signal component A - signal component B) + (environmental interference A - environmental interference B)]. Based on the above analysis, since environmental interference A ≈ environmental interference B, (environmental interference A - environmental interference B) ≈ 0. Therefore, environmental interference is suppressed, i.e., the common-mode interference term (environmental interference A - environmental interference B) is eliminated. The baseline difference, (original A - original B), is a relatively stable DC offset (because the skin's resting reflectivity differs for different wavelengths), which can be eliminated through subsequent filtering such as high-pass filtering. The vibration signal, (signal component A - signal component B), represents the required throat vibration information. Although signal components A and B caused by physical vibration are very similar (almost in the same direction and with equal amplitude changes), the difference (signal component A - signal component B) is not zero; it is simply smaller, only a fraction or a tenth of the amplitude of signal component A or B. The differential amplifier can amplify this small difference signal (signal component A - signal component B) based on the gain.
[0095] Finally, after passing through the differential amplifier, the output signal is mainly Gain*(signal component A - signal component B) + Gain*(original A - original B). This signal then passes through a high-pass filter (HPF) circuit to filter out the low-frequency, stable baseline difference Gain*(original A - original B). Ultimately, a significantly amplified electrical signal is obtained, mainly containing throat vibration information (corresponding to the difference between signal components A and B), while ambient light interference is greatly suppressed.
[0096] The above utilizes the commonalities (common variation characteristics) of ambient light interference and the differences in characteristics of the target signal (throat vibration) to suppress ambient light interference and extract weak throat vibration signals through dual-wavelength differential detection.
[0097] In another embodiment, similar to the principle of active noise-canceling headphones, active noise cancellation technology can be used. A reference detector, i.e., an additional non-contact photodetector, can be used to collect ambient background light noise and eliminate ambient light interference through a real-time subtraction circuit.
[0098] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the voice interaction control method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0099] This application also provides a voice interaction control device applied to a smart safety helmet. The smart safety helmet integrates a throat vibration sensor and a lip movement detection camera. (Please refer to...) Figure 4 The voice interaction control device includes:
[0100] The acquisition module 10 is used to acquire the laryngeal vibration signal collected by the laryngeal vibration sensor and the lip movement signal collected by the lip movement detection camera;
[0101] The judgment module 20 is used to determine whether the wearer of the smart safety helmet emits throat vibration based on the throat vibration signal, and to determine whether the wearer of the smart safety helmet emits lip vibration based on the lip movement signal.
[0102] The detection module 30 is used to detect the wearer's speech activity when the wearer emits throat and lip vibrations.
[0103] In one embodiment, the determination module 20 is further configured to:
[0104] The steps of determining whether the wearer of the smart helmet emits throat vibrations based on throat vibration signals and determining whether the wearer of the smart helmet emits lip vibrations based on lip movement signals further include:
[0105] When the wearer emits throat vibrations and lip vibrations, a first score for the throat vibration and a second score for the lip vibration are determined respectively.
[0106] The confidence level of speech activity detection is determined based on the first score and preset first weight of laryngeal vibration, and the second score and preset second weight of lip vibration.
[0107] When the confidence level is greater than a preset confidence threshold, the wearer's voice activity is detected;
[0108] When the confidence level is less than or equal to a preset confidence threshold, a reminder is issued to the wearer to speak again.
[0109] In one embodiment, the determination module 20 is further configured to:
[0110] The step of determining the confidence level of speech activity detection based on a first score and a preset first weight for laryngeal vibration, and a second score and a preset second weight for lip vibration, includes the following prior steps:
[0111] To obtain the wearer's current environment;
[0112] Based on the context, a preset first weight for throat vibration and a preset second weight for lip vibration are determined.
[0113] In one embodiment, the determination module 20 is further configured to:
[0114] After the step of determining whether the wearer of the smart helmet emits a throat vibration based on the throat vibration signal:
[0115] Vibration features of the laryngeal vibration signal were extracted, including vibration time-domain features, vibration frequency-domain features, and vibration physiological features.
[0116] Based on vibration characteristics and a pre-trained classification model, the throat vibration signal is determined to be either semantic or non-semantic vibration.
[0117] In one embodiment, the determination module 20 is further configured to:
[0118] The step of determining whether a throat vibration signal is a semantic or non-semantic vibration based on vibration features and a pre-trained classification model includes the following prior steps:
[0119] Acquire historical throat vibration signals;
[0120] The vibration pattern of historical laryngeal vibration signals is identified by semantic or non-semantic vibration to obtain training data, which includes historical laryngeal vibration signals and labels of the vibration patterns of historical laryngeal vibration signals.
[0121] The classifier model is trained based on the training data to obtain the classification model.
[0122] In one embodiment, the acquisition module 10 is further configured to:
[0123] The system emits a first infrared light of a first wavelength and a second infrared light of a second wavelength using a throat vibration sensor, and collects the first reflected light corresponding to the first infrared light and the second reflected light corresponding to the second infrared light. The wavelength difference between the first wavelength and the second wavelength is less than a preset wavelength difference value.
[0124] Differential detection is performed based on the first photoelectric signal corresponding to the first reflected light and the second photoelectric signal corresponding to the second reflected light to obtain the throat vibration signal.
[0125] The voice interaction control device provided in this application, employing the voice interaction control method described in the above embodiments, can solve the technical problem that the voice activity detection function currently integrated into smart safety helmets is not accurate enough and suffers from semantic misjudgment. Compared with the prior art, the beneficial effects of the voice interaction control device provided in this application are the same as those of the voice interaction control method provided in the above embodiments, and other technical features in the voice interaction control device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0126] This application provides a smart safety helmet, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the voice interaction control method in Embodiment 1 above.
[0127] The following is for reference. Figure 5 It shows a structural schematic diagram suitable for implementing the smart safety helmet of the present application embodiments. Figure 5 The smart safety helmet shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.
[0128] like Figure 5 As shown, the smart helmet may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the smart helmet. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the smart helmet to communicate wirelessly or wiredly with other devices to exchange data. While the figures show smart helmets with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0129] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0130] The smart safety helmet provided in this application employs the voice interaction control method described in the above embodiments, which solves the technical problem that the voice activity detection function currently integrated into smart safety helmets is not accurate enough and suffers from semantic misjudgment. Compared with the prior art, the beneficial effects of the smart safety helmet provided in this application are the same as those of the voice interaction control method provided in the above embodiments, and other technical features of this smart safety helmet are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0131] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0132] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0133] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the voice interaction control method in the above embodiments.
[0134] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0135] The aforementioned computer-readable storage medium may be included in the smart helmet; or it may exist independently and not assembled into the smart helmet.
[0136] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the smart helmet, cause the smart helmet to: acquire a laryngeal vibration signal collected by a laryngeal vibration sensor and a lip movement signal collected by a lip movement detection camera; determine whether the wearer of the smart helmet has emitted a laryngeal vibration based on the laryngeal vibration signal, and determine whether the wearer of the smart helmet has emitted a lip vibration based on the lip movement signal; and detect the wearer's speech activity when the wearer emits a laryngeal vibration and a lip vibration.
[0137] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0139] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0140] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described voice interaction control method. This addresses the technical problem that the voice activity detection function currently integrated into smart helmets is not accurate enough and suffers from semantic misjudgment. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the voice interaction control method provided in the above embodiments, and will not be elaborated upon here.
[0141] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the voice interaction control method described above.
[0142] The computer program product provided in this application can solve the technical problem that the voice activity detection function currently integrated into smart safety helmets is not accurate enough and has semantic misjudgment. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the voice interaction control method provided in the above embodiments, and will not be repeated here.
[0143] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A voice interaction control method, characterized in that, The voice interaction control method is applied to a smart safety helmet, which integrates a throat vibration sensor and a lip movement detection camera. The voice interaction control method includes: Acquire the laryngeal vibration signal collected by the laryngeal vibration sensor and the lip movement signal collected by the lip movement detection camera; The system determines whether the wearer of the smart helmet emits a throat vibration based on the throat vibration signal, and determines whether the wearer of the smart helmet emits a lip vibration based on the lip movement signal. When the wearer emits throat and lip vibrations, the wearer's speech activity is detected; The steps of determining whether the wearer of the smart helmet emits throat vibrations based on the throat vibration signal and determining whether the wearer of the smart helmet emits lip vibrations based on the lip movement signal further include: When the wearer emits throat vibrations and lip vibrations, a first score for the throat vibration and a second score for the lip vibration are determined, respectively. Obtain the scene in which the wearer is located; Based on the aforementioned scenario, a preset first weight for throat vibration and a preset second weight for lip vibration are determined. The confidence level of speech activity detection is determined based on the first score and preset first weight of laryngeal vibration, and the second score and preset second weight of lip vibration. When the confidence level is greater than a preset confidence threshold, voice activity detection is performed on the wearer; When the confidence level is less than or equal to a preset confidence threshold, a reminder is issued to the wearer to speak again.
2. The voice interaction control method as described in claim 1, characterized in that, The step of determining whether the wearer of the smart helmet emits a throat vibration based on the throat vibration signal includes: The vibration characteristics of the laryngeal vibration signal are extracted, wherein the vibration characteristics include vibration time-domain characteristics, vibration frequency-domain characteristics, and vibration physiological characteristics; Based on the vibration characteristics and the pre-trained classification model, the throat vibration signal is determined to be either semantic vibration or non-semantic vibration.
3. The voice interaction control method as described in claim 2, characterized in that, Prior to the step of determining whether the throat vibration signal is a semantic vibration or a non-semantic vibration based on the vibration characteristics and a pre-trained classification model, the following steps are included: Acquire historical throat vibration signals; The historical laryngeal vibration signal is subjected to semantic or non-semantic vibration pattern recognition to obtain training data, wherein the training data includes the historical laryngeal vibration signal and the labels of the vibration patterns of the historical laryngeal vibration signal. The classifier model is trained based on the training data to obtain a classification model.
4. The voice interaction control method as described in claim 1, characterized in that, The step of acquiring the laryngeal vibration signal collected by the laryngeal vibration sensor includes: The throat vibration sensor emits a first infrared light of a first wavelength and a second infrared light of a second wavelength, and collects the first reflected light corresponding to the first infrared light and the second reflected light corresponding to the second infrared light, wherein the wavelength difference between the first wavelength and the second wavelength is less than a preset wavelength difference value. Differential detection is performed based on the first photoelectric signal corresponding to the first reflected light and the second photoelectric signal corresponding to the second reflected light to obtain the throat vibration signal.
5. A voice-interactive control device, characterized in that, The voice interaction control device is applied to a smart safety helmet, which integrates a throat vibration sensor and a lip movement detection camera. The voice interaction control device includes: The acquisition module is used to acquire the throat vibration signal collected by the throat vibration sensor and the lip movement signal collected by the lip movement detection camera; The judgment module is used to determine whether the wearer of the smart safety helmet emits throat vibration based on the throat vibration signal, and to determine whether the wearer of the smart safety helmet emits lip vibration based on the lip movement signal; The detection module is used to detect the wearer's speech activity when the wearer emits throat and lip vibrations; The determination module is further configured to: after the steps of determining whether the wearer of the smart helmet emits throat vibration based on the throat vibration signal, and determining whether the wearer of the smart helmet emits lip vibration based on the lip movement signal: When the wearer emits throat vibrations and lip vibrations, a first score for the throat vibration and a second score for the lip vibration are determined, respectively. Obtain the scene in which the wearer is located; Based on the aforementioned scenario, a preset first weight for throat vibration and a preset second weight for lip vibration are determined. The confidence level of speech activity detection is determined based on the first score and preset first weight of laryngeal vibration, and the second score and preset second weight of lip vibration. When the confidence level is greater than a preset confidence threshold, voice activity detection is performed on the wearer; When the confidence level is less than or equal to a preset confidence threshold, a reminder is issued to the wearer to speak again.
6. A smart safety helmet, characterized in that, The smart helmet includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the voice interaction control method as described in any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the voice interaction control method as described in any one of claims 1 to 4.
8. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the voice interaction control method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Voice interaction method and device for subway ticket buying
CN112363861A
Prevent earphone of taking an examination of practising fraud
CN206977656U