Working mode switching method and system, head-mounted device and medium
By combining the audio signals collected by the bone conduction sensor and microphone, the signal energy, similarity and vocal confidence are calculated to realize intelligent mode switching of head-mounted devices, solving the problem of inaccurate mode switching in the existing technology and improving user experience and device stability.
Patent Information
- Application Number
- CN202510485745.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
When switching noise reduction mode and transparency mode, existing head-mounted devices are prone to misjudging other people's voices to speak for users in a multi-person environment, resulting in insufficient recognition accuracy and affecting system stability and user experience.
By combining the audio signals collected by the bone conduction sensor and microphone, the signal energy, similarity and vocal confidence can be calculated to achieve intelligent mode switching to avoid misjudgments caused by coughing, non-voice actions or the voice of others in the environment.
It significantly improves the accuracy and stability of mode switching, avoids missed switching, improves user experience, and reduces power consumption and extends device battery life.
Smart Images

Figure CN120343469A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice detection, and particularly to a method, a system, a head-mounted device and a medium for switching working modes. Background Art
[0002] Most intelligent earphone products on the current market generally have two functions: a noise reduction mode and a transparency mode. The noise reduction mode can effectively reduce the interference of environmental noises such as traffic noise, keyboard tapping sounds, and subway running sounds on users; while the transparency mode allows external sounds to enter, helping users to still be able to communicate with the surrounding environment while wearing the earphones and have conversations without having to take off the earphones. Although these two modes can meet the auditory needs in different usage scenarios, in actual use, users still need to manually switch the earphone mode, repeatedly operating between frequent communication and quiet listening, resulting in a poor user experience. Although some existing solutions achieve automatic recognition and mode switching of the conversation scenario by performing feature extraction and model recognition on the microphone signal, however, this method is prone to misidentifying other people's voices as the user's voice in a multi-person environment, and has a low signal-to-noise ratio in a noisy scenario and insufficient recognition accuracy, affecting the stability and practicality of the system. Therefore, there is a need to provide a method, a system, a head-mounted device and a medium for switching working modes. Summary of the Invention
[0003] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method, a system, a head-mounted device and a medium for switching working modes, which improve the problem that the robustness of the head-mounted device to switch the working mode is poor due to the low accuracy of voice source recognition in the prior art.
[0004] To achieve the above object and other related objects, the present invention provides a method for switching working modes, including: acquiring and saving a first audio signal sequence and a second audio signal sequence of the current sampling period; wherein, the first audio signal is collected by a bone conduction sensor of the head-mounted device, and the second audio signal is collected by a microphone of the head-mounted device; calculating the signal energy of the first audio signal sequence; performing similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the similarity corresponding to the current sampling period; inputting the second audio signal sequence into a voice recognition model to obtain the voice confidence of the presence of human voices in the current sampling period; based on the signal energy, the correlation, and the voice confidence, and in combination with the current working mode, switching the working mode between the noise reduction mode and the transparency mode.
[0005] In an embodiment of the present invention, performing similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the correlation corresponding to the current sampling period includes: calculating the mean of the first audio signal sequence and the mean of the second audio signal sequence; calculating the covariance between the first audio signal sequence and the second audio signal sequence based on the mean of the first audio signal sequence and the mean of the second audio signal sequence; calculating the standard deviation of the first audio signal sequence based on the mean of the first audio signal sequence, and calculating the standard deviation of the second audio signal sequence based on the mean of the second audio signal sequence; calculating the Pearson correlation coefficient based on the covariance and the two standard deviations, and using the Pearson correlation coefficient as the correlation corresponding to the current sampling period.
[0006] In an embodiment of the present invention, calculating the mean of the first audio signal sequence and the mean of the second audio signal sequence includes: determining whether the lengths of the first audio signal sequence and the second audio signal sequence are the same: if they are the same, then calculate the mean of the first audio signal sequence and the mean of the second audio signal sequence respectively; otherwise, perform post-processing on the first audio signal sequence to make its length consistent with that of the second audio signal sequence, and calculate the mean of the first audio signal sequence obtained after post-processing and the mean of the second audio signal sequence respectively; where the post-processing is downsampling or interpolation.
[0007] In an embodiment of the present invention, performing similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the similarity corresponding to the current sampling period includes: calculating the dynamic time warping distance between the first audio signal sequence and the second audio signal sequence, and using it as the correlation.
[0008] In an embodiment of the present invention, when the current working mode is the noise reduction mode, based on the signal energy, correlation, and voice confidence, and in combination with the current working mode, switching the working mode between the noise reduction mode and the transparent mode includes: determining whether the signal energy is less than or equal to a preset first threshold:
[0009] If so, keep the current working mode as the noise reduction mode unchanged; otherwise, based on the correlation and voice confidence, and in combination with the current working mode, switch the working mode between the noise reduction mode and the transparent mode.
[0010] In an embodiment of the present invention, based on the correlation and voice confidence, and in combination with the current working mode, switching the working mode between the noise reduction mode and the transparent mode includes: comparing the correlation corresponding to the current sampling period with a preset second threshold, and comparing the voice confidence with a preset third threshold: if the correlation is greater than the second threshold and the voice confidence is greater than the third threshold, then switch the current working mode from the noise reduction mode to the transparent mode; otherwise, keep the current working mode as the noise reduction mode unchanged.
[0011] In an embodiment of the present invention, based on the relevance and voice confidence, and in combination with the current working mode, the working mode is switched between the noise reduction mode and the transparent mode, further including: obtaining the relevance and voice confidence of the historical sampling period; determining the corresponding conversation mode based on the relevance and voice confidence corresponding to the sampling period; and switching the corresponding working mode according to the conversation mode.
[0012] In an embodiment of the present invention, after switching the working mode between the noise reduction mode and the transparent mode, it further includes: continuously monitoring the relevance and voice confidence of each sampling period after the switch; determining the corresponding conversation mode based on the relevance and voice confidence corresponding to the sampling period; and switching the corresponding working mode according to the conversation mode.
[0013] In an embodiment of the present invention, for each sampling period, determining the corresponding conversation mode based on the relevance and voice confidence corresponding to the sampling period includes: comparing the relevance corresponding to the sampling period with a preset second threshold, and comparing the voice confidence corresponding to the sampling period with a preset third threshold: determining whether the relevance is greater than the second threshold and the voice confidence is greater than the third threshold: if so, marking the conversation mode of the sampling period as conversation; otherwise, marking the conversation mode of the sampling period as non-conversation.
[0014] In an embodiment of the present invention, switching the corresponding working mode according to the conversation mode includes: counting the proportion of the conversation mode as conversation in all sampling periods, and determining whether the proportion of conversation is greater than a preset proportion threshold: if so, switching the current working mode from the noise reduction mode to the transparent mode; otherwise, keeping the working mode as the noise reduction mode.
[0015] In an embodiment of the present invention, switching the corresponding working mode according to the conversation mode further includes: counting the conversation mode, and determining whether the number of non-conversations is greater than a preset conversation volume threshold; if so, switching the current working mode from the transparent mode to the noise reduction mode; otherwise, keeping the current working mode as the transparent mode unchanged.
[0016] In an embodiment of the present invention, a switching system for working modes is further provided. The system includes: a data acquisition module, configured to acquire and save a first audio signal sequence and a second audio signal sequence of the current sampling period; wherein, the first audio signal is acquired by a bone conduction sensor of a head-mounted device, and the second audio signal is acquired by a microphone of the head-mounted device; a signal energy determination module, configured to calculate the signal energy of the first audio signal sequence; a similarity analysis module, configured to perform similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the similarity corresponding to the current sampling period; a voice recognition module, configured to input the second audio signal sequence into a voice recognition model to obtain the voice confidence level of the presence of voice in the current sampling period; a mode switching module, configured to switch the working mode between a noise reduction mode and a transparent mode based on the signal energy, the correlation degree, and the voice confidence level, and in combination with the current working mode.
[0017] In an embodiment of the present invention, a head-mounted device is further provided, including: a bone conduction sensor; a microphone; one or more processors; a storage device, configured to store one or more programs, which, when executed by the one or more processors, cause the head-mounted device to implement the switching method of the working mode according to any one of the above.
[0018] In an embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored, which, when executed by a processor of a computer, causes the computer to execute the switching method of the working mode according to any one of the above.
[0019] As described above, a switching method, system, head-mounted device, and medium for working modes of the present invention have the following beneficial effects: By acquiring a first audio signal sequence acquired by a bone conduction sensor and a second audio signal sequence acquired by a microphone in a head-mounted device, and based on the signal energy of the first audio signal sequence, the similarity of the current sampling period, and the voice confidence level, comprehensively considering the voice intensity, the consistency of the voice source, and the voice significance, and in combination with the current working mode, the switching between the noise reduction mode and the transparent mode is intelligently realized. It effectively avoids misjudgment caused by coughing, non-voice actions, or the voice of others in the environment. It significantly improves the accuracy and stability of mode switching. Description of the Drawings
[0020] Figure 1 It is a schematic flowchart of a switching method for working modes provided by an embodiment of the present invention;
[0021] Figure 2 It is a structural block diagram of a switching system for working modes provided by an embodiment of the present invention;
[0022] Figure 3 It is a schematic structural diagram of a head-mounted device provided by an embodiment of the present invention. Detailed implementation manners
[0023] The following uses specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0024] It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0025] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0026] Existing head-mounted devices such as (headphone products) generally support noise reduction mode and transparency mode to meet the auditory needs in different environments. However, in actual use, users still need to frequently switch modes manually, resulting in a poor user experience. To improve the convenience of device use, some technical solutions propose to analyze the environmental sound signals collected by the microphone, extract time-domain, frequency-domain, or time-frequency domain features, and combine traditional machine learning or neural network models to identify the current scene, so as to automatically switch between the noise reduction mode and the transparency mode. However, such methods only rely on the air-conducted sound signals collected by the microphone and have certain limitations. For example, in a multi-person scenario, it is easy to misjudge the speech of others as the speech of the wearer, and in a noisy environment with a low signal-to-noise ratio, it may be difficult to recognize human voices. These factors are likely to cause incorrect switching or switching delay of the working mode, affecting the final use effect of the head-mounted device.
[0027] To address the above issues, the present invention provides a method for switching working modes. By obtaining a first audio signal sequence collected by a bone conduction sensor in a head-mounted device and a second audio signal sequence collected by a microphone, and based on the signal energy of the first audio signal sequence, the similarity of the current sampling period, and the voice confidence level, comprehensively considering the voice intensity, the consistency of the voice source, and the voice saliency, combined with the current working mode, the intelligent switching between the noise reduction mode and the transparency mode is achieved. This effectively avoids misjudgments caused by coughing, non-speech actions, or the voices of others in the environment. It significantly improves the accuracy and stability of mode switching.
[0028] As Figure 1 shown, the method for switching working modes includes the following steps:
[0029] S1. Obtain and save the first audio signal sequence and the second audio signal sequence of the current sampling period; wherein, the first audio signal is collected by the bone conduction sensor of the head-mounted device, and the second audio signal is collected by the microphone of the head-mounted device.
[0030] In the current sampling period, audio signals from two different acoustic sensors are obtained. Among them, the first audio signal is collected by the bone conduction sensor installed in the head-mounted device and is used to reflect the voice signal generated by the wearer through bone vibration. Different from traditional microphones that rely on air conduction, the bone conduction sensor directly picks up the vibration signal of the human bone (such as the mandible). This vibration is caused by vocal cord vibration and is transmitted to the sensor through the bone, avoiding the wind noise and environmental noise interference in air-conducted sound. The second audio signal is collected by the microphone installed in the head-mounted device and is used to capture the voice signal generated by the wearer's voice and the surrounding environment. Among them, the first audio signal is a VPU (Voice Pick-up Sensor) signal, and the second audio signal is a mic (Microphone) signal. The first audio signal and the second audio signal achieve data acquisition and caching processing within the same sampling window, corresponding to form the first audio signal sequence and the second audio signal sequence for subsequent feature analysis. It can be understood that the head-mounted device of the present invention includes but is not limited to headphones, smart glasses, head-mounted display terminals, communication helmets, etc. Any multifunctional wearable terminal integrated with a bone conduction sensor and a microphone module can apply this solution, which is not limited herein. It can be understood that since the audio signals collected by the bone conduction sensor and the microphone are both analog signals, they need to be digitally processed, and the analog audio signals are converted into discrete digital audio signal sequences through analog-to-digital conversion for subsequent feature extraction and analysis.
[0031] It should be noted that the present invention is applicable to the operating scenario of the head-mounted device when the wearer sets it to the noise reduction mode. When the wearer sets the initial mode of the head-mounted device to the transparent mode, considering the user's intention of actively setting the transparent mode, the automatic switching scheme will not be executed. This is because the user usually turns on the transparent mode to continuously obtain the external environmental sounds to maintain the perception of the surrounding information. If it automatically switches back to the noise reduction mode because the user does not make a sound for a period of time, it may interrupt the user's perception of the environment and affect the user experience. Therefore, not executing the switching logic in the transparent mode is more in line with the user's usage expectations.
[0032] S2. Calculate the signal energy of the first audio signal sequence.
[0033] Set the duration of the current sampling period to N milliseconds (such as 32 ms) and the sampling rate to SkHz (such as 48 kHz). Then the total number of sampling points in this sampling period is L = N * S. Thus, the first audio signal sequence corresponding to this sampling period is denoted as X = {x1, x2, ……, x L}, where L is the number of sampling points of the first audio signal sequence in the current sampling period, which is jointly determined by the duration and sampling rate of the current sampling period. The signal energy Energy vpu is as shown in formula (1):
[0034]
[0035] where x i is the amplitude of the i-th first audio signal in the first audio signal sequence. The calculated signal energy is used to characterize the overall intensity of the bone conduction signal in this sampling period, and can be used as a basis for judging whether the user has a vocal behavior, and is used to control the subsequent working mode switching logic. If the signal energy is less than the preset first threshold, the working mode is directly kept as the noise reduction mode unchanged. Otherwise, it is necessary to obtain the second audio signal sequence collected by the microphone and perform subsequent analysis.
[0036] S3. Perform similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the similarity corresponding to the current sampling period.
[0037] Considering that the first audio signal obtained by the bone conduction sensor mainly reflects the bone vibration caused by the wearer's own voice, while the second audio signal obtained by the microphone can also reflect the voices of other people in the environment, the similarity between the two can effectively characterize the consistency of the voice source. Based on this, the present invention performs similarity analysis on the first audio signal sequence and the second audio signal sequence, calculates their similarity, and is used to determine whether the voice captured by the microphone comes from the wearer himself. Among them, the method for determining the similarity includes, but is not limited to, measuring by Euclidean distance, Pearson correlation coefficient or Dynamic Time Warping (DTW) distance, and those skilled in the art can adaptively select based on actual needs.
[0038] Specifically, in order to improve the recognition accuracy of the wearer's voice behavior, in an embodiment of the present invention, S3 includes: calculating the dynamic time warping distance between the first audio signal sequence and the second audio signal sequence, and using it as the correlation degree. The present invention compares the first audio signal sequence collected by the bone conduction sensor with the second audio signal sequence collected by the microphone, and calculates the minimum cumulative alignment distance between the two signals by constructing a cost matrix and performing optimal path matching. This minimum cumulative alignment distance can reflect the morphological similarity between the two signals in the time dimension. After normalizing the calculated minimum cumulative alignment distance, it can be used as the correlation degree of the current sampling period, so as to measure whether the microphone signal comes from the wearer himself, thereby supporting subsequent dialogue scene recognition and working mode control.
[0039] In another embodiment of the present invention, the Pearson correlation coefficient is used as the correlation degree. Specifically, S3 includes S31 to S34 (not shown in the figure):
[0040] S31. Calculate the mean value of the first audio signal sequence and the mean value of the second audio signal sequence.
[0041] Let the first audio signal sequence be X = {x1, x2, ……, x L}, and the second audio signal sequence be Y = {y1, y2, ……, y L}, then the mean value of the first audio signal sequence The mean value of the second audio signal sequence Among them, x i is the i-th first audio signal in the first audio signal sequence, and y i is the i-th second audio signal in the second audio signal sequence.
[0042] Further, in an embodiment of the present invention, S31 includes the following process:
[0043] First, determine whether the lengths of the first audio signal sequence and the second audio signal sequence are the same: If they are the same, calculate the mean of the first audio signal sequence and the mean of the second audio signal sequence respectively. If they are different, perform post-processing on the first audio signal sequence to make its length consistent with that of the second audio signal sequence, and calculate the mean of the first audio signal sequence obtained after post-processing and the mean of the second audio signal sequence respectively; where the post-processing is downsampling or interpolation. Specifically, considering that the bone conduction sensor and the microphone may have inconsistent sampling rates or different sampling periods during the actual sampling process, resulting in different lengths of the signal sequences obtained by the two, in order to ensure the stability of subsequent values, it is necessary to analyze the first audio signal sequence and the second audio signal sequence to determine whether their lengths are consistent during the current sampling period. If the lengths of the two are the same, the means of the first audio signal sequence and the second audio signal sequence in the current sampling period can be directly calculated respectively. On the contrary, if the lengths of the two are inconsistent, post-processing needs to be performed on the first audio signal sequence to match the length of the second audio signal sequence. Specifically, the post-processing can include downsampling or interpolation. For example, when the length of the first audio signal sequence is less than the length of the second audio signal sequence, interpolation processing can be performed on the first audio signal sequence, and vice versa, downsampling can be performed on the first audio signal sequence. It should be noted that this post-processing method is also applicable to the processing process of the first audio signal sequence when the above dynamic time warping distance is used as the correlation degree.
[0044] S32. Calculate the covariance between the first audio signal sequence and the second audio signal sequence based on the mean of the first audio signal sequence and the mean of the second audio signal sequence.
[0045] In order to determine whether the change trends of the first audio signal collected by the bone conduction sensor and the second audio signal collected by the microphone are consistent during the current sampling period, it is necessary to calculate the covariance conv(X, Y) between the first audio signal sequence X and the second audio signal sequence Y in the current sampling period according to formula (2) on the basis of calculating the mean:
[0046]
[0047] The covariance is used to reflect the consistency of the amplitude changes of two signals in the same sampling period. If the covariance is greater than zero, it means that the two trends are consistent and have a strong correlation, and it can be inferred that the microphone contains the wearer's own voice signal.
[0048] S33. Calculate the standard deviation of the first audio signal sequence based on the mean of the first audio signal sequence, and calculate the standard deviation of the second audio signal sequence based on the mean of the second audio signal sequence.
[0049] Denote the first audio signal sequence as X = {x1, x2, ……, xL}, with the mean denoted as Its standard deviation is The second audio signal sequence is denoted as Y = {y1, y2, ……, y L}, with the mean denoted as Its standard deviation is
[0050] S34. Calculate the Pearson correlation coefficient based on the covariance and the two standard deviations, and use the Pearson correlation coefficient as the correlation degree corresponding to the current sampling period.
[0051] After obtaining the covariance and the standard deviations corresponding to the two audio signal sequences, use formula (3) to calculate the Pearson correlation coefficient between the two audio signals:
[0052]
[0053] Among them, ρ(X, Y) is the Pearson correlation coefficient between the first audio signal sequence X and the second audio signal sequence Y. The value range of this Pearson correlation coefficient is [-1, 1]. The closer the value is to 0, the less correlated the two signals are. The Pearson correlation coefficient is used as the correlation degree corresponding to the current sampling period to determine whether the microphone contains the voice of the wearer himself, so as to provide a basis for the subsequent automatic switching of the working mode.
[0054] S4. Input the second audio signal sequence into the speech recognition model to obtain the voice confidence of the presence of human voice in the current sampling period.
[0055] Extract features from the second audio signal sequence collected by the microphone to obtain acoustic features. Among them, the feature extraction methods include, but are not limited to, Mel-Frequency Cepstral Coefficients (MFCC), Short-Time Fourier Transform (STFT), etc. As long as the audio signal feature extraction can be effectively achieved, it is not limited here. Input the extracted acoustic features into the pre-trained speech recognition model to analyze the key discriminant information in the acoustic features, so as to obtain the voice confidence. This voice confidence is used to represent the probability of the existence of human voice in the current sampling period, and its value is in the range of [0,1]. The higher the voice confidence, the higher the probability that there is clear and stable human voice information in the microphone signal in the current sampling period. This voice confidence is used as the key information for identifying the wearer's vocal behavior and is used to judge whether to enter the transparent mode subsequently. It should be noted that the speech recognition model can be any model that can recognize the existence of speech, including but not limited to Convolutional Neural Network (CNN), Gated Recurrent Unit (GRU), Temporal Convolutional Network (TCN), etc., which is not limited here.
[0056] S5. Based on the signal energy, correlation and voice confidence, and combined with the current working mode, switch the working mode between the noise reduction mode and the transparent mode.
[0057] Specifically, when the current working mode is the noise reduction mode, S5 includes the following process:
[0058] First, judge whether the signal energy is less than or equal to a preset first threshold: if the signal energy is less than or equal to the first threshold, keep the current working mode as the noise reduction mode unchanged. On the contrary, if the signal energy is greater than the first threshold, based on the correlation and voice confidence, and combined with the current working mode, switch the working mode between the noise reduction mode and the transparent mode.
[0059] In the scenario where the initial working mode applicable to the head-mounted device is the noise reduction mode in the present invention, in this mode, if the calculated signal energy is less than or equal to a preset first threshold, it indicates that the wearer is in a silent or non-speaking state. At this time, the head-mounted device will keep the noise reduction mode unchanged and there is no need to analyze the second audio signal. On the contrary, if the signal energy is higher than the first threshold, it indicates that the wearer may be speaking or having a conversation, and then it is necessary to further obtain the second audio signal collected by the microphone for subsequent correlation and voice confidence analysis to confirm whether it is necessary to switch the current noise reduction mode to the transparent mode.
[0060] Further, in an embodiment of the present invention, based on the correlation and voice confidence, and combined with the current working mode, switching the working mode between the noise reduction mode and the transparent mode includes the following process:
[0061] First, compare the correlation corresponding to the current sampling period with a preset second threshold, and compare the voice confidence with a preset third threshold. If the correlation is greater than the second threshold and the voice confidence is greater than the third threshold, then switch the current working mode from the noise reduction mode to the transparent mode. On the contrary, keep the current working mode as the noise reduction mode unchanged.
[0062] When the energy of the first audio signal collected by the bone conduction sensor is higher than the first threshold and the device is currently in the noise reduction mode, the correlation between the first audio signal and the second audio signal and the voice confidence in the second audio signal will be further comprehensively analyzed to determine whether it is necessary to switch to the transparent mode. Specifically, compare the correlation of the current sampling period with the second threshold, and compare the voice confidence with the third threshold: if the correlation is greater than the second threshold and the voice confidence is greater than the third threshold, it indicates that there is not only a relatively clear voice signal in the current environment, and this voice highly matches the bone conduction signal characteristics of the wearer himself, indicating that the wearer may be in a speaking environment. At this time, switch the current working mode from the noise reduction mode to the transparent mode. On the contrary, keep the noise reduction mode unchanged. By the above method, it can be ensured that the transparent switch is only triggered when the wearer himself actually has a speaking behavior, thus effectively avoiding the mis-switching situation caused by misidentifying other people's voices or background noises, and improving the user experience.
[0063] In an embodiment of the present invention, based on the correlation and voice confidence, and combined with the current working mode, switching the working mode between the noise reduction mode and the transparent mode further includes: first obtaining the correlation and voice confidence of the historical sampling period, then determining the corresponding conversation mode based on the correlation and voice confidence corresponding to the sampling period, and finally switching the corresponding working mode according to the conversation mode.
[0064] Considering that the information in a single sampling period is not sufficient to accurately reflect the user's true usage scenario. For example, the user only briefly replies with one or two words such as "here" or "mm-hmm", but does not continue the conversation. Such instantaneous behaviors can easily cause the device to frequently switch between the noise reduction mode and the transparency mode, thus affecting the user experience. To improve the stability of the working mode switch and avoid frequent switching caused by instantaneous recognition errors, the present invention not only makes a judgment based on the relevance and voice confidence of the current sampling period, but also introduces historical sampling information for auxiliary decision-making. Specifically, the relevance and voice confidence corresponding to multiple consecutive historical sampling periods before the current sampling period are obtained by means of a sliding window, and combined with the relevance and voice confidence of the current sampling period, the conversation mode of the wearer is comprehensively judged. According to the determined conversation state and combined with the current working mode, a mode switching operation is performed: if it is judged to be in a continuous conversation state and the current is in the noise reduction mode, then switch to the transparency mode; otherwise, if it is judged to be in a non-conversation state, automatically fallback to the noise reduction mode. In this way, false switching caused by short-time voice interference is avoided, thereby improving the user experience of the wearer in the voice interaction scenario.
[0065] Further, in an embodiment of the present invention, for each sampling period, determining the corresponding conversation mode based on the relevance and voice confidence corresponding to the sampling period includes: comparing the relevance corresponding to the sampling period with a preset second threshold, and comparing the voice confidence corresponding to the sampling period with a preset third threshold: judging whether the relevance is greater than the second threshold and the voice confidence is greater than the third threshold: if so, marking the conversation mode of the sampling period as conversation; otherwise, marking the conversation mode of the sampling period as non-conversation.
[0066] For each sampling period: First, obtain the relevance and voice confidence corresponding to the sampling period, and compare them with the preset second threshold and third threshold respectively. If the relevance of the sampling period is greater than the second threshold and the voice confidence is greater than the third threshold, it indicates that the current voice signal is very likely to come from the wearer himself, and the conversation mode of the sampling period can be marked as conversation; otherwise, it is marked as non-conversation. The above recognition results based on the sampling period can be used as the input basis for subsequent trend judgment of consecutive periods to improve the accuracy and stability of the overall working mode switching strategy.
[0067] In an embodiment of the present invention, switching the corresponding working mode according to the conversation mode includes: counting the proportion of the sampling periods in which the conversation mode is conversation, and judging whether the proportion of conversation is greater than a preset proportion threshold: if so, switching the current working mode from the noise reduction mode to the transparency mode; otherwise, keeping the working mode as the noise reduction mode.
[0068] When the current working mode is the noise reduction mode, the proportion of the number of sampling periods marked as conversations in the current and historical sampling periods to all sampling periods will be counted. If the conversation proportion is greater than the proportion threshold, it means that the wearer is in a continuous or frequent voice interaction state, and the current working mode will be switched from the noise reduction mode to the transparent mode so that the wearer can communicate. On the contrary, if the conversation proportion does not reach the proportion threshold, it means that the wearer is in a silent or non-voice state, and the current noise reduction mode remains unchanged.
[0069] In an embodiment of the present invention, switching the corresponding working mode according to the conversation mode further includes: counting the number of conversation modes and determining whether the number of non-conversations is greater than a preset conversation volume threshold; if so, switching the current working mode from the transparent mode to the noise reduction mode; otherwise, keeping the current working mode as the transparent mode unchanged.
[0070] In the present invention, after the above process, after switching from the noise reduction mode to the transparent mode, in order to avoid frequent switching of the working mode due to a short pause of the user, the conversation mode of each sampling period will be continuously monitored to determine whether it is necessary to switch back from the transparent mode to the noise reduction mode. Specifically, continuously count the number of non-conversations marked in consecutive sampling periods and compare this number with the preset conversation volume threshold. It can be understood that the conversation volume threshold is used to determine whether the conversation has ended and can be adaptively set according to actual interaction requirements, which is not limited here. If the counted number of non-conversations exceeds the conversation volume threshold, it means that the user has stopped speaking or the conversation behavior has ended, and at this time, the current working mode will be automatically switched from the transparent mode to the noise reduction mode. On the contrary, if the conversation volume does not exceed the conversation volume threshold, it is considered that the user may still be in the conversation, and the current transparent mode will remain unchanged. In this way, the stability and user experience in the voice interaction scenario are significantly improved. In addition, after counting the number of conversation modes, it can also be determined whether the number of conversations is less than a preset non-conversation volume threshold. If so, the current working mode will be switched from the transparent mode to the noise reduction mode; otherwise, the current working mode will be kept as the transparent mode unchanged.
[0071] Such as Figure 2As shown, the switching system 100 for the working mode includes: a data acquisition module 110, a signal energy determination module 120, a similarity analysis module 130, a voice recognition module 140, and a mode switching module 150. The data acquisition module 110 is configured to acquire and save a first audio signal sequence and a second audio signal sequence during the current sampling period; wherein, the first audio signal is collected by a bone conduction sensor of the head-mounted device, and the second audio signal is collected by a microphone of the head-mounted device. The signal energy determination module 120 is configured to calculate the signal energy of the first audio signal sequence. The similarity analysis module 130 is configured to perform a similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the similarity corresponding to the current sampling period. The voice recognition module 140 is configured to input the second audio signal sequence into a voice recognition model to obtain the voice confidence level indicating the presence of voice during the current sampling period. The mode switching module 150 is configured to switch the working mode between a noise reduction mode and a transparent mode based on the signal energy, the correlation degree, and the voice confidence level, and in combination with the current working mode.
[0072] For the specific limitations of the switching system for the working mode, reference can be made to the limitations of the switching method for the working mode in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned switching system for the working mode can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of a processor in a computer device in a hardware format, or stored in a memory in the computer device in a software format, so as to facilitate the processor to call the corresponding operations of the above-mentioned modules.
[0073] It should be noted that, in order to highlight the innovative part of the present invention, modules that are not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other modules in this embodiment.
[0074] As Figure 3 shown, the head-mounted device 1 may include a memory 11, a processor 12, a bone conduction sensor 13, a microphone 14, and a bus, and may further include a computer program stored in the memory 11 and executable on the processor 12, such as a switching program for the working mode.
[0075] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 can be an internal storage unit of the head-mounted device 1, such as the mobile hard disk of the head-mounted device 1. In some other embodiments, the memory 11 can also be an external storage device of the head-mounted device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the head-mounted device 1. Further, the memory 11 can also include both the internal storage unit of the head-mounted device 1 and the external storage device. The memory 11 can be used not only to store application software and various types of data installed on the head-mounted device 1, such as the code for switching work modes, etc., but also to temporarily store data that has been output or will be output.
[0076] In some embodiments, the processor 12 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged together, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 12 is the control core (Control Unit) of the head-mounted device 1, connecting all components of the entire head-mounted device 1 through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as the program for switching work modes, etc.), and calling data stored in the memory 11, to execute various functions of the head-mounted device 1 and process data.
[0077] The processor 12 executes the operating system of the head-mounted device 1 and various installed application programs. The processor 12 executes the application programs to implement the steps in the above-mentioned method for switching work modes.
[0078] Exemplarily, the computer program can be divided into one or more modules, and one or more modules are stored in the memory 11 and executed by the processor 12 to complete this application. One or more modules can be a series of computer program instruction segments that can complete specific functions, and these instruction segments are used to describe the execution process of the computer program in the head-mounted device 1. For example, the computer program can be divided into a data acquisition module 110, a signal energy determination module 120, a similarity analysis module 130, a voice recognition module 140, and a mode switching module 150.
[0079] The integrated unit implemented in the form of software functional modules described above can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above software functional modules are stored in a storage medium and include several instructions to enable a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the functions of the method for switching working modes in various embodiments of the present application.
[0080] In summary, a method, a system, a head-mounted device, and a medium for switching working modes disclosed in the present invention can accurately identify the scene where the user is located and automatically match the corresponding headphone working mode, thereby improving the overall user experience. By fusing the information collected by the microphone and the bone conduction sensor, it is possible to comprehensively judge whether the user is in an actual conversation state, effectively avoiding mis-switching caused by non-conversation behaviors such as coughing, clearing the throat, and "um". At the same time, in the presence of environmental voice interference (such as others' conversations), the first audio signal can be used to determine whether the voice comes from the wearer himself, so as to avoid misidentifying the voice of others as the user's voice and keep the noise reduction mode running stably. In addition, the present invention also has the advantage of low power consumption. In a quiet or silent state, a preliminary screening can be performed through the first audio signal collected by the bone conduction sensor, and the subsequent recognition algorithm is only started when a possible vocalization behavior is detected, significantly reducing the calculation frequency in unnecessary situations, thereby effectively reducing power consumption and extending the battery life of the device. Therefore, the present invention effectively overcomes various disadvantages in the prior art and has high industrial utilization value.
[0081] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A method for switching working modes, characterized in that, Applied to a head-mounted device, the switching method includes: Obtain and save the first audio signal sequence and the second audio signal sequence of the current sampling period; wherein, the first audio signal is collected by a bone conduction sensor of the head-mounted device, and the second audio signal is collected by a microphone of the head-mounted device; Calculate the signal energy of the first audio signal sequence; Perform similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the correlation corresponding to the current sampling period; Input the second audio signal sequence into a speech recognition model to obtain the voice confidence of the presence of voice in the current sampling period; Based on the signal energy, the correlation, and the voice confidence, and in combination with the current working mode, switch the working mode between a noise reduction mode and a transparent mode.
2. The method for switching the working mode according to claim 1, wherein The performing similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the correlation corresponding to the current sampling period includes: Calculate the mean of the first audio signal sequence and the mean of the second audio signal sequence; Based on the mean of the first audio signal sequence and the mean of the second audio signal sequence, calculate the covariance between the first audio signal sequence and the second audio signal sequence; Calculate the standard deviation of the first audio signal sequence based on the mean of the first audio signal sequence, and calculate the standard deviation of the second audio signal sequence based on the mean of the second audio signal sequence; Based on the covariance and the two standard deviations, calculate the Pearson correlation coefficient, and use the Pearson correlation coefficient as the correlation corresponding to the current sampling period.
3. The method for switching the working mode according to claim 2, wherein The calculating the mean of the first audio signal sequence and the mean of the second audio signal sequence includes: Judge whether the lengths of the first audio signal sequence and the second audio signal sequence are the same: If they are the same, calculate the mean of the first audio signal sequence and the mean of the second audio signal sequence respectively; Otherwise, perform post-processing on the first audio signal sequence to make its length consistent with that of the second audio signal sequence, and calculate the mean of the first audio signal sequence and the mean of the second audio signal sequence obtained after post-processing respectively; wherein, the post-processing is downsampling or interpolation.
4. The method for switching the working mode according to claim 1, wherein The performing similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the correlation corresponding to the current sampling period includes: calculating the dynamic time warping distance between the first audio signal sequence and the second audio signal sequence, and using it as the correlation.
5. The method for switching the working mode according to claim 1, wherein When the current working mode is the noise reduction mode, the switching the working mode between the noise reduction mode and the transparent mode based on the signal energy, the correlation, and the voice confidence, and in combination with the current working mode includes: Judge whether the signal energy is less than or equal to a preset first threshold: If so, keep the current working mode as the noise reduction mode unchanged; Otherwise, based on the correlation and the voice confidence, and in combination with the current working mode, switch the working mode between the noise reduction mode and the transparent mode.
6. The method for switching the working mode according to claim 5, characterized in that Based on the relevance and the voice confidence level, and in combination with the current working mode, switching the working mode between a noise reduction mode and a transparent mode, including: Comparing the relevance corresponding to the current sampling period with a preset second threshold, and comparing the voice confidence level with a preset third threshold: If the relevance is greater than the second threshold and the voice confidence level is greater than the third threshold, then switching the current working mode from the noise reduction mode to the transparent mode; Otherwise, keeping the current working mode as the noise reduction mode unchanged.
7. The method for switching the working mode according to claim 5, characterized in that, Based on the relevance and the voice confidence level, and in combination with the current working mode, switching the working mode between a noise reduction mode and a transparent mode, further includes: Obtaining the relevance and the voice confidence level of a historical sampling period; Determining a corresponding conversation mode based on the relevance and the voice confidence level corresponding to the sampling period; Switching the corresponding working mode according to the conversation mode.
8. The method for switching the working mode according to claim 1, characterized in that, After switching the working mode between the noise reduction mode and the transparent mode, further includes: Continuously monitoring the relevance and the voice confidence level of each sampling period after the switch; Determining a corresponding conversation mode based on the relevance and the voice confidence level corresponding to the sampling period; Switching the corresponding working mode according to the conversation mode.
9. The method for switching the working mode according to claim 7 or 8, characterized in that, For each sampling period, determining a corresponding conversation mode based on the relevance and the voice confidence level corresponding to the sampling period, including: Comparing the relevance corresponding to this sampling period with a preset second threshold, and comparing the voice confidence level corresponding to this sampling period with a preset third threshold: Judging whether the relevance is greater than the second threshold and the voice confidence level is greater than the third threshold: If so, marking the conversation mode of this sampling period as conversation; Otherwise, marking the conversation mode of this sampling period as non - conversation.
10. The method for switching the working mode according to claim 7 or 8, characterized in that, Switching the corresponding working mode according to the conversation mode, including: Counting the proportion of sampling periods with the conversation mode as conversation, and judging whether the proportion of conversation is greater than a preset proportion threshold: If so, switching the current working mode from the noise reduction mode to the transparent mode; Otherwise, keeping the working mode as the noise reduction mode.
11. The method for switching the working mode according to claim 7 or 8, characterized in that, Switching the corresponding working mode according to the conversation mode, further includes: Counting the conversation mode and judging whether the number of non - conversations is greater than a preset conversation quantity threshold; If so, switching the current working mode from the transparent mode to the noise reduction mode; Otherwise, keeping the current working mode as the transparent mode unchanged.
12. A switching system for working modes, characterized in that, Applied to a head - mounted device, the system includes: A data acquisition module, configured to acquire and save a first audio signal sequence and a second audio signal sequence of the current sampling period; wherein, the first audio signal is collected by a bone conduction sensor of the head - mounted device, and the second audio signal is collected by a microphone of the head - mounted device; A signal energy determination module, configured to calculate the signal energy of the first audio signal sequence; A similarity analysis module, configured to perform a similarity analysis on the first audio signal sequence and the second audio signal sequence to obtain the relevance corresponding to the current sampling period; A voice recognition module, configured to input the second audio signal sequence into a speech recognition model to obtain a voice confidence level indicating the presence of voice in the current sampling period; A mode switching module, configured to switch the working mode between a noise reduction mode and a transparency mode based on the signal energy, the correlation, and the voice confidence level, in combination with the current working mode.
13. A head-mounted device, characterized in that, The head-mounted device includes: A bone conduction sensor; A microphone; One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the head-mounted device to implement the working mode switching method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which, when executed by a processor of a computer, causes the computer to execute the working mode switching method according to any one of claims 1 to 11.
Citation Information
Cited By
Intelligent glasses voice interaction method and system, equipment and storage medium
CN122285168A