Wearable device and voice processing method and device

By using a multi-axis bone conduction sensor in wearable devices, the optimal voice signal can be selected for processing based on the usage scenario, thus solving the design freedom and wearing comfort issues of single-axis bone conduction sensor solutions and achieving more stable sound pickup performance.

CN121905162APending Publication Date: 2026-04-21SHANGHAI QIANWEN ZHILIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI QIANWEN ZHILIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
Filing Date
2026-01-13
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing single-axis bone conduction sensor solutions have high requirements for device layout, structural process and wearing fit, which limits the freedom of industrial design and wearing comfort, and is difficult to adapt to different head shapes or usage scenarios, resulting in unstable sound pickup performance.

Method used

Using a multi-axis bone conduction sensor, multiple bone conduction sensor axes are set in the wearable device to collect voice signals from each axis. The optimal voice signal from the best axis is selected according to the usage scenario for the voice recognition process, and microphone signals are combined for auxiliary processing in noisy environments.

Benefits of technology

It enhances design freedom and wearing comfort, adapts to different head shapes or scenarios, improves the stability of sound pickup performance, and reduces the requirements for component layout and structural processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121905162A_ABST
    Figure CN121905162A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses wearable equipment and a voice processing method and device. A multi-axis bone conduction sensor is arranged in the wearable device, first voice signals of all axes are collected by using the multi-axis bone conduction sensor, and after a use scene of the wearable device is determined, the first voice signal of one axis is selected from the first voice signals of all the axes according to the use scene to be determined as a target first voice signal; and executing a speech recognition process based on the target first speech signal. Therefore, the better target first voice signal is selected from the first voice signals of all the axes according to the use scene of the wearable device for subsequent processing, the requirements of the bone conduction sensor for the device layout, the structure process and the wearing fitting degree can be reduced, the design freedom degree and the wearing comfort can be improved, the wearable device can adapt to different head types or scenes, and the application range is wide. And the stability of the pickup performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a wearable device, a voice processing method, and an apparatus. Background Technology

[0002] With the increasing demand for voice interaction from wearable devices such as smart glasses and headphones, accurately picking up user voice in high-noise environments and effectively suppressing environmental interference has become key to improving user experience.

[0003] Currently, the most common solution is the single-axis bone conduction sensor, which is usually integrated into a specific location such as the nose pad to transmit the user's voice signal by utilizing the vibration of the skull. However, the single-axis bone conduction sensor solution has very high requirements for device layout, structural manufacturing process, and fit, which limits the freedom of industrial design and wearing comfort. It is also difficult to adapt to different head shapes or usage scenarios, resulting in unstable sound pickup performance. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a wearable device, a voice processing method and apparatus that can reduce the requirements of bone conduction sensors on device layout, structural process and wearing fit, improve design freedom and wearing comfort, adapt to different head shapes or scenarios, and improve the stability of sound pickup performance.

[0005] In a first aspect, embodiments of the present invention provide a voice processing method for a wearable device, the wearable device including a multi-axis bone conduction sensor, the method comprising: Acquire the first speech signal of each axis collected by the multi-axis bone conduction sensor; Determine the usage scenarios for the wearable device; Based on the usage scenario, select one axis's first speech signal from the first speech signals of each axis and determine it as the target first speech signal; The speech recognition process is performed based on the target first speech signal.

[0006] In some embodiments, determining the current usage scenario based on the voice playback status of the wearable device includes: Obtain the voice playback status of the wearable device; In response to the wearable device being in a voice playback state, the current usage scenario is determined as the playback scenario; In response to the fact that the wearable device is not currently in voice playback mode, the current usage scenario is determined to be a non-playback scenario.

[0007] In some embodiments, determining the first speech signal of one axis from the first speech signals of each axis as the target first speech signal according to the usage scenario includes: Obtain a pre-set mapping relationship, which is the correspondence between the usage scenario and the axis; Determine the target axis corresponding to the current usage scenario based on the mapping relationship; The first speech signal of the target axis is determined as the first speech signal of the target.

[0008] In some embodiments, the speech recognition process based on the target first speech signal includes: In response to the wearable device being in standby mode, the target first voice signal is identified to determine whether it includes a wake word; In response to a wake word, the wearable device is switched to activation mode.

[0009] In some embodiments, the speech recognition process based on the target first speech signal further includes: In response to the wearable device being in an active mode, an operation command is obtained based on the target first voice signal; Execute the operation instructions.

[0010] In some embodiments, the wearable device further includes a microphone, wherein, prior to acquiring the first speech signals of each axis collected by the multi-axis bone conduction sensor, the method further includes: Acquire the second voice signal collected by the microphone; Determine whether the current environment is a noisy environment based on the second voice signal; Since the current environment is not a noisy environment, the speech recognition process is performed based on the second speech signal.

[0011] In some embodiments, acquiring the first speech signals of each axis collected by the multi-axis bone conduction sensor includes: In response to the fact that the current environment is a noisy environment, the first speech signal of each axis collected by the multi-axis bone conduction sensor is acquired.

[0012] In some embodiments, the speech recognition process based on the target first speech signal includes: In response to the wearable device being in standby mode, the target first voice signal and second voice signal are identified to determine whether a wake word is included; In response to a wake word, the wearable device is switched to activation mode.

[0013] In some embodiments, the speech recognition process based on the target first speech signal further includes: In response to the wearable device being in an active mode, an operation command is obtained based on the target first voice signal and second voice signal; Execute the operation instructions.

[0014] In some embodiments, switching the wearable device to an activation mode in response to including a wake word includes: In response to including a wake word, detect whether it is a false wake-up based on the first voice signal of two predetermined axes in the multi-axis bone conduction sensor; The system responds to the false wake-up and keeps the wearable device in standby mode. In response to the absence of a false wake-up, the wearable device is switched to active mode.

[0015] In some embodiments, the multi-axis bone conduction sensor includes a first axis, a second axis, and a third axis; when the wearable device is worn correctly, the first axis is in the vertical direction, the second axis is in the left-right direction, and the third axis is in the front-back direction.

[0016] In a second aspect, embodiments of the present invention provide a wearable device, the wearable device comprising: Equipment body; A multi-axis bone conduction sensor is installed at a designated location on the device body to collect the first voice signal of each axis; A memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to perform the following steps: Acquire the first speech signal of each axis collected by the multi-axis bone conduction sensor; Determine the usage scenarios for the wearable device; Based on the usage scenario, the first speech signal of one axis in the first speech signals of each axis is selected as the target first speech signal; The speech recognition process is performed based on the target first speech signal.

[0017] In some embodiments, the wearable device further includes: A microphone is used to collect a second voice signal.

[0018] In some embodiments, the multi-axis bone conduction sensor includes a first axis, a second axis, and a third axis; when the wearable device is worn correctly, the first axis is in the vertical direction, the second axis is in the left-right direction, and the third axis is in the front-back direction.

[0019] Thirdly, embodiments of the present invention provide a voice processing device for a wearable device, the wearable device including a multi-axis bone conduction sensor, the device comprising: The first speech signal acquisition unit is used to acquire the first speech signals of each axis collected by the multi-axis bone conduction sensor; The usage scenario determination unit is used to determine the usage scenario of the wearable device. A target speech signal determination unit is configured to select a first speech signal of one axis from the first speech signals of each axis according to the usage scenario and determine it as the target first speech signal. A speech signal processing unit is used to perform a speech recognition process based on the target first speech signal.

[0020] Fourthly, embodiments of the present invention provide a computer program product comprising a computer program, wherein when the computer program is run on a computer, the computer executes the method described in the first aspect above.

[0021] Fifthly, embodiments of the present invention provide a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the method described in the first aspect.

[0022] The technical solution of this invention involves setting a multi-axis bone conduction sensor in a wearable device. This sensor collects first speech signals along each axis. After determining the usage scenario of the wearable device, a first speech signal from each axis is selected as the target first speech signal based on the usage scenario. A speech recognition process is then performed based on this target first speech signal. Therefore, by selecting a superior target first speech signal from each axis for subsequent processing based on the usage scenario of the wearable device, the requirements for device layout, structural manufacturing, and fit of the bone conduction sensor can be reduced. This increases design freedom and wearing comfort, allowing for adaptation to different head shapes or scenarios and improving the stability of sound pickup performance. Attached Figure Description

[0023] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which: Figure 1 This is a schematic diagram of a wearable device according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a wearable device according to another embodiment of the present invention; Figure 3 This is a circuit diagram of a wearable device according to an embodiment of the present invention; Figure 4 This is a flowchart of the voice processing method of the wearable device according to the first embodiment of the present invention; Figure 5 This is a flowchart illustrating the determination of usage scenarios in an embodiment of the present invention; Figure 6 This is a flowchart illustrating the determination of the target first speech signal according to an embodiment of the present invention; Figure 7 This is a schematic diagram of signal quality according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the mapping relationship in an embodiment of the present invention; Figure 9 This is a flowchart of the voice processing method of the wearable device according to the second embodiment of the present invention; Figure 10 This is a schematic diagram of the voice processing device of a wearable device according to an embodiment of the present invention. Detailed Implementation

[0024] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0025] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0026] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".

[0027] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0028] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be processed within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0029] With the widespread application of wearable devices (such as smart glasses, headphones, AR / VR headsets, hearing aids, and wearable health monitoring terminals) in various fields, voice interaction, as one of the commonly used user input methods, is becoming increasingly important. To achieve highly reliable voice pickup and recognition, devices need to be able to accurately capture the wearer's own voice in complex acoustic environments while effectively suppressing environmental noise interference. Bone conduction technology, due to its unique physical mechanism, is gradually becoming a key means to improve the near-field voice signal-to-noise ratio. Compared to traditional air conduction microphones, bone conduction devices (such as vibrating piezoelectric units or vibration acceleration sensors) are highly sensitive to the wearer's own voice but almost unresponsive to external environmental sound waves, thus naturally possessing excellent acoustic isolation characteristics. They are particularly suitable for applications such as voice communication, voice wake-up, and voice command recognition in noisy environments.

[0030] Currently, commonly used bone conduction voice pickup solutions generally employ a single-axis sensing architecture, meaning they detect vibration signals only in a single direction. While this type of solution can achieve a certain level of voice pickup in specific wearing positions (such as nose pads on glasses), its performance is highly dependent on the precise structural positioning of the device and the consistency of the assembly process. However, this requirement not only greatly limits the freedom of industrial design but also seriously affects wearing comfort. For example, to accommodate the installation of bone conduction sensors, the nose pads must be fixed, which significantly reduces wearing comfort compared to detachable nose pads. Especially in prolonged use, concentrated local pressure can easily cause user discomfort. At the same time, individual differences in facial anatomy (such as nose bridge height) make it difficult for fixed designs to achieve stable and reliable signal pickup across all population groups, leading to fluctuations or even failures in voice recognition performance.

[0031] Therefore, embodiments of the present invention provide a wearable device and a corresponding voice processing method to solve the above problems.

[0032] It should be noted that, for ease of explanation, the following description uses smart glasses as an example of wearable device. However, the embodiments of the present invention do not limit the type of wearable device. Other types of wearable devices (such as smart helmets, headphones, AR / VR headsets, hearing aids, and wearable health monitoring terminals) are also applicable to the technical solutions of the embodiments of the present invention.

[0033] Figure 1 This is a schematic diagram of a wearable device according to an embodiment of the present invention. Figure 1 As shown, the wearable device 1 of this embodiment includes a device body 11 and a multi-axis bone conduction sensor 12.

[0034] When the wearable device 1 is a smart glasses, the device body 11 includes a frame 11a and temples 11b, etc.

[0035] A multi-axis bone conduction sensor 12 is installed at a designated position on the device body 11 to collect the first voice signal of each axis.

[0036] Specifically, the multi-axis bone conduction sensor 12 can be implemented using bone conduction devices such as a multi-axis VPU (Voice Pick Up) or a multi-axis VCC (vibration and conduction component).

[0037] When a user speaks, the skull and / or jawbone generate complex vibrations with energy distributed in multiple directions. A multi-axis VPU or multi-axis VCC can decompose these vibrations into accelerations in multiple different directions. The acquired acceleration signals are then processed through high-pass filtering, integral approximation, and other techniques to obtain the speech waveform, i.e., the speech signal.

[0038] The multi-axis bone conduction sensor can be a biaxial, triaxial, or more-axis bone conduction sensor. This embodiment of the invention uses a multi-axis bone conduction sensor 12 comprising three axes as an example, namely a first axis, a second axis, and a third axis. When the wearable device is worn correctly, the direction of the first axis is vertical, such as... Figure 1 The x-axis in the diagram. The second axis is located in the left-right direction, as shown in the diagram. Figure 1 The y-axis in the diagram, where the third axis is located in the front-back direction, such as... Figure 1 The z-axis is also relevant. It should be noted that the bone conduction sensors for both axes are also applicable to the technical solutions of this invention, which will not be elaborated upon here.

[0039] Furthermore, this embodiment of the invention does not limit the placement of the multi-axial bone conduction sensor 12; it can be placed at any location on the device body according to actual needs. For example, the multi-axial bone conduction sensor 12 can be placed at such locations as... Figure 1 The position shown, for example, can also be set as follows: Figure 2 The locations of the three dashed boxes are shown. Furthermore, the multi-axis bone conduction sensor 12 can be placed in other possible locations, which will not be listed here in this embodiment of the invention.

[0040] It should be noted that, Figure 1 Only the location of the multi-axis bone conduction sensor of the wearable device is shown. Other components of the wearable device (such as microphones, processors, memory, communication components, etc.) are not shown. These components can be installed in any location on the wearable device.

[0041] Figure 3 This is a circuit diagram of a wearable device according to an embodiment of the present invention. Figure 3As shown, the wearable device of this embodiment includes a multi-axis bone conduction sensor 12, a microphone 13, a processor 14, a memory 15, and a communication component 16.

[0042] The multi-axis bone conduction sensor 12 is installed at a designated position on the device body to collect the first voice signal of each axis.

[0043] Microphone 13 is used to collect the second voice signal.

[0044] The memory 15 is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor 14 to implement the speech processing method of the present invention.

[0045] Specifically, Figure 4 This is a flowchart of the voice processing method for a wearable device according to the first embodiment of the present invention. Figure 4 As shown, the voice processing method of the wearable device in this embodiment of the invention includes the following steps: Step S110: Acquire the first speech signal of each axis collected by the multi-axis bone conduction sensor.

[0046] In this embodiment, a multi-axis bone conduction sensor is used to collect the first voice signal of each axis. The multi-axis bone conduction sensor can be a two-axis, three-axis or more-axis bone conduction sensor. This embodiment of the invention uses a three-axis bone conduction sensor as an example. The multi-axis bone conduction sensor includes a first axis, a second axis and a third axis. When the wearable device is worn correctly, the direction of the first axis is vertical, the direction of the second axis is left and right, and the direction of the third axis is front and back.

[0047] Step S120: Determine the usage scenario of the wearable device.

[0048] In this embodiment, the current usage scenario is determined based on the voice playback status of the wearable device, and the usage scenario includes a playback scenario and a non-playback scenario.

[0049] Figure 5 This is a flowchart illustrating the determination of usage scenarios in an embodiment of the present invention. For example... Figure 5 As shown, determining the use case of the wearable device includes the following steps: Step S121: Obtain the voice playback status of the wearable device.

[0050] In this embodiment, after acquiring the first voice signals of each axis collected by the multi-axis bone conduction sensor, the voice playback status of the wearable device is obtained.

[0051] Specifically, wearable devices can play voice messages in certain scenarios, such as voice messages for user interaction, music, and other audio data. This voice playback is often controlled by a processor. Therefore, when the processor receives the first voice signals from each axis collected by the multi-axis bone conduction sensor, it can detect whether voice playback is currently in progress.

[0052] Step S122: In response to the wearable device being in voice playback mode, the current usage scenario is determined as the playback scenario.

[0053] In this embodiment, when the wearable device is currently in voice playback mode, the current usage scenario is determined as the playback scenario.

[0054] Step S123: In response to the fact that the wearable device is not currently in voice playback mode, the current usage scenario is determined to be a non-playback scenario.

[0055] In this embodiment, when the wearable device is not currently in a voice playback state, the current usage scenario is determined as a non-playback scenario.

[0056] Step S130: Based on the usage scenario, select one axis's first speech signal from the first speech signals of each axis and determine it as the target first speech signal.

[0057] In this embodiment, the signal quality of each axis will be different under different usage scenarios. Therefore, it is necessary to select the first voice signal of one axis from the first voice signals of each axis according to the usage scenario and determine it as the target first voice signal.

[0058] Specifically, Figure 6 This is a flowchart illustrating the determination of the target first speech signal according to an embodiment of the present invention. Figure 6 As shown, the step of selecting the first speech signal of each axis from the first speech signals of each axis according to the usage scenario and determining it as the target first speech signal includes the following steps: Step S131: Obtain the pre-set mapping relationship.

[0059] In this embodiment, the mapping relationship is the correspondence between the usage scenario and the axis.

[0060] Specifically, the signal quality of each axis will vary depending on the application scenario. Figure 7 This is a schematic diagram of signal quality according to an embodiment of the present invention. Figure 7 This illustrates the degree of influence of different scenarios on different axes, such as Figure 7As shown, the more asterisks, the stronger the signal of that axis. This embodiment of the invention divides signal strength into three levels: strong, medium, and weak. Three asterisks indicate a strong signal, two asterisks indicate a medium signal, and one asterisk indicates a weak signal.

[0061] like Figure 7 As shown, the signal quality along the x-axis is strong, while the signal quality along the y-axis and z-axis is medium. In other words, if we do not consider the influence of other factors, the processing result obtained by using the first speech signal along the x-axis for speech processing is the most accurate.

[0062] Regarding the impact on speaker playback, different speaker mounting methods have different effects on different axes. If the speaker diaphragm's movement direction is perpendicular to the x-axis, the signal quality along the x-axis is strong, the signal quality along the y-axis is weak, and the signal quality along the z-axis is strong. If the speaker diaphragm's movement direction is perpendicular to the y-axis, the signal quality along the x-axis is weak, the signal quality along the y-axis is strong, and the signal quality along the z-axis is weak.

[0063] Regarding the influence of external noise, the signal quality of the x-axis is weak, the signal quality of the y-axis is weak, and the signal quality of the z-axis is strong.

[0064] Therefore, embodiments of the present invention can select different axes in different scenarios based on the signal differences of different axes in different scenarios, which can effectively pick up the wearer's speech while isolating speech interference in that scenario.

[0065] For example, in non-playback scenarios, if only the influence of external noise is considered, and the x-axis is used as the target axis, the first speech signal on the x-axis can be used for subsequent processing to achieve the highest sensitivity to the wearer's speech signal and the strongest suppression of external noise.

[0066] For example, in a playback scenario, considering both external noise and speaker playback, if the speaker diaphragm's movement direction is perpendicular to the x-axis, using the y-axis as the target axis, the wearer's speech signal, while lower than the x-axis, will have significantly lower interference from the speaker playback, and the suppression of external noise will be strongest. This prevents clipping in AEC (Acoustic Echo Cancellation) at high volumes, resulting in the highest overall benefit in terms of SNR (Signal-to-Noise Ratio). If the speaker diaphragm's movement direction is perpendicular to the y-axis, using the y-axis as the target axis, the wearer's speech signal sensitivity will be highest, speaker playback interference will be lowest, and the suppression of external noise will be strongest.

[0067] Therefore, embodiments of the present invention can set corresponding effective axes for different scenarios and establish corresponding mapping relationships. Taking the example of a speaker diaphragm moving perpendicular to the x-axis, the corresponding mapping relationship is as follows: Figure 8 As shown. In Figure 8 In the embodiments shown, the usage scenarios include playback scenarios and non-playback scenarios, where the axis corresponding to the playback scenario is the y-axis and the axis corresponding to the non-playback scenario is the x-axis.

[0068] Step S132: Determine the target axis corresponding to the current usage scenario based on the mapping relationship.

[0069] In this embodiment, the target axis corresponding to the current usage scenario is determined according to the mapping relationship. For example, if the current usage scenario is a playback scenario, the target axis is the x-axis; if the current usage scenario is a non-playback scenario, the target axis is the y-axis.

[0070] Step S133: Determine the first speech signal of the target axis as the target first speech signal.

[0071] In this embodiment, after obtaining the target axis corresponding to the usage scenario, the first voice signal of the target axis is determined as the target first voice signal.

[0072] Step S140: Perform a speech recognition process based on the target first speech signal.

[0073] In this embodiment, after obtaining the target first voice signal, a voice recognition process can be performed based on the target first voice signal.

[0074] The speech recognition process can be divided into two stages: the wake-up stage and the instruction execution stage.

[0075] For the wake-up phase, the speech recognition process based on the target first speech signal includes the following steps: Step S141: In response to the wearable device being in standby mode, the target first voice signal is identified to determine whether it includes a wake word.

[0076] In this embodiment, if the current wearable device is in standby mode, the current stage is determined to be the wake-up stage, and the target first voice signal needs to be identified to determine whether the target first voice signal includes a wake-up word.

[0077] Step S142: In response to the wake word, switch the wearable device to activation mode.

[0078] In this embodiment, if the target first voice signal includes a wake-up word, the wearable device is woken up, that is, the wearable device is switched to active mode. If the target first voice signal does not include a wake-up word, the wearable device remains in standby mode, and the system operates based on the detected voice signal.

[0079] Furthermore, during the wake-up phase, the first speech signals of each axis acquired by the multi-axis bone conduction sensor can be used to detect false alarms. Specifically, in response to a wake-up word, whether it is a false wake-up is detected based on the first speech signals of two predetermined axes in the multi-axis bone conduction sensor. If it is a false wake-up, the wearable device is kept in standby mode; if it is not a false wake-up, the wearable device is switched to active mode. Specifically, if the difference between the first speech signals of the two predetermined axes is greater than or equal to a predetermined threshold, it is considered a false wake-up; if the difference between the first speech signals of the two predetermined axes is less than the predetermined threshold, it is considered not a false wake-up.

[0080] Specifically, in combination Figure 7 Since the difference in signal amplitude between the x-axis (or y-axis) and z-axis is small when the wearer speaks, while the difference in signal amplitude between the x-axis (or y-axis) and z-axis is significant under external noise, the wake word can be determined as belonging to the wearer or a non-wearer based on the amplitude difference characteristics between the x-axis (or y-axis) and z-axis, thereby optimizing the false wake-up rate.

[0081] During the instruction execution phase, the speech recognition process based on the target first speech signal includes the following steps: Step S143: In response to the wearable device being in active mode, obtain operation instructions based on the target first voice signal.

[0082] In this embodiment, when the wearable device is woken up, that is, when the wearable device is in active mode, the user can input commands via voice. For example, the user can say voice commands such as "take a photo," "take a video," "make a call to xxx," or "play music." The target first voice signal is recognized to determine the user's intention, and corresponding operation commands are generated based on the user's intention.

[0083] Step S144: Execute the operation command.

[0084] In this embodiment, after obtaining the operation instruction, the corresponding operation instruction is executed, such as sending a shooting instruction to the image acquisition module.

[0085] This invention, through the integration of a multi-axis bone conduction sensor into a wearable device, collects first speech signals along each axis. After determining the usage scenario of the wearable device, it selects a first speech signal from each axis as the target first speech signal based on that scenario, and then performs a speech recognition process based on this target first speech signal. Therefore, by selecting a superior target first speech signal from each axis for subsequent processing based on the wearable device's usage scenario, the requirements for bone conduction sensor layout, structural manufacturing, and fit can be reduced, increasing design freedom and wearing comfort. This allows for adaptation to different head shapes or scenarios and improves the stability of sound pickup performance.

[0086] It should be noted that the above-described speech processing method can be deployed as a complete speech processing solution in wearable devices. That is, the wearable device can rely on multi-axis bone conduction sensors to collect the first speech signals of each axis to execute the complete speech recognition process. Furthermore, the above-described speech processing method can also be used as an auxiliary process for speech recognition, implemented as part of the speech recognition process.

[0087] When the above speech processing method is used as an auxiliary process for speech recognition, and implemented as part of the speech recognition process, it mainly processes the speech signal collected by the microphone. When the noise is relatively high, it uses the speech signal collected by the multi-axis bone conduction sensor for auxiliary processing. The corresponding speech processing method is as follows: Figure 9 As shown. Among them, Figure 9 This is a flowchart of a voice processing method for a wearable device according to a second embodiment of the present invention. Figure 9 As shown, the speech processing method of this invention includes the following steps: Step S210: Acquire the second voice signal collected by the microphone.

[0088] In this embodiment, the first voice signal of each axis is acquired by the multi-axis bone conduction sensor, and the second voice signal is acquired by the microphone. The acquired voice signals are stored in memory. The second voice signal acquired by the microphone is then retrieved from memory.

[0089] Step S220: Determine whether the current environment is a noisy environment based on the second voice signal.

[0090] In this embodiment, it is determined whether the current environment is a noisy environment based on the acquired second speech signal. This determination can be achieved using various existing methods, such as short-time energy of speech frames, zero-crossing rate, speech activity detection-assisted judgment, signal-to-noise ratio estimation, spectral feature analysis, machine learning, and deep learning.

[0091] For the short-term energy of the speech frame, the energy of the second speech signal is calculated. When there is no speech, the energy mainly comes from noise. If there is a continuous high energy but no clear speech structure, it is determined to be a noisy environment.

[0092] Regarding the zero-crossing rate, noisy environments typically have a higher zero-crossing rate, while human speech has a lower zero-crossing rate. Therefore, by detecting the zero-crossing rate of the second speech signal, it can be determined whether it is a noisy environment. If the zero-crossing rate is greater than or equal to a predetermined zero-crossing rate threshold, it is considered a noisy environment; otherwise, it is considered not a noisy environment.

[0093] For speech activity detection-assisted judgment, speech activity detection can be used to determine whether someone is speaking. Based on this principle, if the speech activity detection output is non-speech for a long time, but the microphone has a continuous signal, it is considered to be a noisy environment.

[0094] For signal-to-noise ratio estimation, the signal-to-noise ratio of the second speech signal is calculated to determine whether it is a noisy environment.

[0095] For spectral characteristic analysis, the sound of human speech is usually concentrated between 300Hz and 3.4kHz. If there are abnormally low or high frequencies, it is judged to be a noisy environment.

[0096] For machine learning and deep learning, a detection model is trained to output "whether it is a noisy environment". This detection model can then directly determine whether it is a noisy environment based on the second speech signal.

[0097] Furthermore, if the current environment is not a noisy environment, the second speech signal is directly used for speech recognition processing, and the process proceeds to step S230.

[0098] If the current environment is noisy, directly using the second speech signal for speech recognition processing may result in inaccurate processing results. Therefore, it is necessary to process the first speech signal collected by the bone conduction sensor and proceed to step S240.

[0099] Step S230: Perform a speech recognition process based on the second speech signal.

[0100] In this embodiment, in response to the fact that the current environment is not a noisy environment, a speech recognition process is performed based on the second speech signal.

[0101] Specifically, the speech recognition process based on the second speech signal includes: determining the current operating mode of the wearable device; in response to the wearable device being in standby mode, recognizing the second speech signal to determine whether it contains a wake-up word; if the second speech signal contains a wake-up word, waking up the wearable device, i.e., switching the wearable device to active mode; if the target second speech signal does not contain a wake-up word, keeping the wearable device in standby mode, and based on the detected second speech signal; in response to the wearable device being in active mode, obtaining an operation command based on the second speech signal, and executing the operation command.

[0102] Step S240: Acquire the first speech signal of each axis collected by the multi-axis bone conduction sensor.

[0103] In this embodiment, given that the current environment is noisy, it is necessary to process the first speech signal acquired by the multi-axis bone conduction sensor. First, the first speech signals of each axis acquired by the multi-axis bone conduction sensor are retrieved from the aforementioned memory.

[0104] Step S250: Determine the usage scenario of the wearable device.

[0105] In this embodiment, the current usage scenario is determined based on the voice playback status of the wearable device, and the usage scenario includes a playback scenario and a non-playback scenario.

[0106] Specifically, after acquiring the first voice signals of each axis collected by the multi-axis bone conduction sensor, the voice playback status of the wearable device is obtained. When the wearable device is currently in voice playback mode, the current usage scenario is determined as a playback scenario. When the wearable device is not currently in voice playback mode, the current usage scenario is determined as a non-playback scenario.

[0107] Step S260: Based on the usage scenario, select one axis's first speech signal from the first speech signals of each axis and determine it as the target first speech signal.

[0108] In this embodiment, the signal quality of each axis will be different under different usage scenarios. Therefore, it is necessary to select the first voice signal of one axis from the first voice signals of each axis according to the usage scenario and determine it as the target first voice signal.

[0109] Specifically, a pre-set mapping relationship is obtained, which is a correspondence between usage scenarios and axes. For example, the axis corresponding to a playback scenario is the y-axis, and the axis corresponding to a non-playback scenario is the x-axis. Based on this mapping relationship, a target axis corresponding to the current usage scenario is determined. For example, if the current usage scenario is a playback scenario, the target axis is the x-axis; if the current usage scenario is a non-playback scenario, the target axis is the y-axis. After obtaining the target axis corresponding to the usage scenario, the first audio signal of the target axis is determined as the target first audio signal.

[0110] Step S270: Perform a speech recognition process based on the target first speech signal and the second speech signal.

[0111] In this embodiment, a speech recognition process is performed based on the first target speech signal and the second speech signal obtained above.

[0112] Specifically, in response to the wearable device being in standby mode, the target first voice signal and second voice signal are identified to determine whether a wake-up word is included. In response to the inclusion of a wake-up word, the wearable device is switched to active mode. In response to the wearable device being in active mode, an operation command is obtained based on the target first voice signal and second voice signal, and the operation command is executed.

[0113] When processing the first and second target speech signals simultaneously, they can be fused to obtain a fused signal, which is then used for subsequent processing. The fusion of the first and second target speech signals can be implemented using various methods, such as signal-level fusion and feature-level fusion.

[0114] Signal-level fusion involves weighting and superimposing or filtering the target first and second speech signals in the time or frequency domains to obtain the fused signal.

[0115] For feature-level fusion, acoustic features are extracted from the first and second speech signals of the target, respectively, and then the acoustic features are spliced ​​or weighted and fused.

[0116] Furthermore, during the wake-up phase, the first speech signals of each axis acquired by the multi-axis bone conduction sensor can be used to detect false alarms. Specifically, in response to a wake-up word, whether it is a false wake-up is detected based on the first speech signals of two predetermined axes in the multi-axis bone conduction sensor. If it is a false wake-up, the wearable device is kept in standby mode; if it is not a false wake-up, the wearable device is switched to active mode. Specifically, if the difference between the first speech signals of the two predetermined axes is greater than or equal to a predetermined threshold, it is considered a false wake-up; if the difference between the first speech signals of the two predetermined axes is less than the predetermined threshold, it is considered not a false wake-up.

[0117] The calculation of the difference between the first speech signals of the two predetermined axes can be achieved using various existing methods. For example, the first speech signals of the two axes can be aligned in time, and a predetermined number of points can be sampled at the same time point. The difference between each pair of sampling points can be calculated, and then the average value of the difference between all sampling points can be calculated. The absolute value of the average value can be taken as the difference between the first speech signals of the two axes.

[0118] Specifically, in combination Figure 7 Since the difference in signal amplitude between the x-axis (or y-axis) and z-axis is small when the wearer speaks, while the difference in signal amplitude between the x-axis (or y-axis) and z-axis is significant under external noise, the wake word can be determined as belonging to the wearer or a non-wearer based on the amplitude difference characteristics between the x-axis (or y-axis) and z-axis, thereby optimizing the false wake-up rate.

[0119] This invention, through the integration of a multi-axis bone conduction sensor into a wearable device, collects first speech signals along each axis. After determining the usage scenario of the wearable device, it selects a first speech signal from each axis as the target first speech signal based on that scenario, and then performs a speech recognition process based on this target first speech signal. Therefore, by selecting a superior target first speech signal from each axis for subsequent processing based on the wearable device's usage scenario, the requirements for bone conduction sensor layout, structural manufacturing, and fit can be reduced, increasing design freedom and wearing comfort. This allows for adaptation to different head shapes or scenarios and improves the stability of sound pickup performance.

[0120] Figure 10 This is a schematic diagram of the voice processing device of a wearable device according to an embodiment of the present invention. Figure 10 As shown, the device includes a first speech signal acquisition unit 101, a usage scenario determination unit 102, a target speech signal determination unit 103, and a speech signal processing unit 104. The first speech signal acquisition unit 101 acquires the first speech signals of each axis collected by the multi-axis bone conduction sensor. The usage scenario determination unit 102 determines the usage scenario of the wearable device. The target speech signal determination unit 103 selects one first speech signal from the first speech signals of each axis according to the usage scenario and determines it as the target first speech signal. The speech signal processing unit 104 performs a speech recognition process based on the target first speech signal.

[0121] This invention, through the integration of a multi-axis bone conduction sensor into a wearable device, collects first speech signals along each axis. After determining the usage scenario of the wearable device, it selects a first speech signal from each axis as the target first speech signal based on that scenario, and then performs a speech recognition process based on this target first speech signal. Therefore, by selecting a superior target first speech signal from each axis for subsequent processing based on the wearable device's usage scenario, the requirements for bone conduction sensor layout, structural manufacturing, and fit can be reduced, increasing design freedom and wearing comfort. This allows for adaptation to different head shapes or scenarios and improves the stability of sound pickup performance.

[0122] exist Figure 3 In the illustrated embodiment, the wearable device includes at least one processor 14; a memory 15 communicatively connected to at least one processor 14; and a communication component 16 communicatively connected to a scanning device, the communication component 16 receiving and transmitting data under the control of the processor 14; wherein the memory 15 stores instructions executable by at least one processor 14, the instructions being executed by at least one processor 14 to implement the above-described voice processing method.

[0123] Specifically, the wearable device includes: one or more processors 14 and memory 15. Figure 3 Taking a processor 14 as an example, the processor 14 and the memory 15 can be connected via a bus or other means. Figure 3 Taking a bus connection as an example, memory 15, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Processor 14 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in memory 15, thereby realizing the above-mentioned voice processing method.

[0124] The memory 15 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. Furthermore, the memory 15 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 15 may optionally include memory remotely located relative to the processor 14, and these remote memories may be connected to external devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0125] One or more modules are stored in memory 15 and, when executed by one or more processors 14, perform the speech processing method in any of the above method embodiments.

[0126] The above-mentioned products can perform the methods provided in the embodiments of this application, and have the corresponding functional modules and beneficial effects of performing the methods. For technical details not described in detail in this embodiment, please refer to the methods provided in the embodiments of this application.

[0127] This invention, through the integration of a multi-axis bone conduction sensor into a wearable device, collects first speech signals along each axis. After determining the usage scenario of the wearable device, it selects a first speech signal from each axis as the target first speech signal based on that scenario, and then performs a speech recognition process based on this target first speech signal. Therefore, by selecting a superior target first speech signal from each axis for subsequent processing based on the wearable device's usage scenario, the requirements for bone conduction sensor layout, structural manufacturing, and fit can be reduced, increasing design freedom and wearing comfort. This allows for adaptation to different head shapes or scenarios and improves the stability of sound pickup performance.

[0128] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.

[0129] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0130] This invention also provides A1, a voice processing method for a wearable device, the wearable device including a multi-axis bone conduction sensor, the method comprising: Acquire the first speech signal of each axis collected by the multi-axis bone conduction sensor; Determine the usage scenarios for the wearable device; Based on the usage scenario, the first speech signal of one axis in the first speech signals of each axis is selected as the target first speech signal; The speech recognition process is performed based on the target first speech signal.

[0131] A2. As described in A1, determining the current usage scenario based on the voice playback status of the wearable device includes: Obtain the voice playback status of the wearable device; In response to the wearable device being in a voice playback state, the current usage scenario is determined as a playback scenario; In response to the fact that the wearable device is not currently in voice playback mode, the current usage scenario is determined to be a non-playback scenario.

[0132] A3. As described in A2, the step of selecting one axis's first speech signal from the first speech signals of each axis as the target first speech signal according to the usage scenario includes: Obtain a pre-set mapping relationship, which is the correspondence between the usage scenario and the axis; Determine the target axis corresponding to the current usage scenario based on the mapping relationship; The first speech signal of the target axis is determined as the first speech signal of the target.

[0133] A4. As described in A1, the speech recognition process based on the target first speech signal includes: In response to the wearable device being in standby mode, the target first voice signal is identified to determine whether it includes a wake word; In response to a wake word, the wearable device is switched to activation mode.

[0134] A5. As described in A4, the speech recognition process based on the target first speech signal further includes: In response to the wearable device being in an active mode, an operation command is obtained based on the target first voice signal; Execute the operation instructions.

[0135] A6. The method as described in A1, wherein the wearable device further includes a microphone, and the method further includes, before acquiring the first speech signals of each axis collected by the multi-axis bone conduction sensor: Acquire the second voice signal collected by the microphone; Determine whether the current environment is a noisy environment based on the second voice signal; Since the current environment is not a noisy environment, the speech recognition process is performed based on the second speech signal.

[0136] A7. As described in A6, acquiring the first speech signals of each axis collected by the multi-axis bone conduction sensor includes: In response to the fact that the current environment is a noisy environment, the first speech signal of each axis collected by the multi-axis bone conduction sensor is acquired.

[0137] A8. As described in A7, the speech recognition process based on the target first speech signal includes: In response to the wearable device being in standby mode, the target first voice signal and second voice signal are identified to determine whether a wake word is included; In response to a wake word, the wearable device is switched to activation mode.

[0138] A9. As described in A8, the speech recognition process based on the target first speech signal further includes: In response to the wearable device being in an active mode, an operation command is obtained based on the target first voice signal and second voice signal; Execute the operation instructions.

[0139] A10, as described in A4 or 8, wherein switching the wearable device to an activation mode in response to including a wake word includes: In response to including a wake word, detect whether it is a false wake-up based on the first voice signal of two predetermined axes in the multi-axis bone conduction sensor; The system responds to the false wake-up and keeps the wearable device in standby mode. In response to the absence of a false wake-up, the wearable device is switched to active mode.

[0140] A11. As described in A3, the multi-axis bone conduction sensor includes a first axis, a second axis, and a third axis; when the wearable device is worn correctly, the first axis is in the vertical direction, the second axis is in the left-right direction, and the third axis is in the front-back direction.

[0141] This invention also provides A12, a wearable device, the wearable device comprising: Equipment body; A multi-axis bone conduction sensor is installed at a designated location on the device body to collect the first voice signal of each axis; A memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to perform the following steps: Acquire the first speech signal of each axis collected by the multi-axis bone conduction sensor; Determine the usage scenarios for the wearable device; Based on the usage scenario, the first speech signal of one axis in the first speech signals of each axis is selected as the target first speech signal; The speech recognition process is performed based on the target first speech signal.

[0142] A13. The wearable device as described in A12, further comprising: A microphone is used to collect a second voice signal.

[0143] A14. The wearable device as described in A12, wherein the multi-axis bone conduction sensor includes a first axis, a second axis, and a third axis; when the wearable device is worn correctly, the direction of the first axis is vertical, the direction of the second axis is left-right, and the direction of the third axis is front-back.

[0144] This invention also provides A15, a voice processing device for a wearable device, the wearable device including a multi-axis bone conduction sensor, the device comprising: The first speech signal acquisition unit is used to acquire the first speech signals of each axis collected by the multi-axis bone conduction sensor; The usage scenario determination unit is used to determine the usage scenario of the wearable device. A target speech signal determination unit is configured to select a first speech signal of one axis from the first speech signals of each axis according to the usage scenario and determine it as the target first speech signal. A speech signal processing unit is used to perform a speech recognition process based on the target first speech signal.

[0145] The present invention also provides A16, a computer program product, the computer program product comprising a computer program, wherein when the computer program is run on a computer, the computer executes the method described in any one of A1-A11 above.

[0146] The present invention also provides 17, a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the method as described in any one of A1-A11.

[0147] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A voice processing method for a wearable device, characterized in that, The wearable device includes a multi-axis bone conduction sensor, and the method includes: Acquire the first speech signal of each axis collected by the multi-axis bone conduction sensor; Determine the usage scenarios for the wearable device; Based on the usage scenario, select one axis's first speech signal from the first speech signals of each axis and determine it as the target first speech signal; The speech recognition process is performed based on the target first speech signal.

2. The method according to claim 1, characterized in that, Determining the current usage scenario based on the voice playback status of the wearable device includes: Obtain the voice playback status of the wearable device; In response to the wearable device being in a voice playback state, the current usage scenario is determined as the playback scenario; In response to the fact that the wearable device is not currently in voice playback mode, the current usage scenario is determined to be a non-playback scenario.

3. The method according to claim 2, characterized in that, The step of selecting a first speech signal of each axis from the first speech signals of each axis as the target first speech signal according to the usage scenario includes: Obtain a pre-set mapping relationship, which is the correspondence between the usage scenario and the axis; Determine the target axis corresponding to the current usage scenario based on the mapping relationship; The first speech signal of the target axis is determined as the first speech signal of the target.

4. The method according to claim 1, characterized in that, The speech recognition process based on the target first speech signal includes: In response to the wearable device being in standby mode, the target first voice signal is identified to determine whether it includes a wake word; In response to a wake word, the wearable device is switched to activation mode.

5. The method according to claim 4, characterized in that, The speech recognition process based on the target first speech signal further includes: In response to the wearable device being in an active mode, an operation command is obtained based on the target first voice signal; Execute the operation instructions.

6. A wearable device, characterized in that, The wearable device includes: Equipment body; A multi-axis bone conduction sensor is installed at a designated location on the device body to collect the first voice signal for each axis. A memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to perform the following steps: Acquire the first speech signal of each axis collected by the multi-axis bone conduction sensor; Determine the usage scenarios for the wearable device; Based on the usage scenario, select one axis's first speech signal from the first speech signals of each axis and determine it as the target first speech signal; The speech recognition process is performed based on the target first speech signal.

7. The wearable device according to claim 6, characterized in that, The wearable device also includes: A microphone is used to collect a second voice signal.

8. A voice processing device for a wearable device, characterized in that, The wearable device includes a multi-axis bone conduction sensor, and the device includes: The first speech signal acquisition unit is used to acquire the first speech signals of each axis collected by the multi-axis bone conduction sensor; The usage scenario determination unit is used to determine the usage scenario of the wearable device. A target speech signal determination unit is configured to select a first speech signal of one axis from the first speech signals of each axis according to the usage scenario and determine it as the target first speech signal. A speech signal processing unit is used to perform a speech recognition process based on the target first speech signal.

9. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is run on a computer, the computer performs the method according to any one of claims 1-5.

10. A computer-readable storage medium storing computer program instructions thereon, characterized in that, The computer program instructions, when executed by a processor, implement the method as described in any one of claims 1-5.