Head-mounted device, voice pickup method, and readable storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GEER TECH CO LTD
- Filing Date
- 2023-11-21
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]本申请的主要目的在于提供一种头戴设备,旨在解决目前VPU拾取得到的语音信号较差,包括大量噪声的技术问题
[0026]本申请实施例提出的一种头戴设备、方法、头戴设备及可读存储介质。在本申请实施例中,包括骨传导传感器的语音拾取部件被安装在头戴设备的支架上,而在语音拾取部件中骨传导传感器将通过弹性减震机构与转动机构连接,也即骨传导传感器与头戴设备的连接过程中存在有的弹性减震机构,且弹性减震机构存在最大弹性方向。实际通过骨传导传感器拾取外界的语音时,也即拾通过骨骼和皮肤传播的震动时,震动将通过弹性减震机构传递至骨传导传感器,由于弹性减震机构具有弹性,故可减少震动在弹性减震机构中的衰减,尤其是在所述弹性减震机构最大弹性方向上震动衰减最小,相应的,最大弹性方向上的震动被拾取并产生的信号就越多,反之,非最大弹性方向的震动衰减就更多,相应的,非最大弹性方向的震动别拾取并产生的信号就越少,如此可通过弹性减震机构对震动进行针对性过滤,而且在实际应用中,可通过转动机构控制弹性减震机构转动,从而变更最大弹性方向的指向,使得最大弹性方向的指向与期望的震动(例如,说话声或语音播放部件的声音,可预先确定话声或语音播放部件的声音的震动方向)的震动方向相同,从而实现针对性拾取语音,提高语音拾取结果的信噪比,降低后续音频算法对音频的处理难度。
Smart Images

Figure CN117539064B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of head-mounted device technology, and more particularly to a head-mounted device, a voice pickup method, and a readable storage medium. Background Technology
[0002] Currently, head-mounted devices such as AR (Augmented Reality) glasses or VR (Virtual Reality) glasses are equipped with VPUs (Voice Pickup Units, bone conduction sensors) to pick up voice features from the environment. These are typically fixed in place using FPCs (Flexible Printed Circuits) attached to the outer casing, or the VPU being attached to the motherboard. Because the VPU is attached to the motherboard or casing, the motherboard or casing transmits vibration signals from various external directions to the VPU, resulting in poor voice data picked up by the VPU. This means the obtained voice signal includes a large amount of noise, increasing the difficulty of subsequent audio processing algorithms. Summary of the Invention
[0003] The main objective of this application is to provide a head-mounted device that addresses the technical problem of poor voice signal pickup by current VPUs, including a large amount of noise.
[0004] To achieve the above objectives, this application provides a head-mounted device, including a support frame, and a voice playback component and a voice pickup component mounted on the support frame, wherein the voice pickup component includes:
[0005] A bone conduction sensor, which is used to pick up voice feature data in the environment;
[0006] An elastic damping mechanism, wherein the bone conduction sensor is located at one end of the elastic damping mechanism in the direction of maximum elasticity, and the elastic damping mechanism has the maximum elastic modulus in the direction of maximum elasticity;
[0007] A rotating mechanism is provided to control the rotation of the elastic damping mechanism to change the direction of the maximum elasticity.
[0008] Optionally, the elastic damping mechanism includes a first connection end connected to the bone conduction sensor and a second connection end connected to the rotation mechanism, wherein the direction of the line connecting the first connection end and the second connection end is the direction of maximum elasticity.
[0009] Optionally, the elastic damping mechanism is a conical helical spring, wherein the spring diameter at the first connecting end of the conical helical spring is smaller than the spring diameter at the second connecting end.
[0010] Optionally, the rotating mechanism of the head-mounted device includes a stepper motor and a drive shaft connected to the stepper motor, the drive shaft being fixed to the second connecting end of the conical helical spring.
[0011] Optionally, the angle between the mechanical vibration direction of the voice playback component and the preset voice vibration direction is greater than or equal to a first preset angle threshold, wherein the mechanical vibration direction is the vibration direction of the mechanical vibration generated by the voice playback component.
[0012] Optionally, the voice pickup component may also include a pneumatic sensor.
[0013] To achieve the above objectives, this application provides a voice pickup method, which is applied to the head-mounted device described above;
[0014] The speech pickup method includes:
[0015] Based on the ambient noise intensity of the head-mounted device, determine the target orientation of the maximum elastic direction of the elastic damping mechanism in the head-mounted device;
[0016] Based on the rotation mechanism connected to the elastic damping mechanism in the head-mounted device, the rotation of the elastic damping mechanism is controlled to adjust the maximum elastic direction of the elastic damping mechanism to the target direction;
[0017] Voice feature data is obtained by picking up voice features in the environment through a bone conduction sensor located on the elastic damping mechanism.
[0018] Optionally, the target orientation includes a first target orientation, and the step of determining the target orientation of the maximum elastic direction of the elastic damping mechanism in the head-mounted device based on the ambient noise intensity of the head-mounted device includes:
[0019] If the ambient noise intensity is greater than a preset intensity threshold, then the target direction is the first target direction, wherein after the maximum elastic direction of the elastic damping mechanism is adjusted to the first target direction, the angle between the maximum elastic direction and the preset voice vibration direction is less than or equal to a second preset angle threshold.
[0020] Optionally, the target pointing further includes a second target pointing, and the step of determining the target pointing of the maximum elastic direction of the elastic damping mechanism in the head-mounted device based on the ambient noise intensity of the head-mounted device further includes:
[0021] If the ambient noise intensity is less than or equal to a preset sound intensity threshold, then the target direction is the second target direction, wherein, after adjusting the maximum elastic direction of the elastic damping mechanism to the second target direction, the angle between the maximum elastic direction and the mechanical vibration direction of the voice playback component in the head-mounted device is less than or equal to a second preset angle threshold.
[0022] Optionally, the type of the speech feature data includes first speech feature data picked up by the bone conduction sensor in the posture corresponding to the first target direction, and second speech feature data picked up by the bone conduction sensor in the posture corresponding to the second target direction. After the step of obtaining speech feature data by picking up speech features in the environment through the bone conduction sensor located on the elastic damping mechanism, the method includes:
[0023] The voice data collected by the air conduction sensor in the head-mounted device is subjected to noise reduction processing using the first voice feature data.
[0024] Alternatively, echo cancellation processing can be performed on the voice data collected by the air conduction sensor in the head-mounted device using the second voice feature data.
[0025] To achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium, and stores a speech pickup program thereon. When the speech pickup program is executed by a processor, it implements the steps of the speech pickup method described above.
[0026] This application provides a head-mounted device, method, head-mounted device, and readable storage medium. In this application embodiment, a voice pickup component including a bone conduction sensor is mounted on the bracket of the head-mounted device. In the voice pickup component, the bone conduction sensor is connected to a rotation mechanism through an elastic damping mechanism. That is, an elastic damping mechanism exists during the connection process between the bone conduction sensor and the head-mounted device, and the elastic damping mechanism has a maximum elastic direction. When a bone conduction sensor actually picks up external speech, that is, when it picks up vibrations transmitted through bones and skin, the vibrations are transmitted to the bone conduction sensor through an elastic damping mechanism. Because the elastic damping mechanism is elastic, the attenuation of vibrations within it is reduced, especially in the direction of maximum elasticity. Consequently, more vibrations in the direction of maximum elasticity are picked up and generate more signals. Conversely, vibrations in directions other than maximum elasticity attenuate more, and consequently, fewer vibrations in these directions are picked up and generate fewer signals. This allows for targeted filtering of vibrations through the elastic damping mechanism. In practical applications, the rotation of the elastic damping mechanism can be controlled by a rotating mechanism to change the direction of maximum elasticity, making it the same as the direction of the desired vibration (e.g., the sound of speech or the sound of a voice playback device, the vibration direction of which can be predetermined). This achieves targeted speech pickup, improves the signal-to-noise ratio of the speech pickup results, and reduces the processing difficulty of subsequent audio algorithms. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the state structure of the head-mounted device of this application;
[0028] Figure 2 This is a side view of one embodiment of the voice pickup component in the head-mounted device of this application;
[0029] Figure 3 This is a side view of a posture of another embodiment of the voice pickup component in the head-mounted device of this application;
[0030] Figure 4 This is a side view of another embodiment of the voice pickup component in the head-mounted device of this application, in another pose.
[0031] Figure 5 This is a schematic diagram of another embodiment of the head-mounted device of this application;
[0032] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of this application;
[0033] Figure 7 This is a flowchart illustrating the first embodiment of the speech pickup method in this application;
[0034] Figure 8 This is a schematic diagram of the control framework of the rotating mechanism in the speech pickup method of this application;
[0035] Figure 9 This is a schematic diagram of another state structure of the head-mounted device in the voice pickup method of this application;
[0036] Figure 10 This is a schematic diagram of the speech feature data processing framework in the speech pickup method of this application.
[0037] Explanation of reference numerals in the accompanying drawings of the embodiments:
[0038] 10 Voice pickup component 1 Bone conduction sensor 2 Elastic damping mechanism 3 Rotating mechanism 21 Cylindrical helical spring 22 Conical helical spring 31 Stepper motor 32 transmission shaft 201 First connection end 202 Second connection end 4 Gas conduction sensor 20 Voice playback component 30 support 01 Head-mounted devices
[0039] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0040] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0041] Reference Figure 1 , Figure 2 and Figure 3 , Figure 1 This is a schematic diagram of the state structure of the head-mounted device in this application. Figure 1 As shown, the head-mounted device 01 includes a support 30, and a voice playback component 20 and a voice pickup component 10 mounted on the support 30. The voice pickup component 10 includes:
[0042] Bone conduction sensor 1, the bone conduction sensor 1 being used to pick up voice feature data in the environment;
[0043] The elastic damping mechanism 2 has the bone conduction sensor 1 located at one end of the elastic damping mechanism 2 in the maximum elastic direction, and the elastic damping mechanism 2 has the largest elastic modulus in the maximum elastic direction.
[0044] Rotating mechanism 3 is used to control the rotation of the elastic damping mechanism 2 to change the direction of the maximum elasticity.
[0045] For example, the head-mounted device 01 can be AR glasses, VR glasses, or voice glasses, etc. A voice playback component 20 and a voice pickup component 10 are mounted on the support 30 of the head-mounted device 01. The voice playback component 20 can refer to a speaker or horn, etc. The bone conduction sensor 1 in the voice pickup component 10 is typically a VPU, and the elastic damping mechanism 2 can be a spring, for example... Figure 2 The cylindrical helical spring 21 or Figure 3A conical helical spring 22, etc. It is understood that for a spring, the direction in which its elastic modulus is greatest, or its direction of maximum elasticity, is along its cylindrical or conical central axis. The bone conduction sensor 1 is mainly used to pick up speech feature data in the environment, for example, picking up vibration signal data generated by vibrations transmitted through bones or skin when a user speaks. The bone conduction sensor 1 is located at one end of the elastic damping mechanism 2 in the direction of maximum elasticity, for example, fixed at one end of the central axis of the cylindrical helical spring 21 or the conical helical spring 22. The rotating mechanism 3 may include a transmission rod and a motor or micro-motor. The rotating mechanism 3 is also connected to the elastic damping mechanism 2, and the rotation of the elastic damping mechanism 2 can be controlled by the rotating mechanism 3, thereby changing the direction of the maximum elasticity of the elastic damping mechanism 2.
[0046] Understandably, compared to existing traditional solutions that attach the VPU to the outer shell of the head-mounted device or to the motherboard, in this embodiment, the voice pickup component 10, including the bone conduction sensor 1, is mounted on the head-mounted device's bracket 30. In the voice pickup component 10, the bone conduction sensor 1 is connected to the rotation mechanism 3 through the elastic damping mechanism 2. That is, there is an elastic damping mechanism 2 in the process of connecting the bone conduction sensor 1 to the head-mounted device, and the elastic damping mechanism 2 has a maximum elastic direction. When the bone conduction sensor 1 actually picks up external speech, that is, when it picks up vibrations transmitted through bones and skin, the vibrations are transmitted to the bone conduction sensor 1 through the elastic damping mechanism 2. Since the elastic damping mechanism 2 is elastic, the attenuation of vibrations in the elastic damping mechanism 2 can be reduced. In particular, the vibration attenuation is the least in the direction of maximum elasticity of the elastic damping mechanism 2. Accordingly, more vibrations in the direction of maximum elasticity are picked up and more signals are generated. Conversely, the vibration attenuation in the non-maximum elasticity direction is greater, and accordingly, less vibrations in the non-maximum elasticity direction are picked up and fewer signals are generated. Thus, the elastic damping mechanism 2 can be used to selectively filter vibrations. Moreover, in practical applications, the rotation mechanism 3 can be used to control the rotation of the elastic damping mechanism 2, thereby changing the direction of the maximum elasticity direction so that the direction of the maximum elasticity direction is the same as the vibration direction of the desired vibration (e.g., the sound of speech or the sound of the voice playback component, the vibration direction of the speech or the sound of the voice playback component can be predetermined). This achieves targeted speech pickup, improves the signal-to-noise ratio of the speech pickup result, and reduces the processing difficulty of subsequent audio algorithms.
[0047] like Figure 2 The image shown is a side view of one embodiment of the voice pickup component in this application. In the voice pickup component 10, the elastic damping mechanism 2 includes a first connection end 201 connected to the bone conduction sensor 1 and a second connection end 202 connected to the rotation mechanism 3, wherein the direction of the line connecting the first connection end 201 and the second connection end 202 is the maximum elastic direction.
[0048] For example, in this embodiment, the elastic damping mechanism 2 of the voice pickup component 10 is a cylindrical helical spring 21. The cylindrical helical spring 21 includes a first connecting end 201 connected to the bone conduction sensor 1 and a second connecting end 202 connected to the rotation mechanism 3. In practical applications, the connection direction between the center point of the first connecting end 201 and the center point of the second connecting end 202 (i.e., the central axis of the cylindrical helical spring 21) is the direction of maximum elasticity. It can be understood that since the bone conduction sensor 1 and the rotation mechanism 3 are respectively connected to the two ends of the cylindrical helical spring 21 in the direction of maximum elasticity, the vibration can be guaranteed to propagate along the direction of maximum elasticity, thereby reducing the attenuation of vibration.
[0049] like Figure 3 The image shown is a side view of another embodiment of the voice pickup component in this application. In this embodiment, the elastic damping mechanism 2 is a conical helical spring 22. Similarly, the conical helical spring 22 includes a first connecting end 201 connected to the bone conduction sensor 1 and a second connecting end 202 connected to the rotation mechanism 3. The spring diameter of the first connecting end 201 is smaller than the spring diameter of the second connecting end 202. That is, the bone conduction sensor 1 is fixed to the end of the conical helical spring 22 with the smaller diameter, and the rotation mechanism 3 is fixed to the end of the conical helical spring 22 with the larger diameter. It can be understood that when the bone conduction sensor 1 and the rotation mechanism 3 are connected via the conical helical spring 22, since the bone conduction sensor 1 is fixed to the end of the conical helical spring 22 with the smaller diameter, the bending of the conical helical spring 22 due to gravity when the bone conduction sensor 1 is close to a "suspended state" can be reduced, thus reducing the impact on the pickup effect of the bone conduction sensor 1 when picking up external speech.
[0050] Reference Figure 3 The rotating mechanism 3 of the voice pickup component 10 includes a stepper motor 31 and a drive shaft 32 connected to the stepper motor. The drive shaft 32 is fixed to the second connecting end 202 of the conical helical spring 22.
[0051] For example, the transmission mechanism 3 includes a stepper motor 31 and a transmission shaft 32 connected to the stepper motor 31. The transmission shaft 32 is also fixed to the second connection end 202 of the conical helical spring 22. In practical applications, the stepper motor 31 drives the transmission shaft 32 to rotate, and the conical helical spring 22 rotates with the transmission shaft 32, with the maximum elastic direction of the conical helical spring 22 changing synchronously. Furthermore, referring to… Figure 4 This is a side view of another posture of the voice pickup component in another embodiment of this application. (See attached image.) Figure 4 As shown, the maximum elastic direction of the conical helical spring 22 is perpendicular to the current viewing plane.
[0052] Reference Figure 1 The head-mounted device 01 also includes a voice playback component 20 mounted on a bracket 30. The angle between the mechanical vibration direction of the voice playback component 20 and a preset voice vibration direction is greater than or equal to a first preset angle threshold, wherein the mechanical vibration direction is the vibration direction of the mechanical vibration generated by the voice playback component 20.
[0053] For example, the aforementioned voice playback component 20 is used for audio output. The second angle between the vibration direction of the mechanical vibration generated by the voice playback component 20 (e.g., the vibration direction of the vibrating sheet in the voice playback component 20) and a preset voice vibration direction is greater than a second preset angle threshold. In practical applications, it is preferable to configure the voice playback component 20 such that the vibration direction of its generated mechanical vibration is perpendicular to the preset voice vibration direction (e.g., ...). Figure 1 As shown, the mechanical vibration direction of the voice playback component is perpendicular to the current viewing angle plane, while the preset voice vibration direction is located in the current viewing angle plane. Therefore, the mechanical vibration direction of the voice playback component is perpendicular to the preset voice vibration direction (where the preset voice vibration direction usually refers to the vibration direction caused by the user speaking). That is, the angle between the above-mentioned mechanical vibration direction and the preset voice vibration direction is 90°. Therefore, the second preset angle threshold can be within a preset range around 90°. It can be understood that the explanation uses an angle of 90° between the mechanical vibration direction and the preset voice vibration direction. Since the mechanical vibration direction and the preset voice vibration direction are perpendicular, when picking up voice feature data, it can avoid the maximum elastic direction of the elastic damping mechanism 2 being consistent with multiple sound vibration directions at the same time, thereby improving the signal-to-noise ratio of the voice feature data picked up by the bone conduction sensor.
[0054] Reference Figure 5 The voice pickup component 10 in the head-mounted device 01 also includes an air conduction sensor 4. It is understood that the air conduction sensor 4 can be a microphone. Typically, the air conduction sensor 4 is used to pick up sound wave vibrations propagating through the air and generate corresponding voice data. In this embodiment, the air conduction sensor 4 can be set at any position on the support.
[0055] like Figure 6 As shown, Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of this application.
[0056] The head-mounted devices in this application include, but are not limited to, Mixed Reality (MR) devices (e.g., MR glasses or MR helmets), Augmented Reality (AR) devices (e.g., AR glasses or AR helmets), Virtual Reality (VR) devices (e.g., VR glasses or VR helmets), Extended Reality (XR) devices, or some combination thereof.
[0057] like Figure 6 As shown, the head-mounted device may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0058] Optionally, the device may also include a camera, RF (Radio Frequency) circuitry, audio circuitry, a WiFi module, and so on. These will not be elaborated upon further here.
[0059] Those skilled in the art will understand that Figure 6 The head-mounted device structure shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0060] like Figure 6 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a voice pickup program.
[0061] exist Figure 6 In the head-mounted device shown, the network interface 1004 is mainly used to connect to the backend server and communicate data with it; the user interface 1003 is mainly used to connect to the client (user terminal) and communicate data with it; and the processor 1001 can be used to call the voice pickup program stored in the memory 1005 and perform the following operations:
[0062] Based on the ambient noise intensity of the head-mounted device, determine the target orientation of the maximum elastic direction of the elastic damping mechanism in the head-mounted device;
[0063] Based on the rotation mechanism connected to the elastic damping mechanism in the head-mounted device, the rotation of the elastic damping mechanism is controlled to adjust the maximum elastic direction of the elastic damping mechanism to the target direction;
[0064] Voice feature data is obtained by picking up voice features in the environment through a bone conduction sensor located on the elastic damping mechanism.
[0065] Furthermore, the processor 1001 can call the voice pickup program stored in the memory 1005 and also perform the following operations:
[0066] The target orientation includes a first target orientation, and the step of determining the target orientation of the maximum elastic direction of the elastic damping mechanism in the head-mounted device based on the ambient noise intensity of the head-mounted device includes:
[0067] If the ambient noise intensity is greater than a preset intensity threshold, then the target direction is the first target direction, wherein after the maximum elastic direction of the elastic damping mechanism is adjusted to the first target direction, the angle between the maximum elastic direction and the preset voice vibration direction is less than or equal to a second preset angle threshold.
[0068] Furthermore, the processor 1001 can call the voice pickup program stored in the memory 1005 and also perform the following operations:
[0069] The target orientation further includes a second target orientation, and the step of determining the target orientation of the maximum elastic direction of the elastic damping mechanism in the head-mounted device based on the environmental noise intensity of the head-mounted device further includes:
[0070] If the ambient noise intensity is less than or equal to a preset sound intensity threshold, then the target direction is the second target direction, wherein, after adjusting the maximum elastic direction of the elastic damping mechanism to the second target direction, the angle between the maximum elastic direction and the mechanical vibration direction of the voice playback component in the head-mounted device is less than or equal to a second preset angle threshold.
[0071] Furthermore, the processor 1001 can call the voice pickup program stored in the memory 1005 and also perform the following operations:
[0072] The types of the speech feature data include first speech feature data picked up by the bone conduction sensor in the posture corresponding to the first target direction, and second speech feature data picked up by the bone conduction sensor in the posture corresponding to the second target direction. After the step of obtaining speech feature data by picking up speech features in the environment through the bone conduction sensor located on the elastic damping mechanism, the method includes:
[0073] The voice data collected by the air conduction sensor in the head-mounted device is subjected to noise reduction processing using the first voice feature data.
[0074] Alternatively, echo cancellation processing can be performed on the voice data collected by the air conduction sensor in the head-mounted device using the second voice feature data.
[0075] Reference Figure 7 The first embodiment of the voice pickup method of this application is applied to the head-mounted device described above, with reference to... Figure 1 The voice pickup component 10 is mounted on the support 30 of the head-mounted device 01. The head-mounted device 01 can be AR glasses, VR glasses, or voice glasses. The bone conduction sensor 1 in the voice pickup component 10 is connected to the rotation mechanism 3 through the elastic damping mechanism 2. Vibration signals from the environment are transmitted to the bone conduction sensor 1 through the elastic damping mechanism 2. In addition, the bone conduction sensor 1 and the rotation mechanism 3 are connected to the PCB (Printed Circuit Board) in the support 30 by wiring (for power supply or communication).
[0076] The speech pickup method includes:
[0077] Step S10: Based on the ambient noise intensity of the head-mounted device, determine the target direction of the maximum elastic direction of the elastic damping mechanism in the head-mounted device;
[0078] It should be noted that, in this embodiment, the implementing entity of the above method can be a head-mounted device. In addition to bone conduction sensors, the head-mounted device can also be equipped with air conduction sensors, such as a microphone. The microphone can collect ambient sound around the head-mounted device, and the ambient noise intensity can be determined based on the magnitude of the ambient sound (e.g., the decibel level). Then, based on the ambient noise intensity, the target direction of the maximum elastic direction of the elastic damping mechanism in the head-mounted device is determined. It is understood that different directions of the maximum elastic direction of the elastic damping mechanism in the head-mounted device will result in different vibration objects being picked up by the head-mounted device. In practical applications, technicians can pre-determine the vibration direction of the targeted vibration object during propagation, and then set the target direction based on the vibration direction. Similarly, the mapping relationship between ambient noise intensity and target direction can be set by technicians according to actual needs, and no specific limitations are imposed here.
[0079] Step S20: Based on the rotation mechanism connected to the elastic damping mechanism in the head-mounted device, control the rotation of the elastic damping mechanism to adjust the maximum elastic direction of the elastic damping mechanism to the target direction;
[0080] For example, different control commands can be generated based on the ambient noise intensity. The rotating mechanism connected to the elastic damping mechanism in the head-mounted device can control the rotation of the elastic damping mechanism according to the corresponding control command, that is, rotate the maximum elastic direction of the elastic damping mechanism to the target direction.
[0081] For example, refer to Figure 8 This is a schematic diagram of the control framework for the rotating mechanism in this application. For example, the rotating mechanism includes a motor (which can also be a stepper motor). The motor driver controls the motor according to the control signals sent by the SOC (System on Chip), thereby controlling the rotation of the elastic damping mechanism. The control signals are generated by the SOC based on the audio collected by the MIC and its built-in control signal generation logic.
[0082] Step S30: Voice feature data is obtained by picking up voice features in the environment through a bone conduction sensor located on the elastic damping mechanism.
[0083] For example, when the elastic damping mechanism is rotated to a position where the maximum elastic direction is consistent with the target direction, the voice features in the environment (i.e., vibrations in the environment, such as vibrations transmitted through bones or skin, or vibrations transmitted through a head-mounted device) can be picked up by the bone conduction sensor connected to the other end of the elastic damping mechanism to obtain voice feature data.
[0084] It is understood that in this embodiment, the head-mounted device includes a support frame, and a voice playback component and a voice pickup component mounted on the support frame. The bone conduction sensor of the voice pickup component is used to pick up voice feature data from the environment. The bone conduction sensor is located at one end of the elastic damping mechanism of the voice pickup component in the direction of maximum elasticity. That is, an elastic damping mechanism exists during the connection between the bone conduction sensor and the head-mounted device, and this mechanism has a direction of maximum elasticity. When the bone conduction sensor picks up vibrations from external bones, skin, or the head-mounted device, the vibration is transmitted to the bone conduction sensor through the elastic damping mechanism. Because the elastic damping mechanism is elastic, it reduces the attenuation of vibrations within the mechanism. Specifically, the smaller the vibration attenuation in the direction of maximum elasticity, the more signals are generated in that direction. Conversely, the greater the vibration attenuation in directions other than maximum elasticity, the less signals are generated. Thus, the elastic damping mechanism can filter vibrations. In practical applications, the target direction of the maximum elasticity in the elastic damping mechanism is first determined based on the noise level in the environment. Then, the rotation mechanism is used to control the rotation of the elastic damping mechanism so that the maximum elasticity direction in the elastic damping mechanism is consistent with the target direction. This allows for targeted speech pickup, that is, picking up the desired speech as much as possible, thereby improving the signal-to-noise ratio of the speech pickup results and reducing the processing difficulty of subsequent audio algorithms.
[0085] In one feasible implementation, the target orientation includes a first target orientation, and the step of determining the target orientation of the maximum elastic direction of the elastic damping mechanism in the head-mounted device based on the ambient noise intensity of the head-mounted device includes:
[0086] Step S110: If the ambient noise intensity is greater than a preset intensity threshold, then the target direction is the first target direction, wherein after the maximum elastic direction of the elastic damping mechanism is adjusted to the first target direction, the angle between the maximum elastic direction and the preset voice vibration direction is less than or equal to a second preset angle threshold.
[0087] It should be noted that this embodiment is mainly applied to head-mounted devices, taking a two-way call scenario of a head-mounted device as an example. In this scenario, there are two parties involved in the call: the user of the head-mounted device and the user's contact person. It is understood that the ambient noise intensity when the user speaks to the contact person through the head-mounted device is significantly greater than the ambient noise intensity when the user listens to the contact person speaking through the head-mounted device. Therefore, the aforementioned preset intensity threshold can be set by a technician based on the sound intensity difference between the two scenarios. When the ambient noise intensity is greater than the preset intensity threshold, it actually represents the scenario where the user speaks to the contact person through the head-mounted device. Therefore, in this scenario, the expected voice to be picked up should be the user's voice. Correspondingly, the target direction of the maximum elastic direction is the first target direction. Wherein, after adjusting the maximum elastic direction of the elastic damping mechanism to the first target direction, the angle between the maximum elastic direction and the preset voice vibration direction is less than or equal to the second preset angle threshold. It is worth noting that the preset voice vibration direction is the vibration direction of the sound wave when the user speaks. Therefore, the maximum elastic direction and the preset voice vibration direction are preferably parallel, that is, the angle between the maximum elastic direction and the preset voice vibration direction is 0°. Therefore, the first preset angle threshold can be within the preset range around 0°.
[0088] Taking the maximum elastic direction and the preset voice vibration direction as an example, in this posture, the elastic damping mechanism has the least attenuation of the sound wave vibration when the user speaks. Therefore, it can selectively pick up the user's voice, so that the noise ratio in the picked-up user voice is minimized, which improves the signal-to-noise ratio.
[0089] If the head-mounted device is AR glasses, according to actual test results, the vibration caused by the user speaking is strongest in the vertical direction of the temples, while it is relatively weaker in other directions, such as... Figure 1 The direction of the vocal vibration is shown, and at this time, the first target can be... Figure 1 The maximum elastic direction of the conical helical spring shown.
[0090] In one feasible implementation, the target pointing further includes a second target pointing, and the step of determining the target pointing of the maximum elastic direction of the elastic damping mechanism in the head-mounted device based on the ambient noise intensity of the head-mounted device further includes:
[0091] In step S120, if the ambient noise intensity is less than or equal to a preset sound intensity threshold, the target direction is the second target direction. After adjusting the maximum elastic direction of the elastic damping mechanism to the second target direction, the angle between the maximum elastic direction and the mechanical vibration direction of the voice playback component in the head-mounted device is less than or equal to the second preset angle threshold.
[0092] It should be noted that the head-mounted device also includes a voice playback component, and the angle between the mechanical vibration direction of the voice playback component and the preset voice vibration direction is greater than or equal to a first preset angle threshold, wherein the mechanical vibration direction is the vibration direction of the mechanical vibration generated by the voice playback component. Similarly, based on the aforementioned two-way communication scenario, the user is speaking while the voice playback component is outputting, and the expected voice to be picked up while the user is speaking should be the content of the user's speech; therefore, the output of the voice playback component can be considered noise. Taking the example where the direction of the mechanical vibration generated by the voice playback component is perpendicular to the preset voice vibration direction (i.e., the first preset angle threshold is 90°, and the angle between the mechanical vibration direction and the preset voice vibration direction is equal to 90°), since the maximum elastic direction of the elastic damping mechanism in the head-mounted device is parallel to the vibration direction of the user's voice (i.e., the preset voice vibration direction) when speaking to the user, the attenuation of the user's voice vibration is very small during the transmission of the vibration through the elastic damping mechanism. Since the direction of the mechanical vibration generated by the voice playback component is perpendicular to the preset voice vibration direction, the maximum elastic direction of the elastic damping mechanism is also perpendicular to the direction of the mechanical vibration generated by the voice playback component. Therefore, when the elastic damping mechanism transmits the mechanical vibration generated by the voice playback component, it will significantly attenuate the mechanical vibration generated by the voice playback component. In this way, when both the voice playback component and the user vibrate simultaneously, as much of the user's speech content as possible can be picked up while as little of the content output by the voice playback component is picked up, thereby improving the signal-to-noise ratio of the head-mounted device's pickup results.
[0093] Additionally, when the ambient noise intensity is less than or equal to a preset sound intensity threshold, the target direction is the second target direction. After adjusting the maximum elastic direction of the elastic damping mechanism to the second target direction, the angle between the maximum elastic direction and the mechanical vibration direction of the voice playback component in the head-mounted device will be less than or equal to a second preset angle threshold. For example, if the second preset angle threshold is 0°, the maximum elastic direction is parallel to the mechanical vibration direction of the voice playback component. Similarly, the second preset angle threshold can be within a preset range around 0°. For example, based on the above two-way communication scenario, when the user's contact speaks, according to normal life habits, the user mainly listens to the content of the contact's speech and does not speak themselves. Therefore, in this scenario, the ambient noise of the head-mounted device should be smaller than the user's own speech. In this scenario, the expected speech should be the content output by the voice playback component (for echo cancellation), so the maximum elastic direction of the elastic damping mechanism is adjusted to be parallel to the mechanical vibration direction of the voice playback component.
[0094] Taking the direction of maximum elasticity parallel to the vibration direction of the voice playback component as an example, under this posture, the elastic damping mechanism attenuates the mechanical vibration output of the voice playback component to the minimum. Therefore, the output of the voice playback component can be picked up in a targeted manner, which means that the noise ratio in the picked-up content of the voice playback component output is minimized, thus achieving an improvement in the signal-to-noise ratio.
[0095] like Figure 9 The diagram shown is a schematic diagram of another state structure of the head-mounted device in this application. In the diagram, the conical rotary spring 22 is in another posture (i.e., the maximum elastic direction of the conical rotary spring 22 is rotated to be consistent with the direction of the second target), and in this posture, the maximum elastic direction of the conical rotary spring 22 is parallel to the mechanical vibration direction of the voice playback component 20.
[0096] In one feasible implementation, the type of the speech feature data includes first speech feature data picked up by the bone conduction sensor in the posture corresponding to the first target direction, and second speech feature data picked up by the bone conduction sensor in the posture corresponding to the second target direction. After the step of obtaining speech feature data by picking up speech features in the environment through the bone conduction sensor located on the elastic damping mechanism, the method includes:
[0097] Step S310: Using the first voice feature data, perform noise reduction processing on the voice data collected by the air conduction sensor in the head-mounted device;
[0098] Step S320, or, using the second voice feature data, perform echo cancellation processing on the voice data collected by the air conduction sensor in the head-mounted device.
[0099] For example, the voice pickup component in the head-mounted device also includes an air conduction sensor, which can be a microphone, for collecting voice data from the environment. Furthermore, the types of voice feature data collected by the bone conduction sensor include first voice feature data picked up by the bone conduction sensor in a posture corresponding to the first target direction, and second voice feature data picked up by the bone conduction sensor in a posture corresponding to the second target direction. In addition to determining the magnitude of ambient noise, the air conduction sensor can also be used for voice content acquisition. If the voice feature data picked up by the bone conduction sensor is first voice feature data, then the voice data collected by the microphone is subjected to noise reduction processing based on the first voice feature data (which mainly consists of the voice features of the user's speech content) (e.g., the portion of the voice data that does not match the first voice feature data is treated as noise). If the voice feature data picked up by the bone conduction sensor is second voice feature data (which mainly consists of the voice features of the content output by the voice playback component), the voice data collected by the microphone is subjected to echo cancellation processing (e.g., the portion of the voice data that matches the second voice feature data is treated as noise).
[0100] For example, refer to Figure 10 This diagram illustrates the speech feature data processing framework in this application. The input channels for sound mainly include the microphone (MIC) and the voice processing unit (VPU), while the output channels mainly include the sound storage box (SPK). Additionally, the system includes filters, algorithmic fusion processing, AEC (Acoustic Echo Cancel), NS (Noise Suppression), EQ (Equalizer), a Bluetooth module, and an antenna. It is understood that when the VPU picks up the first speech feature data, this data will be input into the algorithmic fusion processing for noise reduction. In this case, the AEC performs echo noise reduction based solely on internal downlink data (e.g., data input to the SPK). Conversely, if the VPU picks up the second speech feature data, this data will be input into the AEC along with the internal downlink data for echo cancellation.
[0101] Furthermore, this application also proposes a readable storage medium, which is a computer-readable storage medium, on which a speech pickup program is stored. When the speech pickup program is executed by a processor, it implements the steps of the speech pickup method described above.
[0102] The specific implementation of the medium in this application is basically the same as the various embodiments of the above-described voice pickup method, and will not be described again here.
[0103] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0104] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0105] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a vehicle, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0106] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A head-mounted device, characterized in that, The system includes a bracket, and a voice playback component and a voice pickup component mounted on the bracket, the voice pickup component comprising: A bone conduction sensor, which is used to pick up voice feature data in the environment; An elastic damping mechanism, wherein the bone conduction sensor is located at one end of the elastic damping mechanism in the direction of maximum elasticity, and the elastic damping mechanism has the maximum elastic modulus in the direction of maximum elasticity; A rotating mechanism is used to control the rotation of the elastic damping mechanism to change the direction of the maximum elasticity. The elastic damping mechanism includes a first connection end connected to the bone conduction sensor and a second connection end connected to the rotation mechanism, wherein the direction of the line connecting the first connection end and the second connection end is the direction of maximum elasticity.
2. The head-mounted device as claimed in claim 1, characterized in that, The elastic damping mechanism is a conical helical spring, and the spring diameter at the first connecting end of the conical helical spring is smaller than the spring diameter at the second connecting end.
3. The head-mounted device as described in claim 2, characterized in that, The rotating mechanism of the head-mounted device includes a stepper motor and a drive shaft connected to the stepper motor, the drive shaft being fixed to the second connecting end of the conical helical spring.
4. The head-mounted device as claimed in claim 1, characterized in that, The angle between the mechanical vibration direction of the voice playback component and the preset voice vibration direction is greater than or equal to a first preset angle threshold, wherein the mechanical vibration direction is the vibration direction of the mechanical vibration generated by the voice playback component.
5. The head-mounted device as claimed in claim 1, characterized in that, The voice pickup component also includes a gas conduction sensor.
6. A speech pickup method, characterized in that, The voice pickup method is applied to the head-mounted device as described in any one of claims 1 to 5; The speech pickup method includes: Based on the ambient noise intensity of the head-mounted device, determine the target orientation of the maximum elastic direction of the elastic damping mechanism in the head-mounted device; Based on the rotation mechanism connected to the elastic damping mechanism in the head-mounted device, the rotation of the elastic damping mechanism is controlled to adjust the maximum elastic direction of the elastic damping mechanism to the target direction; Voice feature data is obtained by picking up voice features in the environment through a bone conduction sensor located on the elastic damping mechanism.
7. The speech pickup method as described in claim 6, characterized in that, The target orientation includes a first target orientation, and the step of determining the target orientation of the maximum elastic direction of the elastic damping mechanism in the head-mounted device based on the ambient noise intensity of the head-mounted device includes: If the ambient noise intensity is greater than a preset intensity threshold, then the target direction is the first target direction, wherein after the maximum elastic direction of the elastic damping mechanism is adjusted to the first target direction, the angle between the maximum elastic direction and the preset voice vibration direction is less than or equal to a second preset angle threshold.
8. The speech pickup method as described in claim 7, characterized in that, The target orientation further includes a second target orientation, and the step of determining the target orientation of the maximum elastic direction of the elastic damping mechanism in the head-mounted device based on the environmental noise intensity of the head-mounted device further includes: If the ambient noise intensity is less than or equal to a preset sound intensity threshold, then the target direction is the second target direction, wherein, after adjusting the maximum elastic direction of the elastic damping mechanism to the second target direction, the angle between the maximum elastic direction and the mechanical vibration direction of the voice playback component in the head-mounted device is less than or equal to a second preset angle threshold.
9. The speech pickup method as described in claim 8, characterized in that, The types of the speech feature data include first speech feature data picked up by the bone conduction sensor in the posture corresponding to the first target direction, and second speech feature data picked up by the bone conduction sensor in the posture corresponding to the second target direction. After the step of obtaining speech feature data by picking up speech features in the environment through the bone conduction sensor located on the elastic damping mechanism, the method includes: The voice data collected by the air conduction sensor in the head-mounted device is subjected to noise reduction processing using the first voice feature data. Alternatively, echo cancellation processing can be performed on the voice data collected by the air conduction sensor in the head-mounted device using the second voice feature data.
10. A readable storage medium, characterized in that, The readable storage medium is a computer-readable storage medium, and a speech pickup program is stored on the readable storage medium. When the speech pickup program is executed by a processor, it implements the steps of the speech pickup method as described in any one of claims 7 to 9.
Citation Information
Patent Citations
Bone conduction loudspeaker, intelligent wearing device and resonance processing method
CN109068248A
Vibration sensor
CN113286213A