Intelligent device adaptive sound pickup methods, equipment, media and software products
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-08-14
AI Technical Summary
[0033]本申请实施例提供的智能设备自适应拾音方法、设备、介质及程序产品,通过姿态传感器采集智能设备的三维姿态数据,并根据该三维姿态数据判断智能设备的佩戴方向,能够及时识别智能设备在正戴、侧戴等不同佩戴状态下的方向变化;进而通过根据佩戴方向调整麦克风阵列的权重矩阵,能够使麦克风阵列的拾音指向与用户嘴部相对位置更好匹配,减少因佩戴方向变化导致的语音灵敏度下降及环境噪声混入,进而提高智能设备拾音一致性和通话稳定性。
Smart Images

Figure CN122579025A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice processing for smart wearable devices, and in particular to an adaptive voice pickup method, device, medium, and program product for smart devices. Background Technology
[0002] With the advancement of technology, smart devices, especially smartwatches, are widely used in fields such as communication, health monitoring, and sports tracking, and their voice call function has become one of the core requirements. Smartwatches need to pick up sound in different wearing orientations to achieve voice call functionality. Therefore, how to improve the consistency of sound pickup and call stability of smart devices under changing wearing orientations has become a pressing technical problem to be solved.
[0003] Currently, existing methods for picking up voice during calls in smart devices typically use fixed microphone arrays. The distance between the user's mouth and the microphone is estimated by the acoustic characteristics of the microphone array (such as sound pressure level differences and time differences), and the gain or filtering parameters are dynamically adjusted.
[0004] However, existing smart device call pickup methods are prone to deviation from the mouth position due to changes in wearing orientation, resulting in decreased voice sensitivity and the introduction of more environmental noise. Some devices rely on manual mode switching, which not only imposes a heavy operational burden but is also susceptible to judgment errors caused by posture fluctuations, resulting in unstable call quality. Summary of the Invention
[0005] This application provides an adaptive voice pickup method, device, medium, and program product for smart devices to solve the aforementioned technical problems. This method addresses application scenarios where the voice pickup performance of smart devices is prone to fluctuations under different wearing orientations. By sensing the current spatial posture of the device and identifying its wearing orientation, the voice pickup strategy can adaptively adjust according to changes in wearing status, thereby improving the consistency of voice acquisition and call stability of smart devices under changing orientations.
[0006] In a first aspect, embodiments of this application provide an adaptive sound pickup method for smart devices, the method comprising:
[0007] Three-dimensional attitude data of intelligent devices are collected through attitude sensors, including gyroscopes and accelerometers; the three-dimensional attitude data includes pitch angle data, roll angle data, and yaw angle data.
[0008] Determine the wearing orientation of smart devices based on three-dimensional posture data;
[0009] Adjust the weight matrix of the microphone array according to the wearing direction.
[0010] In one possible implementation, dynamically adjusting the weight matrix of the microphone array according to the wearing direction includes:
[0011] When the wearing orientation is upright, a weight matrix is loaded to make the microphone closer to the user's mouth have a higher weight than the microphone farther away from the user's mouth;
[0012] When the wearing direction is side-mounted, a side-mounted weight matrix is loaded, so that the microphone closer to the user's mouth has a higher weight than the microphone farther away from the user's mouth.
[0013] In one possible implementation, before dynamically adjusting the weight matrix of the microphone array according to the wearing orientation, the following steps are included:
[0014] Set the duration threshold for determining the wearing direction;
[0015] If the three-dimensional posture data continuously meets the duration threshold, the corresponding wearing direction is determined.
[0016] In one possible implementation, before dynamically adjusting the weight matrix of the microphone array according to the wearing orientation, the following steps are included:
[0017] The correlation data between the user's wearing direction and the human voice test signal is obtained through the user calibration function;
[0018] Optimize the weight matrix based on the associated data.
[0019] In one possible implementation, the three-dimensional attitude data of the smart device is acquired via an attitude sensor, including:
[0020] The Kalman filter algorithm is used to fuse the raw data from the attitude sensor and output three-dimensional attitude data.
[0021] In one possible implementation, dynamically adjusting the weight matrix of the microphone array according to the wearing direction further includes:
[0022] Expand the preset threshold range for multi-mode wearing directions;
[0023] Historical posture data is analyzed using a machine learning model, and the upper and lower limits of the preset threshold range for multi-mode wearing direction are dynamically optimized based on the analysis results; the machine learning model is a posture analysis model based on a neural network.
[0024] In one possible implementation, after dynamically adjusting the weight matrix of the microphone array according to the wearing direction, the method further includes:
[0025] The acoustic signals collected by the microphone array are analyzed by the environmental noise detection module, and the microphone gain is adjusted in conjunction with the voice activity detection module.
[0026] Secondly, embodiments of this application provide an adaptive sound pickup device for smart devices, the device comprising:
[0027] The data acquisition module is used to acquire three-dimensional attitude data of the smart device through attitude sensors, including gyroscopes and accelerometers; the three-dimensional attitude data includes pitch angle data, roll angle data, and yaw angle data.
[0028] The judgment module is used to determine the wearing direction of the smart device based on three-dimensional posture data;
[0029] The adjustment module is used to adjust the weight matrix of the microphone array according to the wearing direction.
[0030] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0031] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0032] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0033] The adaptive sound pickup method, device, medium, and program product for smart devices provided in this application collect three-dimensional posture data of the smart device through a posture sensor, and determine the wearing direction of the smart device based on the three-dimensional posture data. It can promptly identify the directional changes of the smart device in different wearing states such as wearing it upright or sideways. Furthermore, by adjusting the weight matrix of the microphone array according to the wearing direction, the sound pickup direction of the microphone array can be better matched with the relative position of the user's mouth, reducing the decrease in voice sensitivity and the mixing of environmental noise caused by changes in wearing direction, thereby improving the sound pickup consistency and call stability of the smart device. Attached Figure Description
[0034] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0035] Figure 1 A flowchart illustrating an adaptive sound pickup method for a smart device provided in an embodiment of this application;
[0036] Figure 2 A flowchart illustrating an adaptive sound pickup method for side-wearing / front-wearing watches provided in this application embodiment;
[0037] Figure 3 This is a schematic diagram of the structure of an adaptive sound pickup device for a smart device provided in an embodiment of this application;
[0038] Figure 4 This is a schematic diagram of the structure of an electronic device provided in this application.
[0039] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0040] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0041] In the embodiments of this application, the use of terms such as "first" and "second" is to distinguish between identical or similar items that have essentially the same function and effect. For example, "first electronic device" and "second electronic device" are merely used to distinguish different electronic devices and do not limit their order of execution. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.
[0042] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.
[0043] The following is an explanation of some terms used in the embodiments of this application:
[0044] Smart devices refer to portable smart wearable devices that include posture sensors, dual microphone arrays, and data processing modules, and have call pickup capabilities. Examples include smartwatches, smart bracelets with call functionality, and smart wristbands.
[0045] Pitch angle: refers to the angle at which a smart device rotates around the X-axis. For example, in the case of a smart watch, which is the core application carrier of a smart device, the pitch angle can represent the angle between the watch face and the horizontal plane.
[0046] Roll angle: refers to the angle at which a smart device rotates around the Z-axis. In the case of a smart watch, which is the core application carrier of a smart device, the roll angle can indicate the direction in which the watch face is flipped.
[0047] Wearing direction: This refers to the spatial orientation of the smart device when worn by the user, including two basic modes: wearing it upright and wearing it to the side. For example, when the core application carrier of the smart device is a smartwatch, wearing it upright means that the front of the watch face faces the outside of the wrist, while wearing it to the side means that the watch face faces the inside or side of the wrist.
[0048] Smart devices, as wearable communication terminals, are widely used in scenarios such as voice calls, message interaction, health monitoring, and exercise recording. Among these, call pickup capability directly affects the device's practicality and user experience. In actual use, users do not have a fixed way of wearing smart devices. They may wear them upright with the main body facing the outside of their wrist, or they may switch to wearing them sideways with the main body facing the inside or side of their wrist due to adjustments in the tightness of the fixing components, wrist movement habits, clothing obstruction, or changes in call posture. Under different wearing directions, the relative position between the user's mouth and the microphone array will change significantly, thus affecting the pickup direction and voice propagation path. Therefore, how to maintain stable voice acquisition effect under dynamic changes in wearing direction has become a key issue in the call application of smart devices.
[0049] Currently, existing methods for picking up voice signals during calls in smart devices typically employ fixed microphone arrays. These arrays estimate the distance between the user's mouth and the microphone based on their acoustic characteristics (such as sound pressure level differences and time differences), and then dynamically adjust gain or filtering parameters. However, in real-world usage environments, the user's wearing orientation is often not static. As the wearing orientation changes, the relative orientation of the microphone array also shifts away from the user's mouth, resulting in a decrease in the energy of the microphone's voice signal, which should be prioritized for voice pickup, while environmental noise and echoes are amplified.
[0050] While some devices offer the function of manually switching between upright and side-wearing modes, this method relies on the user to actively judge the current wearing status and perform the operation, which not only increases the burden of use, but also easily causes misjudgment when the posture fluctuates for a short time or the wearing status is between upright and side-wearing, resulting in the weight configuration being inconsistent with the actual wearing direction.
[0051] Therefore, existing smart device voice pickup methods are not adaptable enough to the dynamic changes in wearing direction, the increasing demand for automated operation, and the coexistence of personalized wearing differences, resulting in inconsistent voice pickup and poor voice clarity during calls.
[0052] In view of the aforementioned problems with existing smart device voice pickup methods, this application proposes an adaptive voice pickup method for smart devices. This method acquires three-dimensional posture data of the smart device using a posture sensor, determines the wearing direction of the smart device using the three-dimensional posture data, and adjusts the weight matrix of the microphone array according to the wearing direction. This method uses the device's own posture information to complete the wearing state recognition and voice pickup strategy matching, enabling microphones closer to the user's mouth to receive higher weights under different wearing methods, thereby improving the consistency of voice acquisition and call stability of the smart device under changing orientation.
[0053] The technical solutions of this application will be described in detail below with reference to specific embodiments. The specific embodiments described below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0054] The subject executing this adaptive sound pickup method for smart devices can be, for example, a smartwatch, a smart bracelet, a wearable terminal with call functionality, or other smart devices with microphone arrays and posture detection capabilities.
[0055] Figure 1 This is a flowchart illustrating an adaptive sound pickup method for smart devices provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes:
[0056] S101 collects three-dimensional attitude data of the smart device through an attitude sensor, which includes a gyroscope and an accelerometer. The three-dimensional attitude data includes pitch angle data, roll angle data and yaw angle data.
[0057] For example, an attitude sensor can be a sensor used to detect the orientation and motion state of a smart device in three-dimensional space.
[0058] For example, three-dimensional pose data can be parameters of the pose of a smart device in three-dimensional space.
[0059] For example, a gyroscope can be an inertial device used to measure the angular velocity of a device about its reference axes, and its output reflects the rotational changes of the device over a short period of time.
[0060] For example, an accelerometer is a sensor used to measure the acceleration component of a device, which can stably reflect information related to the direction of gravity in everyday wear scenarios.
[0061] For example, pitch angle data can represent the pitch state of the device relative to the longitudinal reference axis, roll angle data can represent the roll state of the device relative to the lateral reference axis, and yaw angle data can represent the orientation change state of the device in the horizontal plane.
[0062] Optionally, the smart device can periodically read the three-axis angular velocity data output by the gyroscope and the three-axis acceleration data output by the accelerometer based on a preset sampling frequency. For example, the preset sampling frequency can be from 10Hz to 50Hz. When the sampling frequency is lower than 10Hz, the response to the dynamic actions of the smart device may be delayed, and when the sampling frequency is higher than 50Hz, it will increase the processor's computing load and system power consumption.
[0063] In one possible implementation, the smart device can use a Kalman filter algorithm to fuse the raw data from the attitude sensor and output three-dimensional attitude data.
[0064] Optionally, the smart device can periodically read the raw data output by the attitude sensor, perform time synchronization, zero-bias compensation, and dimensional unification processing, and then input it into the Kalman filter module. This Kalman filter module can calculate the predicted value of the current attitude based on the system state transition model, and then combine it with accelerometer observations to correct the prediction error, outputting smoother and more continuous three-dimensional attitude data. This three-dimensional attitude data can be directly used for subsequent wearing orientation determination and provides a stable basis for the dynamic switching of the microphone array weight matrix.
[0065] Optionally, the Kalman filter parameters can be set according to the sampling frequency, noise variance, and motion intensity to adapt to different wearers and different motion states. In practical applications, the Kalman filter module can also be implemented in software integrated into the main control chip, or preprocessed by an independent sensor unit before output; this application embodiment does not limit this.
[0066] Optionally, the intelligent device can perform fusion calculations on the original attitude signal to reduce the impact of gyroscope integral drift on the attitude output, while suppressing the instantaneous fluctuations caused by linear acceleration interference of the accelerometer under dynamic motion conditions, and outputting continuous and stable three-dimensional attitude data. The intelligent device can also use a complementary filtering algorithm to combine the high-frequency response characteristics of the gyroscope with the low-frequency stability characteristics of the accelerometer, so that the device can accurately reflect the current orientation and spatial attitude changes of the device body.
[0067] The above methods can significantly improve the accuracy and anti-interference ability of three-dimensional posture data, reduce the impact of gyroscope drift and instantaneous disturbances of accelerometer on posture calculation, and enable smart devices to stably acquire posture information in complex wearing and movement scenarios, thereby improving the reliability of wearing direction recognition and the response consistency of call pickup control.
[0068] Optionally, the smart device can map the calculated pitch, roll, and yaw angles to a preset angle range and perform time-window smoothing. For example, the processor can calculate a moving average, median, or weighted average based on attitude angle samples from the most recent 0.5 to 2 seconds to obtain stable attitude data at the current moment. For scenarios with rapid angle changes, the smart device can also retain both instantaneous attitude angles and smoothed attitude angles simultaneously. The instantaneous attitude angles are used to detect rapid wrist flicking movements, while the smoothed attitude angles are used to output the current stable wearing state.
[0069] Optionally, the smart device can obtain three-dimensional attitude data frames of pitch angle, roll angle and yaw angle data and store them in the runtime cache.
[0070] S102 determines the wearing direction of the smart device based on three-dimensional posture data.
[0071] For example, the wearing direction can refer to the spatial orientation of the smart device on the user's wrist, which reflects the orientation of the device body, microphone opening, and edge of the device relative to the user's mouth.
[0072] For example, wearing the device in the front can be a way of wearing it with the main body of the device facing the outside of the wrist and the microphone array facing the direction of normal call pickup. Wearing it on the side can be understood as a way of wearing the device in the side that has been significantly flipped or rotated relative to wearing it in the front, so that the side that was originally far away from the mouth is now closer to the mouth.
[0073] Optionally, the smart device can compare the currently collected pitch angle, roll angle, and yaw angle with a pre-established wearing posture feature model and output a classification result indicating whether the device is currently worn upright, worn sideways, or in another intermediate state.
[0074] Optionally, the smart device can establish a rule base for determining the wearing direction when the device is shipped from the factory, during system initialization, or when the user first enables the call enhancement function. For example, the smart device can define the combination of pitch angles between -10 degrees and +10 degrees and roll angles between -10 degrees and +10 degrees as the candidate range for wearing the device upright, and define the combination of pitch angles between -10 degrees and +10 degrees and roll angles between 80 degrees and 100 degrees or between -100 degrees and -80 degrees as the candidate range for wearing the device sideways; for postures that exceed the above ranges but are still close to the boundaries, they can be defined as transitional or unstable ranges.
[0075] Optionally, to suppress short-term misjudgments, the intelligent device may not immediately confirm the wearing direction after a single frame of posture data meets a threshold. Instead, it can use a continuous judgment mechanism, that is, to count the percentage of data frames that meet a certain posture category condition within a continuous time window of 1 to 3 seconds. When the percentage reaches a preset threshold, such as 70%, 80%, or 90%, the corresponding wearing direction result is output. If the current posture data does not meet any stable judgment conditions, the intelligent device can maintain the currently confirmed wearing direction state and mark the state as pending observation, thereby ensuring the continuous operation of the audio processing link.
[0076] In another implementation, smart devices can construct posture feature vectors based on 3D posture data and call a classification model to identify the wearing direction. This classification model can be a decision tree classifier, support vector machine model, neural network model, or other machine learning models suitable for lightweight deployment. The model input includes the mean pitch angle, mean roll angle, change in yaw angle, variance of angular velocity, and attitude stability indices within a certain time window. The model output includes category labels such as upright wearing, left-side wearing, right-side wearing, and intermediate wrist-flipping state.
[0077] Optionally, to accommodate differences in wrist size, wearing tightness, and habitual posture among users, the smart device can also record the user's posture parameters when performing standard wearing actions upon first use, generating an individualized calibration center value. This center value will then be used as a benchmark for threshold compensation in subsequent judgments. For dynamic scenarios during calls, the smart device can also determine whether the current wearing phase is stable based on the rate of posture change. When the rate of posture change exceeds a preset threshold, it indicates that the wrist is rapidly rotating. In this case, the output of a new wearing direction is delayed until the rate of change decreases and stabilizes before confirming the direction switch, thus avoiding frequent fluctuations in the sound pickup weights during movement.
[0078] Using the above method, the device wearing status is transformed from invisible changes in user behavior into quantifiable posture classification results based on three-dimensional posture data. After adopting duration threshold, stability detection and individualized calibration, it can effectively suppress the wrong direction recognition caused by short-term fluctuations, improve the accuracy and robustness of wearing direction judgment, and thus provide a reliable basis for subsequent weight matrix switching.
[0079] S103, adjusts the weight matrix of the microphone array according to the wearing direction.
[0080] For example, a microphone array can refer to a collection of multiple microphones arranged at different positions on the body of a smart device that can simultaneously collect acoustic signals. Typical structures include dual microphone arrays set on the left and right sides, top and bottom edges or side walls of the device body, or array structures composed of three or more microphones.
[0081] For example, the weight matrix can be a set of parameters describing the proportion of each microphone channel's contribution to the speech enhancement process, such as channel gain coefficients, beamforming coefficients, or multi-channel filtering weights. By adjusting the weight matrix, the proportion of each microphone input signal in the synthesized output can be changed, giving higher weights to microphone channels closer to the user's mouth and with higher speech signal-to-noise ratios, while giving lower weights to channels farther from the mouth and more susceptible to environmental noise or obstruction, thereby achieving targeted speech enhancement and background noise suppression.
[0082] Optionally, after obtaining the current wearing orientation, the smart device can read the corresponding weight matrix from the pre-stored parameter configuration table and send it to the local audio processing unit.
[0083] In one possible implementation, when the device is worn upright, the smart device can load an upright weight matrix to make the microphones closer to the user's mouth have a higher weight than the microphones farther away from the user's mouth; when the device is worn to the side, the smart device can load a side weight matrix to make the microphones closer to the user's mouth have a higher weight than the microphones farther away from the user's mouth.
[0084] For example, the head-mounted weight matrix and the side-mounted weight matrix can be pre-stored in memory and mapped to the wearing direction recognition result, so that the processor can directly call the corresponding matrix after recognizing the current wearing direction. The microphone closer to the user's mouth corresponds to a larger coefficient in the matrix, while the microphone farther away from the user's mouth corresponds to a smaller coefficient, thereby making the target speech account for a higher proportion in the synthesized audio signal.
[0085] In one possible implementation, the smart device has a dual-microphone structure, with two microphones positioned on the left and right sides of the main body of the device, respectively. The processor loads different two-dimensional weight vectors based on whether the device is worn upright or sideways, and performs weighted superposition on the voice signals collected from each channel. When the device is determined to be worn upright, the weight coefficient of the microphone closer to the mouth is set to 0.7 to 0.9, and the other side is set to 0.1 to 0.3. When the device is determined to be worn sideways, the weight allocation relationship is switched accordingly to ensure that the channel closer to the mouth still dominates sound pickup when the direction of the sound source changes.
[0086] Optionally, if the smart device contains three or more microphones, the smart device can generate a multi-channel weight vector based on the wearing direction, the angle between each microphone and the expected mouth direction, and the array geometry, and expand it into a beamforming matrix to control the delay compensation and gain allocation of different channels.
[0087] Optionally, the loading of the weight matrix is performed by the audio control module, which can be implemented by the processor's software program or by a dedicated audio processing unit. The preset matrix parameters in the memory can be stored in a tabular format, facilitating rapid switching when the posture recognition result changes and smoothing the transition of the current call audio during the switching moment to avoid abrupt changes in auditory perception. This implementation, by mapping the wearing direction one-to-one with the weight matrix, enables the device to automatically match different pickup gain distributions between wearing it upright and sideways.
[0088] Optionally, in the audio processing chain, the adjustment of the weight matrix can be performed using a smooth switching method. Specifically, after confirming the new wearing orientation, the processor does not directly perform a step replacement of the microphone weights, but gradually transitions the old matrix to the new matrix within a preset transition time using linear interpolation, exponential smoothing, or piecewise gradual change.
[0089] For example, weight updates can be completed within 100 to 500 milliseconds to avoid abrupt changes in audio output, volume fluctuations, or abnormal spatial perception. When the current posture recognition result is in a transitional range, the system can also maintain the previous stable matrix or calculate an intermediate matrix between two adjacent weight matrices to adapt to the intermediate state between wearing the head and wearing it on the side.
[0090] In some embodiments, the adjustment of the weight matrix can also be further corrected by combining the results of environmental noise detection, voice activity detection, and historical call quality indicators. For example, when the environmental noise comes from the side away from the main microphone, the weight of the channel on that side can be further reduced based on the directional matching matrix; when no voice activity is detected, the weight switching speed can be slowed down, and the gain of the channel near the mouth can be increased after the user speaks, so as to reduce meaningless parameter jitter in the idle state.
[0091] Using the above method, the microphone array can always be aligned with the side more conducive to speech acquisition, improving the effective energy of the target speech, reducing the impact of environmental noise and echo on call quality, thereby improving the consistency of sound pickup under different wearing orientations and reducing the operational burden of manual mode switching for users. Because the weight matrix is dynamically loaded directly according to the wearing orientation, the system can maintain relatively stable speech clarity and call continuity when posture changes.
[0092] Optionally, before dynamically adjusting the weight matrix of the microphone array according to the wearing direction, the smart device can set a duration threshold for determining the wearing direction. If the three-dimensional posture data continuously meets the duration threshold, the corresponding wearing direction is determined.
[0093] For example, the duration threshold can be used to limit the minimum duration for which three-dimensional attitude data must be continuously maintained to meet a preset wearing orientation condition. The preset wearing orientation condition can be defined by at least one attitude parameter, such as pitch angle, roll angle, and yaw angle, or a combination thereof. The duration threshold can be set to a fixed duration, or it can be adaptively corrected based on the current motion intensity, attitude fluctuation amplitude, and historical judgment results to improve the stability of the wearing orientation determination.
[0094] Optionally, the smart device can continuously receive 3D posture data output from the posture sensor and continuously compare the posture data within a preset sampling window. When the posture data falls within the threshold range corresponding to the correct or side-wearing position for a continuous time period, the device outputs a wearing direction determination result consistent with that time period. When the posture data only briefly meets the threshold range but does not reach the duration threshold, the original determination result remains unchanged to avoid erroneous switching caused by wrist shaking, wrist rotation, or momentary occlusion. The duration threshold can be in the range of seconds, such as 2 seconds. The smart device can use a timer and a sliding window statistical method to complete the continuity confirmation.
[0095] By employing the above method, instantaneous posture changes are distinguished from stable wearing states. The switching of the microphone array weight matrix is only triggered after the wearing direction remains consistently stable. This approach ensures that the loading timing of the weight matrix aligns with the actual wearing state, reducing misjudgments and frequent switching caused by brief posture disturbances. This improves the continuity, stability, and speech clarity of call pickup, while also reducing the probability of environmental noise being incorrectly amplified.
[0096] In one possible implementation, before dynamically adjusting the weight matrix of the microphone array according to the wearing direction, the smart device can obtain correlation data between the user's wearing direction and the human voice test signal through a user calibration function. Then, based on the correlation data, the weight matrix is optimized.
[0097] For example, the user calibration function may refer to the smart device guiding the user to maintain a preset wearing state and emit test voice when the device is initialized, the wearing mode is switched, or the user actively initiates calibration, so as to collect the correspondence between the wearing direction and the human voice test signal.
[0098] For example, the human voice test signal can be a calibration statement read aloud by the user, a human voice response to a prompt tone broadcast by the device, or a test voice segment containing specific frequency band characteristics.
[0099] For example, the associated data can be generated by the microphone array synchronously acquiring audio data from each channel during calibration and combining it with the three-dimensional posture information output by the posture sensor. The associated data includes at least the wearing direction indicator, the amplitude energy of each microphone channel for the test speech, the signal-to-noise ratio, and the speech intelligibility index.
[0100] Optionally, after acquiring the associated data, the smart device can use the test voice acquisition quality under the current wearing direction as the optimization target, and solve for the weights of each channel of the microphone array to obtain a personalized weight matrix that matches the user's wearing habits. The solution process can be based on the minimum mean square error criterion, the maximum signal-to-noise ratio criterion, or the weighted energy criterion. By increasing the weights of microphone channels closer to the user's mouth and decreasing the weights of channels farther from the mouth, the calibrated weight matrix can reflect the difference between the user's actual wearing position and the voice propagation path. The optimized weight matrix can be written to the local storage unit and used as a basic parameter for subsequent dynamic adjustments based on the wearing direction, thereby completing individualized sound pickup configuration without changing the physical arrangement of the microphones.
[0101] Optionally, the smart device can obtain the mapping relationship between the wearing direction and the test speech through calibration. Then, it can use this mapping relationship to correct the general weight configuration, so that the subsequent weight switching in the upright or side-wearing state no longer depends solely on a fixed template, but adaptively adjusts based on the user's own wearing deviation and sound pickup characteristics. This enables the microphone array to maintain a relatively stable target speech enhancement capability under different users and different wearing states, and reduces the impact of environmental noise on call clarity.
[0102] The above method allows calibration results to be directly incorporated into the weight matrix design process, improving the consistency between weight allocation and actual wearing direction, and reducing errors caused by general parameters. Since the associated data originates from actual user speech scenarios, the system is more adaptable to differences in wrist size, strap tightness, and wearing habits, thereby improving the stability, clarity, and personalization of call pickup.
[0103] Based on the above analysis, this application provides an adaptive sound pickup method for smart devices, including collecting three-dimensional attitude data of the smart device through an attitude sensor, wherein the attitude sensor includes a gyroscope and an accelerometer, and the three-dimensional attitude data includes pitch angle data, roll angle data and heading angle data; determining the wearing direction of the smart device based on the three-dimensional attitude data; and adjusting the weight matrix of the microphone array according to the wearing direction.
[0104] In this embodiment, by establishing a sequential control relationship between posture acquisition, wearing direction recognition, and array weight adjustment, the device can automatically switch its microphone pickup strategy based on its perceived spatial posture changes, eliminating the need for the user to manually set the call mode. This allows the side of the microphone array closer to the user's mouth to receive a relatively higher weight, thereby improving the target speech energy, suppressing pickup distortion caused by environmental noise and posture mismatch, and reducing call quality fluctuations caused by dynamic wearing changes.
[0105] In one possible implementation, the smart device can further expand the preset threshold range for multi-mode wearing orientation. Then, historical posture data is analyzed using a machine learning model, and the upper and lower limits of the preset threshold range for multi-mode wearing orientation are dynamically optimized based on the analysis results. Here, the machine learning model is a posture analysis model based on a neural network.
[0106] Optionally, a preset threshold range for multi-mode wearing orientation is used to characterize the posture boundary of the smartwatch between upright, side-mounted, tilted, and flipped wearing. The threshold range can correspond to at least one combination of pitch angle, roll angle, and yaw angle. By expanding the preset threshold range, it is possible to prevent the device posture from being prematurely classified as a single wearing orientation when it is in a critical range, thereby reducing microphone weight matrix switching jitter.
[0107] Optionally, the neural network-based posture analysis model can be deployed on the local processing unit of the smart device. The input is historical posture data, which includes a three-dimensional posture sequence continuously collected during calls, sports, and daily wear, and fused by Kalman filtering. The neural network posture analysis model learns the posture change trend over different time periods by using a multilayer perceptron, recurrent neural network, or a structure with a temporal feature extraction layer, and outputs the confidence score and threshold correction amount corresponding to each wearing direction.
[0108] Optionally, the smart device can dynamically adjust the upper and lower limits of the preset threshold range for the multi-mode wearing direction based on the confidence level and correction amount, so that the current posture can be more easily mapped to a mode that matches the actual wearing state.
[0109] Optionally, the smart device can set initial threshold ranges for each wearing mode based on factory parameters, and then adaptively expand or shrink the threshold boundaries by combining the individualized posture distribution formed by the user over a long period of use. When the neural network detects that a user's posture transitions frequently between side-wearing and tilted wearing, the overlapping range of the corresponding interval can be appropriately expanded to improve the continuity of mode switching; when historical data shows that the user's posture is stable over a long period of time, the boundary range can be narrowed to enhance the judgment accuracy. The smart device can also send the threshold update results to the wearing direction determination module, thereby driving the microphone array to load a weight matrix that matches the current wearing mode.
[0110] In this way, smart devices can automatically correct the judgment boundary of multi-mode wearing direction based on historical posture changes, and use the corrected result for microphone array weight allocation. Since the threshold range can be updated according to user posture habits and scene changes, the continuity and adaptability of wearing direction recognition are enhanced, the probability of misjudgment during weight matrix switching is reduced, thereby improving the sound pickup stability and call clarity under different wearing conditions.
[0111] As one possible implementation, after dynamically adjusting the weight matrix of the microphone array according to the wearing direction, the smart device can also analyze the acoustic signals collected by the microphone array through the environmental noise detection module, and adjust the microphone gain in conjunction with the voice activity detection module.
[0112] For example, the ambient noise detection module can be integrated into the audio processing unit of a smartwatch. Its input is the raw acoustic signal obtained by the microphone array during call acquisition. The acoustic signal enters the noise analysis channel after being converted from analog to digital at the front end.
[0113] For example, the environmental noise detection module performs short-time spectrum analysis on the signal within a continuous time window, extracts the noise main frequency distribution, in-band energy and background noise floor value, and generates environmental noise intensity parameters accordingly.
[0114] For example, the voice activity detection module is used to identify whether there is a valid human voice in the current signal. It can output a voice activity flag based on frame energy, spectral entropy, zero-crossing rate or a preset classification model, and output voice confidence information.
[0115] Optionally, when the ambient noise intensity parameter indicates a low-noise environment, the smart device can maintain or slightly increase the gain of the target microphone channel closest to the user's mouth to ensure natural speech. When a medium-to-high noise environment is detected and the speech activity flag is valid, the smart device can increase the target channel gain and simultaneously reduce the background sound pickup gain of non-target channels, thereby increasing the proportion of effective speech in the mixed signal. If the speech activity detection result indicates no significant speech activity, the smart device can reduce the overall microphone gain and maintain noise suppression, thus preventing ambient sound from being amplified. Gain adjustment can be performed by a digital gain controller. The smart device can generate gain coefficients based on noise intensity parameters and speech confidence information and output them to each microphone channel.
[0116] In practical applications, the environmental noise detection module and the voice activity detection module can also be implemented in other equivalent ways, and this application embodiment does not limit them.
[0117] By using the above method to adjust the microphone gain in conjunction with the ambient noise and the state of speech activity, the probability of background noise masking speech can be reduced, the impact of near and far field changes on sound pickup quality can be mitigated, and call clarity and continuity can be improved.
[0118] It should be understood that the above examples are merely illustrative and not limiting. In some embodiments, the posture calculation algorithm, wearing direction determination method, weight matrix structure, and transition control mechanism can all be adjusted according to the device hardware conditions and application requirements. As long as the technical objective of recognizing the wearing direction based on posture data and adjusting the microphone array weight matrix accordingly can be achieved, they can all be included in the implementation scope of the embodiments of this application.
[0119] Figure 2 This is a flowchart illustrating an adaptive sound pickup method for a watch worn sideways / front-facing, as provided in an embodiment of this application. Figure 2 As shown, this method is applied to a smartwatch with at least two microphones. The method includes:
[0120] (1) Collect the attitude data of the smartwatch. This attitude data is acquired in real time by the attitude sensors (such as gyroscope and accelerometer) built into the watch, including at least one of the watch's pitch angle, roll angle, and yaw angle. The gyroscope and accelerometer collect the watch's pitch angle and roll angle in real time at a frequency of 10-50Hz and transmit them to the processor. For example, if the attitude data meets the corresponding wearing direction threshold for 1-3 seconds, it is determined to be the wearing direction to avoid misjudgment due to instantaneous attitude fluctuations.
[0121] (2) Based on the posture data, determine the current wearing direction of the watch. The wearing direction includes wearing it upright and wearing it to the side. The preset upright posture thresholds are: pitch angle of 0°±10° and roll angle of 0°±10°; the side posture thresholds are: pitch angle of 0°±10° and roll angle of 90°±10° or -90°±10°. For example, if the posture data meets the upright wearing threshold for 2 seconds, it is determined to be wearing it upright; if it meets the side wearing threshold for 2 seconds, it is determined to be wearing it to the side.
[0122] (3) Preset the head-mounted pickup weight matrix and the side-mounted pickup weight matrix. The microphone array is a dual-microphone array, which is set on both sides of the watch face. When the head is mounted, the microphone weight on the side closer to the user's mouth is 0.7-0.9, and the microphone weight on the other side is 0.1-0.3. When the side is mounted, adjust the weight distribution so that the microphone weight on the side closer to the user's mouth is 0.7-0.9, and the microphone weight on the other side is 0.1-0.3.
[0123] (4) Weight calibration: When the user wears the watch for the first time, select "front-facing calibration" or "side-facing calibration" in the watch settings interface. The system records the current posture data and the user's voice test signal, optimizes the weight matrix, and makes the sound pickup effect more in line with the user's personal wearing habits.
[0124] In addition, a weight calibration step can be enabled in practical applications: when a user wears the device for the first time, they can manually select the wearing direction, and the system records the posture data and optimal sound pickup weights in that direction. Subsequently, the weight matrix is optimized based on the calibration data to improve sound pickup accuracy.
[0125] Optionally, the smartwatch may include a watch face, a watch band, a microphone array consisting of at least two microphones, a posture sensor, a processor, and a memory; the posture sensor is used to collect the posture data of the watch; the memory stores a head-on pickup weight matrix, a side-on pickup weight matrix, and a computer program; when the processor executes the computer program, it implements an adaptive pickup method for head-on / side-on wearing of the watch.
[0126] Figure 3 This is a schematic diagram of the structure of an adaptive sound pickup device for a smart device provided in an embodiment of this application, as shown below. Figure 3 As shown, the adaptive sound pickup device 300 of the smart device includes:
[0127] The acquisition module 301 is used to acquire three-dimensional attitude data of the smart device through an attitude sensor; the attitude sensor includes a gyroscope and an accelerometer; the three-dimensional attitude data includes pitch angle data, roll angle data and yaw angle data.
[0128] The judgment module 302 is used to determine the wearing direction of the smart device based on the three-dimensional posture data;
[0129] Adjustment module 303 is used to adjust the weight matrix of the microphone array according to the wearing direction.
[0130] Optionally, the acquisition module 301 is also used to perform fusion processing on the raw data of the attitude sensor using a Kalman filter algorithm to output three-dimensional attitude data.
[0131] Optionally, the adjustment module 303 is further configured to load a forward-facing weight matrix when the wearing direction is upright, so that the weight of the microphone closer to the user's mouth is higher than the weight of the microphone farther from the user's mouth; and to load a side-facing weight matrix when the wearing direction is side-facing, so that the weight of the microphone closer to the user's mouth is higher than the weight of the microphone farther from the user's mouth.
[0132] Optionally, the adjustment module 303 is further configured to set a duration threshold for determining the wearing direction before dynamically adjusting the weight matrix of the microphone array according to the wearing direction; and determine the corresponding wearing direction when the three-dimensional posture data continuously meets the duration threshold.
[0133] Optionally, the adjustment module 303 is further configured to set a duration threshold for determining the wearing direction before dynamically adjusting the weight matrix of the microphone array according to the wearing direction; and determine the corresponding wearing direction when the three-dimensional posture data continuously meets the duration threshold.
[0134] Optionally, the adjustment module 303 is further configured to, before dynamically adjusting the weight matrix of the microphone array according to the wearing direction, obtain correlation data between the user's wearing direction and the human voice test signal through a user calibration function; and optimize the weight matrix based on the correlation data.
[0135] Optionally, the adjustment module 303 is also used to expand the preset threshold range of the multi-mode wearing direction; analyze historical posture data through a machine learning model, and dynamically optimize the upper and lower limits of the preset threshold range of the multi-mode wearing direction based on the analysis results; the machine learning model is a posture analysis model based on a neural network.
[0136] The intelligent device adaptive sound pickup device provided in this application embodiment can be used to execute the technical solutions of any of the above embodiments of this application. Its implementation principle and technical effect are similar, and will not be described again here.
[0137] Figure 4 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 4 As shown, the electronic device 400 provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the device 400 further includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.
[0138] In a specific implementation, at least one processor 401 executes computer execution instructions stored in memory 402, causing at least one processor 401 to perform the above-described method.
[0139] Optionally, the memory 402 can be either standalone or integrated with the processor 401.
[0140] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the foregoing embodiments, and will not be repeated here.
[0141] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method of any of the foregoing embodiments.
[0142] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments.
[0143] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.
[0144] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0145] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU) or other general-purpose processors. The processor can also be a Digital Signal Processor (DSP) or an Application Specific Integrated Circuit (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0146] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device, and may also be various media that can store program code, such as USB flash drives, portable hard drives, read-only memory (ROM), disks or optical discs.
[0147] The aforementioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof. Examples of storage media include Static Random Access Memory (SRAM) or Electrically Erasable Programmable Read Only Memory (EEPROM).
[0148] Storage media can be, for example, erasable programmable read-only memory (EPROM) or programmable read-only memory (PROM). Storage media can also be read-only memory (ROM), magnetic storage, flash memory, magnetic disks, or optical disks. Storage media can be any available medium accessible to general-purpose or special-purpose computers.
[0149] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside within an application-specific integrated circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components within an electronic device or host device.
[0150] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0151] The sequence numbers of the embodiments in this application are merely for description and do not represent the superiority or inferiority of the embodiments. Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0152] Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0153] The above are merely preferred embodiments of this application and do not limit the scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the protection scope of this application.
[0154] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0155] It should be further noted that although the steps in the flowchart are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise explicitly stated in this document, there is no strict order requirement for the execution of these steps, and they can be executed in other orders.
[0156] Furthermore, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0157] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0158] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0159] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An adaptive sound pickup method for intelligent devices, characterized in that, The method includes: Three-dimensional attitude data of the intelligent device is collected by an attitude sensor; the attitude sensor includes a gyroscope and an accelerometer; the three-dimensional attitude data includes pitch angle data, roll angle data and yaw angle data. The wearing orientation of the smart device is determined based on the three-dimensional posture data; Adjust the weight matrix of the microphone array according to the wearing direction.
2. The method according to claim 1, characterized in that, The step of dynamically adjusting the weight matrix of the microphone array according to the wearing direction includes: When the wearing direction is upright, a weight matrix for upright wearing is loaded, so that the weight of the microphone closer to the user's mouth is higher than the weight of the microphone farther away from the user's mouth; When the wearing direction is side-wearing, a side-wearing weight matrix is loaded so that the microphone weight closer to the user's mouth is higher than the microphone weight farther away from the user's mouth.
3. The method according to claim 1, characterized in that, Before dynamically adjusting the weight matrix of the microphone array according to the wearing direction, the following steps are included: Set a duration threshold for determining the wearing direction; If the three-dimensional posture data continuously meets the duration threshold, the corresponding wearing direction is determined.
4. The method according to claim 1, characterized in that, Before dynamically adjusting the weight matrix of the microphone array according to the wearing direction, the following steps are included: The correlation data between the user's wearing direction and the human voice test signal is obtained through the user calibration function; Based on the associated data, the weight matrix is optimized.
5. The method according to claim 1, characterized in that, The acquisition of three-dimensional attitude data of the intelligent device through the attitude sensor includes: The Kalman filter algorithm is used to fuse the raw data from the attitude sensor and output three-dimensional attitude data.
6. The method according to claim 1, characterized in that, The step of dynamically adjusting the weight matrix of the microphone array according to the wearing direction further includes: Expand the preset threshold range for multi-mode wearing directions; Historical posture data is analyzed using a machine learning model, and the upper and lower limits of the preset threshold range for the multi-mode wearing direction are dynamically optimized based on the analysis results; the machine learning model is a posture analysis model based on a neural network.
7. The method according to any one of claims 1-6, characterized in that, After dynamically adjusting the weight matrix of the microphone array according to the wearing direction, the method further includes: The acoustic signals collected by the microphone array are analyzed by the environmental noise detection module, and the microphone gain is adjusted in conjunction with the voice activity detection module.
8. An electronic device, characterized in that, include: Memory and processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-7.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.