Audio processing method, apparatus and electronic device
By adjusting the beamforming protection angle of the open-back wireless headphones in real time, the microphone array offset problem caused by unstable wearing posture was solved, achieving stable noise reduction and clear voice output in different wearing states.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2026-06-12
- Publication Date
- 2026-07-31
AI Technical Summary
Open-back wireless headphones suffer from unstable wearing posture, causing the microphone array to shift relative to the user's mouth, affecting directional sound pickup and reducing audio noise reduction performance.
By acquiring the posture information of the earphone wearing status in real time, combined with preset position information and the distance between the two microphones, the protection angle of beamforming is dynamically adjusted, and a steering vector and weight parameters are generated to achieve adaptive beamforming, ensuring that the microphone array is accurately aligned with the user's mouth.
The headset significantly improves the consistency of call noise reduction and voice clarity under different wearing conditions, solves the problem of noise reduction failure caused by wearing misalignment, and improves the overall call quality of the headset.
Smart Images

Figure CN122493874A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of communication technology, and specifically relates to an audio processing method, apparatus and electronic device. Background Technology
[0002] Audio noise reduction technology aims to suppress ambient noise in noisy frequency signals. While filtering out environmental noise interference, it preserves the clarity and naturalness of human voices and is widely used in scenarios such as voice communication and human-computer interaction. Among them, dual-microphone beamforming technology relies on a microphone array to construct a directional sound pickup area facing the user's mouth and other sound-producing directions, focusing on picking up the user's voice and effectively suppressing surrounding environmental noise.
[0003] However, due to the limitations of its product form, the wearing posture of open-back wireless headphones is easily affected by factors such as ear shape, clamping force, and wearing angle, making it impossible to maintain a stable wearing position like over-ear headphones. When the same user wears them multiple times, the relative distance and angle between the microphone array and the user's mouth will randomly shift, causing the directional pickup area with fixed parameters to be unable to accurately lock onto the direction of the sound source. This results in attenuation of the human voice signal, failure of noise suppression function, and reduced audio noise reduction effect. Summary of the Invention
[0004] The purpose of this application is to provide an audio processing method, apparatus, electronic device, storage medium, chip, and computer program product that can improve audio noise reduction performance.
[0005] In a first aspect, embodiments of this application provide an audio processing method applied to an electronic device, the electronic device including a first microphone and a second microphone, the audio processing method including: Based on the preset position information and the posture information of the electronic device in the first wearing state, the target position information is determined; the target position information is the position information of the user's mouth relative to the origin in the first wearing state; the preset position information is the position information of the user's mouth relative to the origin in the second wearing state, and the origin is the midpoint of the line connecting the first microphone and the second microphone. The steering vector is determined based on the target location information and the distance between the first and second microphones; The weighting parameters are determined based on the steering vector and the noise space matrix. The output signal is determined based on the first audio signal collected by the first microphone, the second audio signal collected by the second microphone, and the weighting parameters.
[0006] Secondly, embodiments of this application provide an audio processing apparatus applied to an electronic device, the electronic device including a first microphone and a second microphone, the audio processing apparatus comprising: The determination module is used to determine the target position information based on the preset position information and the posture information of the electronic device in the first wearing state; the target position information is the position information of the user's mouth relative to the origin in the first wearing state; the preset position information is the position information of the user's mouth relative to the origin in the second wearing state, and the origin is the midpoint of the line connecting the first microphone and the second microphone. The determining module is also used to determine the steering vector based on the target location information and the distance between the first microphone and the second microphone; The determination module is also used to determine the weight parameters based on the steering vector and the noise space matrix; The determining module is also used to determine the output signal based on the first audio signal collected by the first microphone, the second audio signal collected by the second microphone, and the weighting parameters.
[0007] Thirdly, embodiments of this application provide an electronic device, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of the audio processing method as described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the audio processing method as described in the first aspect.
[0009] Fifthly, embodiments of this application provide a chip, which includes a processor and a display interface, the display interface and the processor being coupled together, the processor being used to run programs or instructions to implement the steps of the audio processing method as described in the first aspect.
[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the audio processing method as described in the first aspect.
[0011] In this embodiment, target position information can be determined based on preset position information and the posture information of the electronic device in the first wearing state. The target position information is the position information of the user's mouth relative to the origin in the first wearing state; the preset position information is the position information of the user's mouth relative to the origin in the second wearing state, where the origin is the midpoint of the line connecting the first and second microphones. A steering vector is determined based on the target position information and the distance between the first and second microphones. Weight parameters are determined based on the steering vector and the noise space matrix. The output signal is determined based on the first audio signal collected by the first microphone, the second audio signal collected by the second microphone, and the weight parameters. In this way, by using the real-time posture information of the electronic device and the preset position information of the user's mouth in the standard wearing state, the target position information of the user's mouth relative to the origin of the line connecting the two microphones in the current first wearing state can be accurately calculated through spatial rotation transformation. This effectively overcomes the inherent defects of open-back wireless headphones, such as the variable wearing posture and the easy random shift of the relative position between the dual microphone array and the user's mouth. Furthermore, based on the real-time calculated target mouth position and the distance between the two microphones, a guiding vector is dynamically generated. Combined with the noise space matrix characterizing noise statistics, beamforming weight parameters are adaptively solved to achieve real-time adaptive adjustment of the beamforming pickup guard angle, ensuring the main lobe of the beam is always precisely aligned with the real-time human voice source direction. This avoids problems such as voice attenuation and noise suppression failure caused by voice deviation from the preset direction in traditional fixed-beam solutions. It significantly improves the consistency and robustness of call noise reduction under different wearing angles or postures, stably preserving voice clarity and naturalness under various wearing conditions, and improving the overall call quality of electronic devices. Attached Figure Description
[0012] Figure 1 A schematic diagram of the protection angle of the audio processing method provided in some embodiments of this application; Figure 2 Flowcharts of audio processing methods provided for some embodiments of this application; Figure 3 One of the schematic diagrams of an ear clip-on wireless earphone involved in the audio processing method provided for some embodiments of this application; Figure 4 A second schematic diagram of an ear clip-on wireless earphone related to an audio processing method provided in some embodiments of this application; Figure 5(a) is a schematic diagram of the reference coordinate system involved in the audio processing method provided in some embodiments of this application; Figure 5(b) is a schematic diagram of the reference coordinate system involved in the audio processing method provided in some embodiments of this application; Figure 6 A schematic diagram of the structure of an audio processing apparatus provided for some embodiments of this application; Figure 7 A schematic diagram of the structure of an electronic device is provided for some embodiments of this application; Figure 8 This is a schematic diagram of the hardware structure of an electronic device provided for some embodiments of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0014] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0015] Call functionality is a fundamental feature of headsets and other electronic devices. To achieve high-quality calls, headsets need to handle environmental noise cancellation (ENC), wind noise reduction, and echo cancellation through hardware and algorithms. Among these, the mainstream ENC solution is multi-microphone beamforming technology. This technology uses a microphone array to control the directionality of sound pickup, weighting and synthesizing the amplitude and phase of the signals from each microphone to form a beam pointing in a specific direction in space, suppressing interference from other directions. This technology focuses on creating a pickup guard angle pointing towards the user's mouth to pick up only the user's voice and eliminate environmental noise as much as possible.
[0016] The basic logic of beamforming is to achieve directional filtering by utilizing the near-field and / or far-field differences of a sound source. Specifically, speech from the mouth is a near-field sound source, for example, at a distance of less than 10cm, resulting in a large difference in sound pressure level and a fixed time delay when it reaches the main and auxiliary microphones. Environmental noise, on the other hand, is a far-field sound source, with similar sound pressure levels and variable time delays when it reaches both microphones. Thus, by performing time delay compensation and adaptive weighting on the dual-microphone signals, the signals from the mouth direction are superimposed and enhanced in phase, while noise signals from other directions are suppressed through destructive interference, thereby creating a spatial filtering effect with a high-gain main lobe and a low-gain null.
[0017] However, the above effects depend on a stable wearing posture. For ear clip-on, ear hook, and semi-in-ear wireless headphones with poor wearing consistency, the fixed protective angle is difficult to always be aligned with the user's mouth, resulting in a decrease in noise cancellation effect or even voice attenuation. Therefore, an adaptive noise cancellation solution is urgently needed.
[0018] Specifically, such as Figure 1 As shown, Figure 1 To visually demonstrate the protection angle, we'll use a semi-in-ear wireless earphone as an example; the same principle applies to clip-on wireless earphones. Call noise cancellation in semi-in-ear wireless earphones typically employs the following methods: Method one employs a narrow beamforming strategy with fixed parameters. The algorithm presets a fixed pickup protection angle to enhance speech directly facing the mouth and suppress noise from other directions. However, semi-in-ear and clip-on wireless headphones are easily affected by ear shape, clamping force, and wearing angle, such as... Figure 1 As shown, the relative position of the microphone array and the user's mouth will randomly shift each time the same user wears the microphone, causing the user's mouth direction to exceed the range of the preset narrow directional protection angle, such as shifting from the first relative position 101 to the second relative position 102. This means that the narrow directional protection angle indicated after shifting from the standard wearing angle cannot cover the user's mouth direction, resulting in user voice attenuation and deterioration of the sound pickup effect.
[0019] Method two, in order to be compatible with various wearing conditions, adopts the following... Figure 1 The wide directional guard angle indicated by the diagram shows the range. However, an excessively wide guard angle reduces the directivity of beamforming, making it difficult to effectively distinguish between near-field human voices and far-field noise, sacrificing noise reduction performance and resulting in insufficient background noise suppression during calls.
[0020] Method 3, although equipped with dual-microphone hardware, actually uses a single-microphone call solution. It can be seen that this solution does not make full use of the advantages of dual-microphone beamforming, lacks effective spatial filtering capabilities, and has limited suppression of environmental noise, resulting in limited call noise reduction effect.
[0021] It is evident that the call noise reduction solutions for semi-in-ear and clip-on wireless earphones either suffer from voice attenuation due to wearing misalignment, sacrifice noise reduction effect for wearing compatibility, or fail to fully utilize the dual-microphone hardware capabilities, resulting in poor overall noise reduction performance.
[0022] To address the aforementioned technical issues, embodiments of this application provide an audio processing method, apparatus, electronic device, and storage medium. This method analyzes the triaxial acceleration data collected by the headphone's G-Sensor to obtain the headphone's posture information in the current wearing state (first wearing state) in real time. Simultaneously, based on the posture information and the dual-microphone layout, the method determines the relative position of the dual microphones to the user's mouth and then adaptively adjusts the protection angle direction formed by the dual microphone beams to match the current wearing state in real time, thereby achieving consistent call noise reduction performance under different wearing states.
[0023] It should be noted that the terminology used in the implementation section of this application is only used to explain the specific embodiments of this application and is not intended to limit this application.
[0024] The following is in conjunction with the appendix Figures 2 to 8 The audio processing method provided in this application will be described in detail through specific embodiments and application scenarios.
[0025] The audio processing method provided in this application can be executed by electronic devices such as mobile phones, tablets, laptops, PDAs, and wearable devices. Some embodiments of this application use electronic devices as the executing entity to illustrate the audio processing method provided in this application.
[0026] The audio processing method provided in this application can be applied to the following scenarios: In call noise reduction scenarios, such as voice calls, video conferencing, and game voice, the protection angle formed by the dual microphone beams of electronic devices, such as open-back wireless headphones, is adjusted in real time to ensure that the beams are always pointed at the user's mouth, thereby improving the consistency of noise reduction and voice clarity under different wearing conditions. In voice interaction scenarios, such as waking up a smart assistant and inputting commands, dynamic beamforming is used to aim at the user's mouth to improve the signal-to-noise ratio of voice commands and reduce false wake-up and recognition error rates. In recording and live streaming scenarios, when users wear open-back wireless headphones for recording or live streaming, the system adaptively adjusts the pickup direction to reduce environmental noise interference and ensure the purity and naturalness of the recorded sound. Antenna beam adjustment scenarios involve adjusting the beam direction based on the user's posture information when wearing open-back wireless headphones. This can be extended to Bluetooth or Wi-Fi antennas, dynamically adjusting the antenna beam direction based on the real-time posture of the headphones to optimize RF link quality and improve signal strength and connection stability. Multimodal audio processing scenarios combine other sensors on the headphones, such as gyroscopes and accelerometers, to further integrate posture information, which is used to dynamically adjust equalizer parameters or spatial audio rendering direction, thereby improving the listening experience at different wearing angles.
[0027] The open-back wireless headphones in this application include, but are not limited to, ear clip-on wireless headphones, ear hook wireless headphones, ear hook wireless headphones, and behind-the-ear wireless headphones, etc., whose wearing posture is easily affected by factors such as ear shape structure, clamping force, and wearing angle.
[0028] The following is combined Figure 2 This application provides a detailed description of an audio processing method based on an embodiment.
[0029] Figure 2 A flowchart of an audio processing method provided for some embodiments of this application.
[0030] like Figure 2 As shown, the audio processing method provided in this application embodiment can be applied to an electronic device, which includes a first microphone and a second microphone. Based on this, the audio processing method may include steps 210 to 240, as detailed below.
[0031] Step 210: Determine the target position information based on the preset position information and the posture information of the electronic device in the first wearing state; the target position information is the position information of the user's mouth relative to the origin in the first wearing state; the preset position information is the position information of the user's mouth relative to the origin in the second wearing state, where the origin is the midpoint of the line connecting the first microphone and the second microphone; Step 220: Determine the steering vector based on the target position information and the distance between the first microphone and the second microphone; Step 230: Determine the weighting parameters based on the steering vector and the noise space matrix; Step 240: Determine the output signal based on the first audio signal collected by the first microphone, the second audio signal collected by the second microphone, and the weighting parameters.
[0032] In this embodiment of the application, the first wearing state refers to any posture of the user when actually wearing an electronic device, such as an open-back wireless headset, such as the posture of the user looking down to make a call, turning their head to the side, or slightly adjusting the position of the ear clip, which is different from the second wearing state.
[0033] The second wearing state refers to the typical wearing posture of a large group of people collected in the laboratory, that is, the preset standard wearing posture. In this state, the headphone posture is stable and the relative position of the user's mouth and the two microphones is fixed.
[0034] The origin point is the midpoint of the line connecting the first and second microphones on an open-back wireless headset, such as a clip-on wireless headset. In some embodiments, such as... Figure 3As shown, the dual microphones on the clip-on wireless earphones can be respectively mounted on the acoustic component located in front of the ear and the wearing component located behind the ear. The microphone 31 mounted on the acoustic component and close to the user's mouth can be called the "main microphone," and the microphone 32 mounted on the wearing component and far from the user's mouth can be called the "auxiliary microphone." In this way, the dual microphones can form a linear array, and beamforming utilizes the time difference or phase difference and intensity difference of the signals received by the dual microphones to determine the direction of the sound source, i.e., the direction of the user's mouth. In other embodiments, such as... Figure 4 As shown, the dual microphones on the ear clip-on wireless earphone can also be respectively set at both ends of the connecting bridge connecting the acoustic component in front of the ear and the wearing component behind the ear. The microphone 41 set near the user's mouth can be called the "main microphone", and the microphone 42 set away from the user's mouth can be called the "auxiliary microphone". Similarly, the dual microphones can form a linear array. Beamforming uses the time difference or phase difference and intensity difference of the signals received by the dual microphones to determine the direction of the sound source, i.e., the direction of the user's mouth.
[0035] The preset position information is the position of the user's mouth relative to the origin in the second wearing state. This position information can be represented by coordinates, such as the preset coordinates (x0, y0, z0). The preset position information can be obtained by modeling based on the wearing statistics of a large population.
[0036] The target position information is the position of the user's mouth relative to the origin in the first wearing state, and this position information can also be represented by coordinates.
[0037] The steering vector represents the relative phase difference between the signals collected by the two microphones when the sound wave emitted from the user's mouth reaches the dual microphones of the electronic device. It provides a directional reference for subsequent beamforming weight calculation. The spacing between the two microphones is preset to 10mm, which is a fixed hardware parameter of the ear clip wireless earphone and is pre-stored in the earphone memory.
[0038] The noise space matrix is used to characterize the statistical relationship between the environmental noise signals received by the two microphones, such as the power of the environmental noise and the correlation between the two noise signals. It reflects the noise distribution of the current environment in real time. The beamforming weight is used to weight the audio collected by the two microphones to achieve directional enhancement of human voice and suppression of noise.
[0039] Based on this, taking an open-back wireless earphone as an example, where the open-back wireless earphone can be selected as an ear clip-on wireless earphone, steps 210 to 240 will be illustrated.
[0040] First, the clip-on wireless earbuds have a built-in G-sensor, which can be used to collect three-axis acceleration data (ax, ay, az) in real time during the earbuds' current wearing state (first wearing state). By processing the three-axis acceleration data through a preset attitude calculation algorithm, the pitch angle of the earbuds in the current wearing state can be obtained. The pitch angle and roll angle Φ are used to reflect the degree of tilt of the headphones. For example, the pitch angle decreases when the user tilts their head down. The roll angle reflects the degree of lateral rotation of the headphones. For example, the roll angle increases when the user tilts their head to the side. Based on this, the pitch angle and roll angle obtained above are integrated to generate attitude information that can characterize the current spatial posture of the headphones. This attitude information can directly reflect the changes in the spatial position of the headphones caused by wearing offset, ear shape differences, and different clamping forces.
[0041] Next, the standard wearing state is defined as the second wearing state, in which the preset position information (x0, y0, z0) of the user's mouth relative to the origin is stored. When the user's actual wearing of the headphones deviates, for example, due to differences in ear shape causing the ear clip to slip slightly, the headphone pitch angle changes from 15° in the standard wearing state to 8°, and the roll angle changes from 0° in the standard wearing state to 5°. Based on the current posture information, a system is constructed using the pitch angle... The rotation matrix R (composed of the roll angle Φ and the roll angle Φ) At this point, the preset position information of the second wearing state, such as (x0, y0, z0), can be substituted into the rotation matrix to complete the spatial rotation transformation, correct the position deviation caused by the wearing posture, and finally obtain the actual position of the user's mouth relative to the origin in the first wearing state, that is, the target position information (x1, y1, z1), so as to achieve accurate positioning of the actual position of the user's mouth.
[0042] Furthermore, based on the target position information (x1, y1, z1) obtained above, a position vector P=[x1, y1, z1] of the user's mouth relative to the origin is constructed. The target orientation angle of the user's mouth relative to the dual-microphone array is then calculated. This target orientation angle can include the horizontal azimuth angle and the vertical pitch angle. Combining the fixed microphone spacing d=10mm of the clip-on wireless earphone and the propagation characteristics of sound waves in air, the propagation delay difference of the user's voice reaching the first and second microphones can be calculated. Then, based on the delay difference and the sound wave frequency, the relative phase difference between the signals collected by the two microphones is determined. Finally, based on this relative phase difference, a steering vector a is generated. This steering vector accurately matches the actual incident direction of the human voice in the current wearing state, providing an accurate directional reference for adaptive beamforming.
[0043] Finally, with the optimization objectives of achieving distortion-free human voice signal in the target direction and minimizing output noise power, the real-time generated steering vector and the updated noise space matrix are jointly input into the adaptive beamforming algorithm to obtain the corresponding weight parameters, namely the beamforming weights for the first and second microphones. Based on these weight parameters, the first and second audio signals are weighted and fused to output the noise-reduced target audio signal. This achieves adaptive enhancement of the human voice signal in the target direction and adaptive suppression of ambient noise in the non-target direction, effectively solving the problems of human voice attenuation, noise reduction failure, and inconsistent call quality caused by varying wearing postures and relative sound source position shifts in clip-on wireless headphones. It significantly improves the stability and clarity of call noise reduction under various wearing conditions.
[0044] Therefore, by using real-time posture information of the electronic device and preset position information of the user's mouth in the standard wearing state, the target position information of the user's mouth relative to the origin of the midpoint of the dual-microphone connection in the current first wearing state is accurately calculated through spatial rotation transformation. This effectively overcomes the inherent defects of open-back wireless headphones, such as the variable wearing posture and the random displacement of the relative orientation of the dual-microphone array and the user's mouth. Furthermore, a guiding vector is dynamically generated based on the real-time calculated target position of the mouth and the distance between the dual microphones. Combined with the noise space matrix representing the statistical characteristics of noise, the beamforming weight parameters are adaptively solved to achieve real-time adaptive adjustment of the beamforming pickup protection angle, ensuring that the main lobe of the beam is always accurately aligned with the direction of the real-time human voice source. This avoids the problems of human voice attenuation and noise suppression failure caused by the deviation of the human voice from the preset direction in traditional fixed beam solutions. It significantly improves the consistency and robustness of call noise reduction effect under different wearing angles or postures, stably preserves the clarity and naturalness of human voice under various wearing conditions, and improves the overall call quality of the electronic device.
[0045] The steps described above are explained in detail below.
[0046] Regarding step 210, in some embodiments of this application, in order to improve the accuracy of obtaining real-time attitude information of electronic devices, avoid attitude data estimation deviation, and ensure the accuracy of subsequent user mouth position calculation, attitude information needs to be determined before step 210. Based on this, the audio processing method may also include steps 2501 and 2502.
[0047] First, the reference coordinate system of the headphones in this embodiment is defined as shown in Figure 5(a). The reference coordinate system is defined with the midpoint of the line connecting the first microphone and the second microphone as the origin, the direction from the first microphone to the second microphone as the first direction (X-axis, denoted as Ax), the side perpendicular to the first direction and facing the user's face when the headphones are worn as the second direction (Y-axis, denoted as Ay), the first direction and the second direction together spanning a first plane, and the direction perpendicular to both the first and second directions as the third direction (Z-axis, denoted as Az), forming a right-handed rectangular coordinate system. This reference coordinate system is used to uniformly describe the spatial attitude, pitch angle, roll angle of the headphones, and the three-dimensional spatial position of the user's mouth relative to the headphones.
[0048] Based on the reference coordinate system described in the embodiments of this application, attitude information is defined, including pitch angle and roll angle. Pitch angle refers to the angle by which the earphone rotates around the X-axis of the reference coordinate system in the electronic device, such as an ear-clip wireless earphone, reflecting the tilt state of the earphone in the front-to-back direction. When the front end of the earphone (closest to the face) is downward and the rear end is upward, the pitch angle is negative; conversely, the pitch angle is positive. Roll angle refers to the angle by which the earphone rotates around the Y-axis in the reference coordinate system, reflecting the left-to-right offset state of the earphone. When the earphone tilts to the right and the right side sinks, the roll angle is positive; when the earphone tilts to the left and the left side sinks, the roll angle is negative.
[0049] Based on this, in step 2501, the three-axis acceleration data of the electronic device in the first wearing state are obtained.
[0050] In this step, the ear-clip wireless earbuds worn on the user's ears use a built-in G-sensor to collect real-time triaxial raw acceleration data of the earbuds in the first wearing state, namely X-axis acceleration ax, Y-axis acceleration ay, and Z-axis acceleration az. When the earbuds are stationary or undergoing slow wearing and fine-tuning, the triaxial acceleration data can approximately represent the components of gravity on each coordinate axis.
[0051] Step 2502: Based on the triaxial acceleration data, determine the pitch and roll angles of the electronic device in the first wearing state.
[0052] In this step, to eliminate high-frequency interference caused by user head movements and limb movements, this embodiment of the application can first perform low-pass filtering preprocessing on the collected triaxial acceleration data to obtain smooth and stable effective acceleration components. Then, based on the filtered triaxial acceleration data, combined with the gravitational acceleration g=9.8m / s², 2 Calculate the pitch angle of the clip-on wireless earphones in the current wearing posture. With roll angle Φ.
[0053] Based on this, using the coordinate system shown in Figure 5(a), the G-sensor mounted on the earcup wireless earphone can collect the triaxial acceleration (ax, ay, az) of the earphone in the first wearing state in real time. When stationary or in slow motion, it can be approximately considered to reflect the direction of gravity. Therefore, the pitch angle of the earcup wireless earphone... The forward and backward tilt can be approximately calculated using the following formula (1), and the roll angle Φ of the ear clip wireless headphones, i.e., the left and right deflection, can be approximately calculated using the following formula (2):
[0054] Where, g = 9.8 m / s 2 In practical applications, the acceleration value after low-pass filtering can be used to eliminate motion interference.
[0055] Therefore, by relying on the raw data from gravity sensing, attitude quantification is completed, enabling precise representation of the headphone's spatial tilt and deflection states, and providing accurate and quantifiable basic attitude parameters for subsequent spatial rotation transformations and sound source location estimation.
[0056] It should be noted that, in this embodiment of the application, step 210 can be triggered in two modes, as shown below.
[0057] Mode 1: Real-time adaptive detection mode, which continuously triggers step 210 at preset time intervals during the call. Each time it is triggered, step 210 is executed in real time, and the beamforming protection angle is dynamically adjusted in combination with the dual microphone layout and preset parameters of the user's mouth position.
[0058] This allows for continuous tracking of changes in wearing posture, such as shifts caused by movement or unintentional touches, and dynamic calibration of the protection angle to ensure that the main lobe of the beam is always aligned with the user's mouth. This results in stable and consistent call noise reduction performance under different wearing conditions, improving robustness.
[0059] Mode 2: Switch mode. This mode requires the user to actively click a preset physical or virtual switch after wearing the ear-clip wireless headset and starting a call, thus triggering step 210 once. After triggering, a wearing status detection and a one-time adaptive beamforming adjustment are performed, and this adjustment result is maintained until the call ends or the switch is manually triggered again.
[0060] This reduces system power consumption and computational overhead, meeting users' need for simple calibration and continuous use, while avoiding minor audio delays that may result from frequent testing, making it suitable for usage scenarios with relatively fixed wearing postures.
[0061] In other words, regardless of the mode used, the wearing status is detected by the G-Sensor and the beamforming is adaptively adjusted by combining geometric parameters. This effectively overcomes the noise cancellation failure caused by differences in the wearing of open-back headphones, and significantly improves the clarity of call voice and the ability to suppress environmental noise.
[0062] In step 220, in this embodiment of the application, the dual-microphone orientation preset parameters can be modeled based on statistical data of a large number of people wearing the microphones. For example, the distance d between the first and second microphones, and the position coordinates (x, y) of the first microphone in the dual microphones. m1 ,y m1 ,z m1 The position coordinates (x, y) of the second microphone in a dual-microphone setup. m2 ,y m2 ,z m2 Furthermore, based on statistical data from a large population of users, model the positional information of the user's mouth, such as the user's mouth coordinates (x...). mouth ,y mouth ,z mouth Based on the above two points, the position information of the user's mouth relative to the origin in the standard second wearing state can be obtained, that is, the preset position information. This preset position information can be represented by a vector: P0=[x0,y0,z0]T, which remains fixed.
[0063] Based on this, when the headphones are rotated to the position indicated by the current posture information, the position information P of the user's mouth relative to the origin in the first wearing state can be obtained by applying a rotation to the preset position information P0 using the following formula (3):
[0064] Among them, R( Φ is a rotation matrix consisting of pitch and roll angles, which rotates a vector about a spatial axis.
[0065] In some embodiments of this application, step 230 is used to convert the actual position coordinates of the user's mouth into azimuth parameters that can be used in beamforming algorithms, thereby achieving an effective conversion from spatial position to acoustic pointing. Based on this, step 230 may specifically include steps 2301 and 2302.
[0066] Step 2301: Based on the target position information, determine the target direction angle of the user's mouth relative to the origin in the first wearing state.
[0067] To clarify the spatial definition and coordinate system benchmark of the target orientation angle and avoid ambiguity in angle determination and orientation calculation, the orientation angle is specifically refined. As shown in Figure 5(b), the target orientation angle includes the horizontal azimuth angle and the vertical pitch angle. In this embodiment, the midpoint of the line connecting the two microphones is taken as the origin, the X-axis is laid out along the line connecting the two microphones, the first direction is the positive direction of the X-axis, the Y-axis is perpendicular to the X-axis, and the second direction is the positive direction of the Y-axis. The first direction and the second direction constitute the first plane. Based on this, the horizontal azimuth angle is the angle between the projection of the user's mouth position vector relative to the origin on the first plane and the first direction in the first wearing state. The vertical pitch angle is the angle between the above position vector and the first plane. In this way, the spatial angle calculation benchmark is standardized, the orientation determination error under different wearing postures is eliminated, the three-dimensional spatial orientation of the user's mouth is refined, and the accuracy of adaptive beam adjustment is further improved.
[0068] In this embodiment of the application, the target position information of the user's mouth relative to the origin in the first wearing state can be represented by a three-dimensional position vector as P=[P x P y P z Based on this position vector, the target orientation angle of the user's mouth relative to the dual microphone array can be calculated, where the target orientation angle includes the horizontal azimuth angle α and the vertical pitch angle β.
[0069] Wherein, the horizontal azimuth angle α is the deflection angle of the projection of the position vector in the first plane relative to the positive Y-axis direction of the headphones, which can be calculated by the following formula (4). The vertical pitch angle β is the pitch angle between the position vector and the first plane, i.e., the horizontal plane, which can be calculated by the following formula (5):
[0070] Here, P x P y P z These are the three components of the target location information P, where ||P|| is the position coordinate (x, y) based on the first microphone. m1 ,y m1 ,z m1 ), the position coordinates of the second microphone (x) m2 ,y m2 ,z m2 ) and the user's mouth coordinates (x mouth ,y mouth ,z mouth The distance from the user's mouth to the origin is calculated based on a preset value. In short, since the user's mouth is typically located below and in front of the earpiece, P... y Always positive, P z Since the ratio is negative, α and β can be directly calculated from the ratio.
[0071] Thus, by using the above formulas (4) and (5), the horizontal azimuth and vertical pitch angles corresponding to the user's mouth in any wearing posture can be accurately calculated, realizing the quantitative solution of the three-dimensional spatial sound source orientation, and providing an accurate angle parameter basis for the subsequent accurate calculation of the dual microphone delay difference, phase difference and steering vector.
[0072] Step 2302: Determine the guidance vector based on the target orientation angle and spacing.
[0073] This allows for the establishment of a correlation between spatial location and sound source directionality, unifying the dimensions of input parameters for beamforming, improving the logic and accuracy of steering vector solving, and ensuring the matching degree of sound pickup directionality adjustment.
[0074] Furthermore, in order to clarify the spatial definition and coordinate system reference of the target direction angle and avoid ambiguity in angle determination and direction calculation, the direction angle is further refined and limited. Based on this, the above step 2302 may specifically include steps 23021 to 23023.
[0075] Step 23021: Determine the time delay difference between the sound waves emitted from the user's mouth and the first and second microphones based on the target direction angle, spacing, and sound speed.
[0076] In this step, for a sound wave with frequency f, the time delay difference between the sound wave emitted from the user's mouth and the first and second microphones can be calculated using the following formula (6):
[0077] Where d is the distance between the first microphone and the second microphone, and c is the speed of sound.
[0078] Step 23022: Based on the time delay difference and the frequency of the sound wave, determine the relative phase difference between the signals collected by the first microphone and the second microphone when the sound wave emitted from the user's mouth reaches the first microphone and the second microphone.
[0079] In this step, when the sound wave emitted by the user's mouth reaches the first microphone and the second microphone, the relative phase difference 'a' between the signals collected by the first microphone and the second microphone can be calculated using the following formula (7):
[0080] It should be noted that the above-mentioned headphones are all described using a single headphone as an example.
[0081] Step 23023: Based on the relative phase difference, generate a steering vector, denoted as a. Here, a represents the direction of the human voice that needs to be preserved and enhanced without distortion.
[0082] Therefore, by quantifying the phase difference of dual-microphone signals based on the actual acoustic propagation law, the guide vector is made to fit the actual sound field transmission characteristics, which greatly improves the directivity of directional sound pickup and the rationality of algorithm calculation.
[0083] In step 240, in some embodiments of this application, an optimal solution constraint criterion is added to limit the calculation logic of the weight parameters, so as to balance the requirements of human voice fidelity and environmental noise suppression, and avoid the problem of excessive noise reduction leading to human voice distortion or insufficient noise reduction leading to residual noise. Based on the steering vector and the noise space matrix, the optimal weight parameters are obtained by the adaptive beamforming algorithm. Based on this, step 240 can specifically adopt the minimum variance distortionless response (MVDR) criterion, and calculate the weight parameters by formula (8):
[0084] Where 'a' is the steering vector, Rnn is the noise space matrix, the superscript -1 indicates matrix inversion, and the superscript H indicates conjugate transpose. The weighting parameter W is obtained by solving the algorithm. W is the final parameter required by the adaptive beamforming algorithm. Applying this weighting parameter to the audio signals from the two microphones in real time outputs the enhanced speech signal and completes transmission, while simultaneously updating the pickup guard angle in real time. The noise space matrix Rnn is estimated in real time from the signals collected by the microphones. Based on the statistical relationships such as autocorrelation and cross-correlation of the noise signals received by the two microphones, this matrix can objectively characterize the environmental noise distribution characteristics in the current scene.
[0085] It should be noted that, in this embodiment of the application, in addition to calculating the weight parameters based on the above formula (8), the calculation can also be equivalently implemented in actual engineering by adaptive iterative algorithms such as generalized sidelobe cancellers, thus avoiding the computational power consumption caused by direct matrix inversion. In addition, in order to further reduce the overall computational complexity, several weight vectors in preset directions can be pre-calculated and interpolated in combination with the real-time posture information of electronic devices such as headphones (wearable devices) in the first wearing state to quickly obtain the weight parameters under the current working condition.
[0086] Therefore, while ensuring the user's voice is output completely without distortion, it maximizes the suppression of environmental noise, balances call clarity and noise reduction performance, and improves the listening experience in different scenarios.
[0087] In addition, in some embodiments, in order to achieve dynamic iterative updates of the noise space matrix, distinguish between speech frames and silence frames, and avoid human voices interfering with noise statistics results, the matrix update logic is optimized. Based on this, before step 240, the audio processing method may also include step 2601 or step 2602.
[0088] Step 2601: When the audio signal includes the user's voice signal, the historical noise space matrix is determined as the noise space matrix. The audio signal includes at least one of the following: a first audio signal and a second audio signal. The noise space matrix provided in the embodiments of this application will be described in detail below.
[0089] When the audio signal composed of the first and second audio signals includes the user's voice signal, the historical noise space matrix is directly used as the noise space matrix for the current calculation. The noise space matrix Rnn can be calculated as follows. Here, the noise space matrix Rnn is a complex matrix describing the statistical relationship between the noise signals received by the first and second microphones, and Rnn can be represented by a complex matrix:
[0090] Among them, the diagonal elements R11 and R22 represent the autocorrelation of the noise signals received by the first microphone and the second microphone, respectively, which characterize the power of the single noise signal itself. The off-diagonal elements R12 and R21 represent the cross-correlation of the two noise signals, respectively, which are used to reflect the similarity and phase correlation characteristics of the two noise signals.
[0091] Rnn reflects the spatial distribution characteristics of noise. Specifically, when noise enters a microphone array from different directions, it will produce different time differences, phase differences, and amplitude differences on the two microphones. Rnn estimates the spatial distribution of noise by statistically analyzing these differences. For example, if the noise comes from directly in front, the noise waveforms received by the two microphones are almost identical, and the amplitude of R12 is close to 1. If the noise comes from the side, the noise received by the two microphones differs significantly, i.e., one arrives first and the other later, or it may even be blocked by a head, and the amplitude of R12 is smaller. If the noise is uniform from all directions, such as wind noise or reverberation, the noise from the two microphones is uncorrelated, and the amplitude of R12 is close to 0.
[0092] Step 2602: If the audio signal does not include the user's speech signal, determine the noise space matrix based on the audio signal and the historical noise space matrix. If the audio does not include the user's speech signal, the historical noise space matrix can be recursively smoothed based on at least one of the current first audio signal and the second audio signal to obtain the noise space matrix, which is then used as the current noise space matrix for determining the weight parameters.
[0093] The triggering conditions for steps 2601 and 2602 above will be explained below.
[0094] In practical headphone use, it is impossible to directly measure pure noise signals because the microphone always mixes speech and noise. Therefore, it is necessary to calculate Rnn during periods of speech silence, i.e., when the audio collected by the dual microphones does not include the user's speech, i.e., when the user is not speaking. Based on this, the historical noise space matrix can be calculated using the following formula (9). Perform recursive smoothing to obtain the current noise space matrix. :
[0095] in, The sound wave frequency point, This represents the current microphone signal spectrum. α is an update coefficient that can take any value from 0.8 to 0.9, controlling the update rate.
[0096] For example, since the noise characteristics may differ at different frequencies, the Rnn estimation is performed frequency-by-frequency in the frequency domain. The processing flow is as follows: Perform Fast Fourier Transform (FFT) on each frame of audio from the first microphone and the second microphone respectively to obtain the first frequency domain signal X1(ω,t) of the first audio signal collected by the first microphone and the second frequency domain signal X2(ω,t) of the second audio signal collected by the second microphone. Run VAD to detect whether the audio signal includes a speech signal. If it does not include a speech signal, i.e., VAD=0, calculate the current noise space matrix for each frequency point ω based on the above formula (9), as shown below:
[0097]
[0098]
[0099]
[0100] Furthermore, in this step, it can be determined whether the audio signals collected by the first and second microphones include the user's voice signal through the following process. Specifically, VAD can be used to determine whether the current audio signal contains only noise. The VAD method can be at least one of the following: energy thresholding method or multi-feature fusion. If the energy thresholding method is used, when the voice signal energy is lower than a preset threshold of the background noise energy, it is determined that the audio signals collected by the first and second microphones do not include the user's voice signal, i.e., they are silent. Multi-feature fusion refers to combining energy, spectral flatness, periodicity, and bone conduction pickup device (VPU) assisted detection, etc.
[0101] This avoids human voice signals from polluting noise statistics, ensures that the noise space matrix matches the current environmental noise distribution in real time, and improves the stability and robustness of noise suppression in complex environments.
[0102] In some embodiments of this application, in order to clarify the frequency domain processing flow of dual-microphone audio signals, improve the actual implementation scheme of beam weighting, and complete the signal noise reduction output link, this post-processing step is added. Based on this, the above step 240 may specifically include steps 2401 to 2403.
[0103] Step 2401: Perform Fourier transform processing on the first audio signal and the second audio signal respectively to obtain the first frequency domain signal of the first audio signal and the second frequency domain signal of the second audio signal.
[0104] In this step, the first and second audio signals are preprocessed, specifically including framing and windowing operations. A single frame duration is set to 20ms, the frame overlap rate to 50%, and a Hanning window of the same size is used for windowing. Preprocessing filters out high-frequency interference and noise from both signals, ensuring overall signal stability. Subsequently, Fourier transforms are performed on the preprocessed first and second audio signals to convert the time-domain signals into frequency-domain signals. The first frequency-domain signal, denoted as X1(k,m), is obtained after the transform. Similarly, the same operation is performed on the second audio signal to obtain the second frequency-domain signal, denoted as X2(k,m). Here, k is the frequency index, taking any number from 1 to 1024, and m is the frame index, taking numbers from 1 to N, covering a frequency range of 20Hz to 20kHz. Based on this, the first and second frequency-domain signals are obtained, providing a data basis for subsequent frequency-domain weighted summation.
[0105] Step 2402: Using weight parameters, the first frequency domain signal and the second frequency domain signal are weighted and summed to obtain the fused frequency domain signal.
[0106] In this step, the first frequency domain signal and the second frequency domain signal can be weighted and summed according to formula (10) to obtain the fused frequency domain signal Y(k,m):
[0107] Where w is the weighting parameter, and * denotes complex conjugation.
[0108] Step 2403: Perform an inverse Fourier transform on the fused frequency domain signal to obtain the output signal.
[0109] In this step, the fused frequency domain signal Y(k,m) undergoes a frame-by-frame inverse transform using a fast inverse Fourier transform to convert the frequency domain complex signal into a time domain analog signal. Each frame's inverse transform duration is ≤10ms to ensure real-time performance. The inverse-transformed time domain signals are then superimposed (frame overlap rate 50%) to form a continuous time domain signal. For example, time domain signals from m=1 to 5 frames are superimposed to eliminate inter-frame gaps and obtain a smooth audio signal. Next, the spliced time domain signal undergoes amplitude calibration, for example, to 0.5 to 0.8Vpp, to remove residual noise interference, resulting in the final noise-reduced output signal. This signal can be directly transmitted to the other end of the call, ensuring that the other party can clearly hear noise-free speech during the call. For example, after processing, the output signal-to-noise ratio increases from 15dB to 35dB, effectively suppressing environmental noise. Through the above steps, the adaptive beamforming algorithm adjusts the coefficients according to the direction weights, thereby adaptively adjusting the beamforming direction based on the real-time wearing status of the headphones. This allows it to effectively pick up the user's lip speech and suppress environmental noise in different wearing states.
[0110] This enables precise weighting and noise cancellation in the frequency domain, adapts to the operational characteristics of beamforming algorithms, and completes the entire process from raw audio reception to noise reduction output, ensuring the feasibility of the noise reduction algorithm.
[0111] In addition, to further improve the stability, robustness and personalized adaptability of call noise reduction for open-back wireless earphones, the embodiments of this application can perform closed-loop verification and learning optimization functions, as detailed below.
[0112] In some embodiments of this application, call quality feedback can monitor relevant indicators of audio signals collected by the first and second microphones in real time, with a focus on monitoring the signal-to-noise ratio and speech clarity index, and provide real-time feedback on the current call quality status, providing data support for subsequent learning optimization and anomaly recovery.
[0113] In some embodiments of this application, adaptive learning can be performed to adapt to the wearing habits of different users. Specifically, when it is detected that the signal-to-noise ratio is continuously at a high level in a certain wearing state, such as SNR≥35dB for more than 5 seconds, the angle thresholds corresponding to the wearing state, such as pitch angle and roll angle reference thresholds, are automatically fine-tuned to achieve personalized optimization that fits the user's personal wearing habits and further improve the call noise reduction effect in the wearing state.
[0114] In some embodiments of this application, when the G-sensor detects an abnormal wearing status of the headset, such as a roll angle absolute value > 20° and lasting for more than 300ms, a pitch angle < 10° and low-frequency small fluctuations in Z-axis acceleration, and simultaneously detects a decrease in call quality such as SNR < 15dB and a speech clarity index below a threshold, a dual abnormal recovery mechanism can be activated: first, the headset prompts the user to fine-tune the wearing posture via voice prompts; second, it automatically switches to omnidirectional pickup mode as a fault-tolerant solution to avoid call interruption or noise reduction failure and ensure call continuity.
[0115] That is, after step 240, the audio processing method may also include steps 2701 and 2702.
[0116] Step 2701: Determine whether the electronic device is in an abnormal wearing state based on the gravity sensor data of the electronic device. The abnormal wearing state includes at least one of the following: deflection state, slippage state, or loose state.
[0117] Specifically, when the absolute value of the roll angle is detected to be greater than 20° and the duration exceeds 300ms, or when the pitch angle is less than 10° and the Z-axis acceleration shows low-frequency small fluctuations, it is judged as an abnormal wearing state.
[0118] Step 2702: When an abnormal wearing condition is determined, the electronic device outputs a voice prompt message to prompt the user to fine-tune the wearing posture.
[0119] Therefore, by monitoring the wearing status of the headphones in real time, such as excessive roll angle, insufficient pitch angle, and low-frequency fluctuations in the Z-axis, as well as call quality such as signal-to-noise ratio and speech intelligibility index, voice prompts are triggered when both are abnormal, guiding users to fine-tune their wearing posture. This can promptly and accurately remind users to correct wearing misalignment, avoiding noise cancellation failure or speech attenuation due to poor wearing posture. It also ensures that the adaptive beamforming algorithm restores the optimal sound pickup direction, significantly improving the consistency and robustness of call noise cancellation under different wearing conditions, while improving the user experience and reducing call quality complaints caused by wearing problems.
[0120] In summary, the audio processing method provided in this application can accurately locate the actual position of the user's mouth by acquiring the headphone's posture information in real time through a G-sensor and combining it with preset geometric parameters. By dynamically adjusting the weight parameters, it can achieve directional enhancement of human voice and effective suppression of environmental noise. It has the following advantages: First, the G-sensor operates with low power consumption, without affecting the headphone's battery life; second, the beam direction can adaptively adjust with the wearing posture, avoiding noise reduction failure caused by wearing misalignment; third, it adapts to different ear shapes and wearing angles, significantly improving the signal-to-noise ratio for calls and ensuring clear, distortion-free human voice; fourth, the algorithm deployment is flexible, and parameters can be adjusted according to hardware computing power, balancing practicality and stability, effectively solving the technical pain point of unstable noise reduction performance in open-back headphones.
[0121] The audio processing method provided in this application can be executed by an audio processing device. This application uses an audio processing device executing a display as an example to illustrate the apparatus of the audio processing method provided in this application.
[0122] This application also provides an audio processing apparatus. (Specifically combined with...) Figure 6 Please provide a detailed explanation.
[0123] Figure 6 This is a schematic diagram of the structure of an audio processing device provided for some embodiments of this application.
[0124] like Figure 6 As shown, the audio processing device 60 can be applied to electronic devices, and the audio processing device 60 may specifically include: The acquisition module 601 is used to determine the target position information based on the preset position information and the posture information of the electronic device in the first wearing state; the target position information is the position information of the user's mouth relative to the origin in the first wearing state; the preset position information is the position information of the user's mouth relative to the origin in the second wearing state, and the origin is the midpoint of the line connecting the first microphone and the second microphone. The determining module 602 is further configured to determine the steering vector based on the target location information and the distance between the first microphone and the second microphone; The determining module 602 is also used to determine the weight parameters based on the steering vector and the noise space matrix; The determining module 602 is further configured to determine the output signal based on the first audio signal collected by the first microphone, the second audio signal collected by the second microphone, and the weighting parameters.
[0125] The audio processing device 60 in the embodiments of this application will be described in detail below.
[0126] In some embodiments of this application, the acquisition module 601 can also be used to acquire triaxial acceleration data of the electronic device in a first wearing state when the attitude information includes pitch angle and roll angle. The determination module 602 is also used to determine the pitch and roll angles of the electronic device in the first wearing state based on triaxial acceleration data.
[0127] In some embodiments of this application, the determining module 602 may be specifically used to determine, based on the target position information, the target orientation angle of the user's mouth relative to the origin in the first wearing state; Determine the guidance vector based on the target orientation angle and spacing.
[0128] In some embodiments of this application, the determining module 602 may specifically be used to determine the time delay difference between the sound waves emitted from the user's mouth and the first and second microphones, based on the target direction angle, spacing and sound speed. Based on the time delay difference and the frequency of the sound wave, determine the relative phase difference between the signals collected by the first microphone and the second microphone when the sound wave emitted from the user's mouth reaches the first microphone and the second microphone; A steering vector is generated based on the relative phase difference.
[0129] In some embodiments of this application, the determining module 602 may be specifically used to determine the historical noise space matrix as a noise space matrix when the audio signal includes a user voice signal, wherein the audio signal includes at least one of the following: a first audio signal and a second audio signal; In cases where the audio signal does not include the user's voice signal, the noise space matrix is determined based on the audio signal and the historical noise space matrix.
[0130] In some embodiments of this application, the determining module 602 may be specifically used to perform Fourier transform processing on the first audio signal and the second audio signal respectively to obtain the first frequency domain signal of the first audio signal and the second frequency domain signal of the second audio signal. By using weight parameters, the first frequency domain signal and the second frequency domain signal are weighted and summed to obtain the fused frequency domain signal; The inverse Fourier transform of the fused frequency domain signal is performed to obtain the output signal.
[0131] The audio processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.
[0132] The audio processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0133] The device coordination apparatus provided in this application embodiment can achieve... Figure 2 The various processes implemented in the audio processing method embodiment shown in Figure 5 achieve the same technical effect, and will not be described again here to avoid repetition.
[0134] Based on this, the audio processing device provided in this application embodiment can determine the target position information according to preset position information and the posture information of the electronic device in the first wearing state; the target position information is the position information of the user's mouth relative to the origin in the first wearing state; the preset position information is the position information of the user's mouth relative to the origin in the second wearing state, where the origin is the midpoint of the line connecting the first microphone and the second microphone; a steering vector is determined based on the target position information and the distance between the first microphone and the second microphone; a weighting parameter is determined based on the steering vector and the noise space matrix; and an output signal is determined based on the first audio signal collected by the first microphone, the second audio signal collected by the second microphone, and the weighting parameter. In this way, by using the real-time posture information of the electronic device and combining it with the preset position information of the user's mouth in the standard wearing state, the target position information of the user's mouth relative to the origin of the midpoint of the line connecting the two microphones in the current first wearing state is accurately calculated through spatial rotation transformation, effectively overcoming the inherent defects of open wireless headphones where the wearing posture is variable and the relative orientation of the dual microphone array and the user's mouth is prone to random shift. Furthermore, based on the real-time calculated target mouth position and the distance between the two microphones, a guiding vector is dynamically generated. Combined with the noise space matrix characterizing noise statistics, the beamforming weight parameters are adaptively solved, enabling real-time adaptive adjustment of the beamforming pickup guard angle. This ensures the main lobe of the beam is always precisely aligned with the direction of the real-time human voice source. This avoids the problems of voice attenuation and noise suppression failure caused by the deviation of the human voice from the preset direction in traditional fixed-beam solutions. It significantly improves the consistency and robustness of call noise reduction under different wearing angles or postures, stably preserving the clarity and naturalness of the human voice under various wearing conditions, thus improving the overall call quality of electronic devices.
[0135] Optional, such as Figure 7 As shown, this application embodiment also provides an electronic device 70, including a processor 701 and a memory 702. The memory 702 stores a program or instructions that can run on the processor 701. When the program or instructions are executed by the processor 701, they implement the various steps of the above-described audio processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0136] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0137] Figure 8 This is a schematic diagram of the hardware structure of an electronic device provided for some embodiments of this application.
[0138] The electronic device 800 includes, but is not limited to, components such as: radio frequency unit 801, network module 802, audio output unit 803, input unit 804, sensor 805, display unit 806, user input unit 807, interface unit 808, memory 809, and processor 810.
[0139] Those skilled in the art will understand that the electronic device 800 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 810 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0140] In this embodiment, the processor 810 is configured to determine target position information based on preset position information and the posture information of the electronic device in a first wearing state; the target position information is the position information of the user's mouth relative to the origin in the first wearing state; the preset position information is the position information of the user's mouth relative to the origin in a second wearing state, where the origin is the midpoint of the line connecting the first microphone and the second microphone. The processor 810 is further configured to determine a steering vector based on the target position information and the distance between the first microphone and the second microphone. The processor 810 is further configured to determine weight parameters based on the steering vector and the noise space matrix. The processor 810 is further configured to determine an output signal based on the first audio signal collected by the first microphone, the second audio signal collected by the second microphone, and the weight parameters.
[0141] Therefore, by using real-time posture information of the electronic device and preset position information of the user's mouth in the standard wearing state, the target position information of the user's mouth relative to the origin of the midpoint of the dual-microphone connection in the current first wearing state is accurately calculated through spatial rotation transformation. This effectively overcomes the inherent defects of open-back wireless headphones, such as the variable wearing posture and the random displacement of the relative orientation of the dual-microphone array and the user's mouth. Furthermore, a guiding vector is dynamically generated based on the real-time calculated target position of the mouth and the distance between the dual microphones. Combined with the noise space matrix representing the statistical characteristics of noise, the beamforming weight parameters are adaptively solved to achieve real-time adaptive adjustment of the beamforming pickup protection angle, ensuring that the main lobe of the beam is always accurately aligned with the direction of the real-time human voice source. This avoids the problems of human voice attenuation and noise suppression failure caused by the deviation of the human voice from the preset direction in traditional fixed beam solutions. It significantly improves the consistency and robustness of call noise reduction effect under different wearing angles or postures, stably preserves the clarity and naturalness of human voice under various wearing conditions, and improves the overall call quality of the electronic device.
[0142] The electronic device 800 is described in detail below.
[0143] In some embodiments of this application, the processor 810 can also be used to acquire triaxial acceleration data of the electronic device in a first wearing state, provided that the attitude information includes pitch angle and roll angle. Based on triaxial acceleration data, the pitch and roll angles of the electronic device in the first wearing state are determined.
[0144] In some embodiments of this application, the processor 810 may specifically be used to determine, based on target position information, the target orientation angle of the user's mouth relative to the origin in a first wearing state; Determine the guidance vector based on the target orientation angle and spacing.
[0145] In some embodiments of this application, the processor 810 may specifically be used to determine the time delay difference between the arrival of the sound waves emitted from the user's mouth at the first microphone and the second microphone, based on the target orientation angle, spacing and sound speed. Based on the time delay difference and the frequency of the sound wave, determine the relative phase difference between the signals collected by the first microphone and the second microphone when the sound wave emitted from the user's mouth reaches the first microphone and the second microphone; A steering vector is generated based on the relative phase difference.
[0146] In some embodiments of this application, the processor 810 may be specifically used to determine a historical noise space matrix as a noise space matrix when the audio signal includes a user voice signal, wherein the audio signal includes at least one of the following: a first audio signal and a second audio signal; In cases where the audio signal does not include the user's voice signal, the noise space matrix is determined based on the audio signal and the historical noise space matrix.
[0147] In some embodiments of this application, the processor 810 may be specifically used to perform Fourier transform processing on the first audio signal and the second audio signal respectively to obtain a first frequency domain signal of the first audio signal and a second frequency domain signal of the second audio signal. By using weight parameters, the first frequency domain signal and the second frequency domain signal are weighted and summed to obtain the fused frequency domain signal; The inverse Fourier transform of the fused frequency domain signal is performed to obtain the output signal.
[0148] It should be understood that the input unit 804 may include a graphics processing unit (GPU) 8041 and a microphone 8042. The GPU 8041 processes image information of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 806 may include a display panel, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 807 includes at least one of a touch panel 8071 and other input devices 8072. The touch panel 8071 is also called a touch screen. The touch panel 8071 may include two parts: a touch detection device and a touch display. Other input devices 8072 may include, but are not limited to, a physical keyboard, function keys (such as volume display buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.
[0149] The memory 809 can be used to store software programs and various information. The memory 809 may primarily include a first storage area for storing programs or instructions and a second storage area for storing information. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 809 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 809 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0150] Processor 810 may include one or more processing units; optionally, processor 810 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless display signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 810.
[0151] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0152] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0153] In addition, this application embodiment provides another chip, which includes a processor and a display interface. The display interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described audio processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0154] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0155] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the audio processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0156] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0157] Furthermore, it should be noted that the scope of the methods and apparatus in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. In addition, features described with reference to certain examples may be combined in other examples.
[0158] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0159] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An audio processing method applied to an electronic device, characterized in that, The electronic device includes a first microphone and a second microphone, and the method includes: Based on the preset position information and the posture information of the electronic device in the first wearing state, the target position information is determined; the target position information is the position information of the user's mouth relative to the origin in the first wearing state; the preset position information is the position information of the user's mouth relative to the origin in the second wearing state, and the origin is the midpoint of the line connecting the first microphone and the second microphone; Based on the target location information and the distance between the first microphone and the second microphone, a guidance vector is determined; The weighting parameters are determined based on the steering vector and the noise space matrix. The output signal is determined based on the first audio signal collected by the first microphone, the second audio signal collected by the second microphone, and the weighting parameters.
2. The method according to claim 1, characterized in that, The attitude information includes pitch angle and roll angle; before determining the target position information based on the preset position information and the attitude information of the electronic device in the first wearing state, the method further includes: Acquire the triaxial acceleration data of the electronic device in the first wearing state; Based on the triaxial acceleration data, the pitch angle and roll angle of the electronic device in the first wearing state are determined.
3. The method according to claim 1, characterized in that, The step of determining the guidance vector based on the target location information and the distance between the first microphone and the second microphone includes: Based on the target position information, determine the target orientation angle of the user's mouth relative to the origin in the first wearing state; The guidance vector is determined based on the target orientation angle and the spacing.
4. The method according to claim 3, characterized in that, Determining the guidance vector based on the target direction angle and the spacing includes: Based on the target direction angle, the spacing, and the speed of sound, determine the time delay difference between the sound waves emitted from the user's mouth and the first and second microphones. Based on the time delay difference and the frequency of the sound wave, determine the relative phase difference between the signals collected by the first microphone and the second microphone when the sound wave emitted from the user's mouth reaches the first microphone and the second microphone; The steering vector is generated based on the relative phase difference.
5. The method according to claim 1, characterized in that, Before determining the weight parameters based on the steering vector and the noise space matrix, the method further includes: In cases where the audio signal includes a user's voice signal, the historical noise space matrix is determined as the noise space matrix, and the audio signal includes at least one of the following: the first audio signal and the second audio signal; If the audio signal does not include the user's voice signal, the noise space matrix is determined based on the audio signal and the historical noise space matrix.
6. The method according to claim 1, characterized in that, The step of determining the output signal based on the first audio signal acquired by the first microphone, the second audio signal acquired by the second microphone, and the weighting parameters includes: The first audio signal and the second audio signal are subjected to Fourier transform processing respectively to obtain the first frequency domain signal of the first audio signal and the second frequency domain signal of the second audio signal; The first frequency domain signal and the second frequency domain signal are weighted and summed using the weight parameters to obtain the fused frequency domain signal. The output signal is obtained by performing an inverse Fourier transform on the fused frequency domain signal.
7. An audio processing device, applied to an electronic device, characterized in that, The electronic device includes a first microphone and a second microphone, and the device includes: The determining module is used to determine target position information based on preset position information and the posture information of the electronic device in a first wearing state; the target position information is the position information of the user's mouth relative to the origin in the first wearing state; the preset position information is the position information of the user's mouth relative to the origin in a second wearing state, and the origin is the midpoint of the line connecting the first microphone and the second microphone; The determining module is further configured to determine a guiding vector based on the target location information and the distance between the first microphone and the second microphone; The determining module is further configured to determine weight parameters based on the guiding vector and the noise space matrix; The determining module is further configured to determine the output signal based on the first audio signal collected by the first microphone, the second audio signal collected by the second microphone, and the weighting parameter.
8. The apparatus according to claim 7, characterized in that, The audio processing device further includes an acquisition module for acquiring triaxial acceleration data of the electronic device in the first wearing state when the attitude information includes pitch angle and roll angle. The determining module is further configured to determine the pitch angle and roll angle of the electronic device in the first wearing state based on the triaxial acceleration data.
9. The apparatus according to claim 7, characterized in that, The determining module is further configured to, when the audio signal includes a user voice signal, determine the historical noise space matrix as the noise space matrix, wherein the audio signal includes at least one of the following: the first audio signal and the second audio signal; If the audio signal does not include the user's voice signal, the noise space matrix is determined based on the audio signal and the historical noise space matrix.
10. An electronic device, characterized in that, include: A processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the audio processing method as described in any one of claims 1-6.