Audio processing method and device, intelligent earphone and storage medium

By detecting the microphone position and adjusting the audio signal, combined with the methods of weighting and phase calibration, the problem of inconsistent voice loudness during microphone switching in headphones was solved, thus improving the stability and comfort of the audio signal.

CN121815142APending Publication Date: 2026-04-07GEER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-07

Smart Images

  • Figure CN121815142A_ABST
    Figure CN121815142A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing method and device, an intelligent earphone and a storage medium, and relates to the technical field of earphones. The intelligent earphone comprises an earphone body and a moving part arranged on the earphone body, the moving part is provided with a first microphone, and the moving part moves to drive the first microphone to be close to or away from a wearer; the method comprises the steps of obtaining a first audio signal collected by a first microphone when it is detected that the first microphone is started, and collecting a current position of the first microphone; the first audio signal is adjusted according to the current position, a target audio signal is obtained, and the target audio signal is used for audio playing; wherein when the current position is gradually close to the wearing user, the first audio signal is weakened; and when the current position is gradually far away from the wearing user, enhancing the first audio signal. The stability of the sound in the first audio signal can be adjusted in real time according to the current position of the first microphone, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of headphone technology, and more particularly to an audio processing method, device, smart headphone, and storage medium. Background Technology

[0002] Currently, in applications such as gaming and remote conferencing, a dual-microphone design—combining a fixed earcup microphone and a telescopic boom microphone—is typically used to achieve a balance between ease of wear and clear sound pickup. When the user pulls the boom to switch microphones, a mechanical switch or Hall effect sensor is generally used to switch the two microphone circuits on and off. For example, when the user pulls out the boom microphone, the switch is triggered, immediately disconnecting the earcup microphone signal and connecting the boom microphone signal.

[0003] However, this switching mechanism struggles to maintain consistent voice loudness during microphone switching. For example, when a user pulls out the microphone stick, the earpiece microphone disconnects. Since the user typically pulls the microphone stick out quickly, according to the inverse square law of acoustics, the difference in voice loudness between the two can exceed 20 decibels. This switching can cause a significant volume jump in the other party's perception (for example, pulling the microphone stick from the earpiece to the mouth will suddenly change the volume from normal to loud, creating a noticeable auditory impact; pushing the microphone stick back into the earpiece will drop the volume from normal to faint, indistinct distant speech), thus affecting the comfort of the call. Summary of the Invention

[0004] The main objective of this application is to provide an audio processing method, device, smart earphone, and storage medium, which aims to solve the technical problem that it is difficult for microphones to maintain consistent voice loudness during switching.

[0005] To achieve the above objectives, this application proposes an audio processing method applied to a smart earphone; the smart earphone includes: an earphone body and a movable component disposed on the earphone body, the movable component being provided with a first microphone, the movable component moving to move the first microphone closer to or further away from the user; The method includes: If the first microphone is detected to be enabled, the first audio signal collected by the first microphone is acquired, and the current position of the first microphone is acquired. The first audio signal is adjusted according to the current position to obtain a target audio signal, which is used for audio playback; Specifically, when the current position gradually moves closer to the user, the first audio signal is attenuated; when the current position gradually moves away from the user, the first audio signal is amplified.

[0006] In one embodiment, the smart earphone is further provided with a second microphone, which is located on the earphone body; The step of adjusting the first audio signal according to the current position to obtain the target audio signal includes: The second audio signal is acquired through the second microphone; The weighting of the first audio signal and the second audio signal is adjusted according to the current position. The adjustment operation includes: increasing the weight of the first audio signal and decreasing the weight of the second audio signal when the current position gradually moves closer to the user; and decreasing the weight of the first audio signal and increasing the weight of the second audio signal when the current position gradually moves away from the user. Based on the adjusted weighting, the first audio signal and the second audio signal are fused to obtain the target audio signal.

[0007] In one embodiment, the step of adjusting the weighting of the first audio signal and the second audio signal according to the current position includes: Obtain the preset transition coefficient and use the current position as the independent variable parameter; Based on the preset transition coefficient and the independent variable parameter, the first proportion weight of the first audio signal is determined, and the second proportion weight of the second audio signal is obtained according to the first proportion weight. The step of fusing the first audio signal and the second audio signal according to the adjusted weighting to obtain the target audio signal includes: The target audio signal is obtained by adding the product of the first audio signal and the first weighted average and the product of the second audio signal and the second weighted average.

[0008] In one embodiment, the step of adjusting the weighting of the first audio signal and the second audio signal according to the current position includes: Based on the current position, determine the current phase difference data corresponding to the current position from a preset phase difference mapping table. The phase difference mapping table includes the phase difference data between the audio signal of the first microphone and the audio signal of the second microphone at different positions. Based on the current phase difference data, the first audio signal is phase-calibrated to obtain the calibrated first audio signal; Based on the current position, the respective weights of the second audio signal and the calibrated first audio signal are adjusted.

[0009] In one embodiment, before the step of acquiring the first audio signal collected by the first microphone when the first microphone is detected to be enabled, the method further includes: When the smart earphone is worn by the test user, the frequency sweep test signal is collected through the second microphone to obtain the second test audio signal, and the frequency sweep test signal is played in the lip area of ​​the test user; The moving component is moved to different test positions according to a preset ratio, and the frequency sweep test signal is collected through the first microphone to obtain the first test audio signal for each test position; Based on the second test audio signal and the first test audio signal, determine the phase difference data between the first microphone and the second microphone at each test position; A phase difference mapping table is generated based on the phase difference data of each test position. The phase difference mapping table is used to query the phase difference between the first microphone and the second microphone at different positions.

[0010] In one embodiment, the step of fusing the first audio signal and the second audio signal according to the adjusted weighting to obtain the target audio signal includes: Based on the adjusted weighting, the first audio signal and the second audio signal are fused to obtain a fused audio signal; Acquire the first frequency response data of the first microphone, and acquire the second frequency response data of the second microphone at the current location; The difference between the first frequency response data and the second frequency response data is used to obtain the gain deviation of the first microphone relative to the second microphone at the current position; The target audio signal is obtained by compensating for the gain deviation of the fused audio signal.

[0011] In one embodiment, the step of compensating the fused audio signal based on the gain deviation to obtain the target audio signal includes: The fused audio signal is subjected to a fast Fourier transform to obtain a frequency domain signal, and the corresponding gain value is determined based on the gain deviation. Multiply each frequency point of the frequency domain signal by the gain value to obtain the compensated frequency domain signal; The inverse fast Fourier transform is performed on the compensated frequency domain signal to obtain the target audio signal.

[0012] Furthermore, to achieve the above objectives, this application also proposes an audio processing apparatus, the audio processing apparatus comprising: The microphone detection module is used to acquire the first audio signal collected by the first microphone and acquire the current position of the first microphone when the first microphone is detected to be enabled. A signal adjustment module is used to adjust the first audio signal according to the current position to obtain a target audio signal, which is used for audio playback; Specifically, when the current position gradually moves closer to the user, the first audio signal is attenuated; when the current position gradually moves away from the user, the first audio signal is amplified.

[0013] In addition, to achieve the above objectives, this application also proposes a smart earphone, which includes an earphone body and a movable component disposed on the earphone body. The movable component is provided with a first microphone, and the movable component moves to move the first microphone closer to or away from the wearer. The smart earphone further includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the audio processing method described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the audio processing method described above.

[0015] One or more technical solutions proposed in this application have at least the following technical effects: the audio processing method of this application is applied to smart headphones; the smart headphones include: a headphone body and a moving part disposed on the headphone body, the moving part is provided with a first microphone, and the moving part moves to move the first microphone closer to or away from the user; The method includes: when the first microphone is detected to be enabled, acquiring a first audio signal collected by the first microphone and acquiring the current position of the first microphone; adjusting the first audio signal according to the current position to obtain a target audio signal, the target audio signal being used for audio playback; wherein, when the current position gradually approaches the user, the first audio signal is weakened; and when the current position gradually moves away from the user, the first audio signal is enhanced.

[0016] Because this application adjusts the first audio signal by weakening or strengthening it according to the current position of the first microphone when the moving part moves the first microphone closer to or away from the user, thereby reducing sound jumps and achieving sound stability during the entire adjustment process of the moving part, thus improving the comfort of the call. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the audio processing method of this application. Figure 2 This is a flowchart illustrating Embodiment 2 of the audio processing method of this application; Figure 3 This is a schematic diagram of the process for testing the first microphone and the second microphone as provided in Embodiment 2 of this application; Figure 4 This is an overall flowchart of the audio processing procedure provided in Embodiment 3 of this application; Figure 5 This is a block diagram of the module structure of the audio processing device according to an embodiment of this application; Figure 6 This is a schematic diagram of the hardware operating environment involved in the smart earphones in this application embodiment.

[0020] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0022] The main solution proposed in this application is as follows: Currently, in the use cases of over-ear headphones (such as gaming, e-sports, and remote conferencing), in order to achieve a balance between convenient wearing and clear sound pickup, a dual-microphone design of a fixed earcup microphone and a telescopic boom microphone is generally adopted. When the user pulls out the boom to switch microphones, the on / off switching of the two microphone circuits is generally achieved through a mechanical switch or Hall sensor. For example, when the user pulls out the microphone boom, the switch is triggered, at which point the earcup microphone signal is immediately disconnected and the boom microphone signal is connected.

[0023] However, this switching mechanism struggles to maintain consistent voice loudness during microphone switching. For example, when a user pulls out the microphone stick, the earpiece microphone disconnects. Since the user typically pulls the microphone stick quickly, according to the inverse square law of acoustics, the difference in voice loudness between the two can exceed 20 decibels. This switching results in a significant volume jump perceived by the other party. For instance, pulling the microphone stick from the earpiece towards the mouth causes a sudden shift from normal volume to a loud volume, creating a noticeable auditory shock; pushing the microphone stick back towards the earpiece causes the volume to drop from normal to a faint, indistinct, distant sound, thus affecting the comfort of the call.

[0024] To address the aforementioned issues, this application provides an audio processing method that, when a moving component moves a first microphone closer to or further away from the user, adjusts the acquired first audio signal by weakening or enhancing it according to the current position of the first microphone. This reduces sound abrupt changes and achieves sound stability throughout the entire adjustment process of the moving component, thereby improving the comfort of the call listening experience.

[0025] It should be noted that the executing entity of this application embodiment can be a smart headset with a moving part, which has a first microphone. The first microphone can be moved closer to or away from the user by moving the headset; for example, gaming headsets, in-ear headphones, over-ear headphones, etc. The following uses a smart headset (hereinafter referred to as a headset) as an example to describe this embodiment and the following embodiments.

[0026] Based on this, this application proposes an audio processing method according to a first embodiment, referring to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the audio processing method of this application. In this embodiment, the audio processing method is applied to a smart headset; the smart headset includes: a headset body and a movable component disposed on the headset body, the movable component being provided with a first microphone, and the movable component moving to move the first microphone closer to or further away from the user; The audio processing method may include steps S10-S20: Step S10: If the first microphone is detected to be enabled, acquire the first audio signal collected by the first microphone and acquire the current position of the first microphone.

[0027] It should be noted that the smart earphone can be an earphone with at least a stick microphone (i.e., a first microphone), which is disposed on a moving part, and the moving part is disposed on the earphone body.

[0028] The headphone body can be the overall structural part of a smart headphone, including components such as earcups, headband, and speakers that directly contact the user's head and provide audio playback functions. It is the basic mechanical structure of a smart headphone.

[0029] It should also be noted that the moving part can be a retractable microphone telescopic rod assembly, equipped with a first microphone, guide groove or gear, position detection sensor, and other components. The first microphone is located at the end of the moving part and can capture the user's voice. The user can move the first microphone from the earcup of the headphones to the user's lip area using the moving part to capture a clearer user voice.

[0030] Understandably, the current position can be the physical position of the first microphone relative to a certain reference point (such as the user's mouth, the earcups of the headphones, or the length of movement of the moving part, which is not limited in this embodiment) during the process of the user moving the first microphone through the moving part, and can be denoted as p.

[0031] The current position can be determined by the length of movement of the moving part, or by using a microphone to collect the speaker or the user's breathing sound to determine the real-time position of the first microphone. This embodiment does not limit this.

[0032] It is also understood that the first audio signal can be the raw audio data emitted by the user collected by the first microphone during movement, including the user's voice and possible ambient noise. Since moving the first microphone from the earcup to the mouth causes a sudden change in volume from normal to loud, creating a noticeable auditory impact; or moving the first microphone back towards the earcup causes a drop in volume from normal to faint, indistinct distant speech, this switching process results in a significant volume jump in the first audio signal. Therefore, the audio processing method of this application is required to subsequently adjust the first audio signal.

[0033] In actual use, when the user moves the first microphone from the earcup to their mouth via the moving part or pushes the first microphone into the earcup, the first microphone is activated during the movement process. The first microphone will collect the first audio signal when the user speaks in real time. At this time, the headphones can obtain the first audio signal collected by the first microphone and determine the current position of the first microphone by the movement length of the moving part or by collecting the signal from the speaker.

[0034] Step S20: Adjust the first audio signal according to the current position to obtain a target audio signal, which is used for audio playback; Specifically, when the current position gradually moves closer to the user, the first audio signal is attenuated; when the current position gradually moves away from the user, the first audio signal is amplified.

[0035] It should be noted that the target audio signal can be an optimized and adjusted version of the first audio signal, ensuring the consistency and stability of audio playback volume when the first microphone is in different positions. The adjustment of the first audio signal can be based on the proportion of the current position and a preset weight. For example, let the current position p = {0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1}, where 0-1 represents the current position gradually moving closer to the user, and 1-0 represents the current position gradually moving away from the user. The preset weight can then be set to {0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1}. As the current position gradually approaches or moves away from the user, the corresponding weights can be used to adjust the first audio signal in ascending order. Other adjustment methods are also possible, and this embodiment does not limit this.

[0036] In practical use, as the first microphone gradually moves closer to the user, the headphones progressively attenuate the acquired audio signal to prevent excessive volume; conversely, as the first microphone moves further away from the user, it progressively amplifies the audio signal to ensure clarity. This solves the problem of abrupt changes in audio loudness caused by distance variations during microphone movement, allowing users to enjoy a more stable and consistent audio experience.

[0037] Furthermore, considering that in addition to the first microphone (i.e., the stick microphone) the earphone also has an earmuff microphone, in this embodiment, the smart earphone also has a second microphone (i.e., the earmuff microphone), which is located on the earphone body; The step of adjusting the first audio signal according to the current position to obtain the target audio signal includes: Step S21: Acquire the second audio signal through the second microphone.

[0038] It should be noted that the second microphone is an auxiliary microphone fixed to the headphone body (e.g., the earcups), forming a dual-microphone system with the movable first microphone. The second audio signal can be the raw audio data emitted by the user collected by the second microphone, including the user's voice and possible ambient noise.

[0039] Since the second microphone is fixed to the earphone body, unlike the first audio signal which exhibits abrupt changes, the sound intensity of the second audio signal remains essentially stable. Therefore, when the position of the first microphone changes, the second audio signal can be used as a stable reference audio signal for mixing the two.

[0040] Step S22: Adjust the weighting of the first audio signal and the second audio signal according to the current position.

[0041] The adjustment operation includes: increasing the weight of the first audio signal and decreasing the weight of the second audio signal when the current position gradually moves closer to the user; and decreasing the weight of the first audio signal and increasing the weight of the second audio signal when the current position gradually moves away from the user.

[0042] Step S23: According to the adjusted proportion weight, the first audio signal and the second audio signal are fused to obtain the target audio signal.

[0043] Understandably, the weighting can be the respective proportional coefficients of the first audio signal and the second audio signal when they are mixed. For example, when the current position is gradually moving closer to the user, the initial weighting of the first audio signal can be 0.1 and the initial weighting of the second audio signal can be 0.9; when the current position is gradually moving away from the user, the initial weighting of the first audio signal can be 0.9 and the initial weighting of the second audio signal can be 0.1.

[0044] In practical use, after the headphones acquire the second audio signal from the second microphone, they can dynamically adjust the weighting of each signal based on the current position of the first microphone. Specifically, when the microphone is closer to the user, the weighting of the first audio signal increases while the weighting of the second audio signal decreases; when the microphone is farther from the user, the weighting of the first audio signal gradually decreases while the weighting of the second audio signal gradually increases. Finally, based on the adjusted weightings, the first and second audio signals are merged to obtain the target audio signal with stable volume.

[0045] Since this embodiment uses the second audio signal from the second microphone as a stable audio reference and fuses the two sets of microphone signals by adjusting their respective weights, it can achieve a smooth signal transition and further improve the stability of the target audio signal.

[0046] This application provides an audio processing method applied to a smart headset. The smart headset includes a headset body and a movable component disposed on the headset body. The movable component has a first microphone, and the movable component moves the first microphone closer to or further away from the user. The method includes: when the first microphone is detected to be enabled, acquiring a first audio signal collected by the first microphone and acquiring the current position of the first microphone; adjusting the first audio signal according to the current position to obtain a target audio signal, the target audio signal being used for audio playback; wherein, when the current position gradually moves closer to the user, the first audio signal is attenuated; when the current position gradually moves away from the user, the first audio signal is amplified. Because this embodiment adjusts the acquired first audio signal accordingly based on the current position of the first microphone as the movable component moves closer to or further away from the user, it reduces sound jumps and achieves sound stability throughout the adjustment process of the movable component, thereby improving the comfort of call listening.

[0047] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the above embodiment can be referred to the above description, and will not be repeated hereafter. On this basis, a second embodiment of the dialogue method of this application is proposed, please refer to... Figure 2 , Figure 2 This is a flowchart illustrating a second embodiment of the audio processing method of this application. To adjust the respective weights of the first and second audio signals, such as... Figure 2 As shown, in this embodiment, the step of adjusting the weighting of the first audio signal and the second audio signal according to the current position may include: Step S221: Obtain the preset transition coefficient and use the current position as the independent variable parameter.

[0048] Step S222: Based on the preset transition coefficient and the independent variable parameter, determine the first proportion weight of the first audio signal, and obtain the second proportion weight of the second audio signal according to the first proportion weight.

[0049] It should be noted that the preset transition coefficient can be a parameter used to control the smoothness of the weight adjustment of the two sets of microphone signals, and determines the rate of weight change.

[0050] It should also be noted that the first proportion weight can be the proportion coefficient of the first audio signal in the mixed signal. Its value can be determined by a preset transition coefficient and the current position. The weight increases when the position is closer to the user and decreases when it is farther away, so as to achieve a dynamic balance of sound loudness. The second proportion weight can be the proportion coefficient of the second audio signal in the mixed signal. Its value is calculated by subtracting the first proportion weight from 1, ensuring that the sum of the two weights is always 1, thereby maintaining the energy stability of the mixed signal.

[0051] The step of fusing the first audio signal and the second audio signal according to the adjusted weighting to obtain the target audio signal includes: Step S231: Add the product of the first audio signal and the first weighting ratio and the product of the second audio signal and the second weighting ratio to obtain the target audio signal.

[0052] For example, to facilitate understanding of the weighted fusion process of the first and second audio signals, an example is given below. The parameter at the current position is... By weighting the first and second audio signals at the current location, the target audio signal can be obtained as follows: ; in, The target audio signal; It is the first audio signal, which is a time-domain signal; It is the first proportion weight of the first audio signal; It is the second proportion weight of the second audio signal; It is the second audio signal, which is a time-domain signal.

[0053] for and The calculation can be performed using Functions achieve smooth transitions: = ; = ; in, The preset transition coefficient can be set according to the actual debugging situation, for example... To facilitate a quick transition, To ensure a smooth transition.

[0054] Since this embodiment uses the current position as the independent variable parameter and associates it with a preset transition coefficient, the rate of change of the weight can be controlled, and the accurate first proportion weight of the first audio signal can be obtained, thereby avoiding the error of subjective adjustment.

[0055] Furthermore, in order to calibrate the first audio signal, reference Figure 3 , Figure 3 This is a schematic diagram illustrating the testing process for the first and second microphones provided in Embodiment 2 of this application. Figure 3 As shown, in this embodiment, before the step of acquiring the first audio signal collected by the first microphone when the first microphone is detected to be enabled, the method further includes: When the smart earphone is worn by the test user, the frequency sweep test signal is collected through the second microphone to obtain the second test audio signal, and the frequency sweep test signal is played in the lip area of ​​the test user; The moving component is moved to different test positions according to a preset ratio, and the frequency sweep test signal is collected through the first microphone to obtain the first test audio signal for each test position; Based on the second test audio signal and the first test audio signal, determine the phase difference data between the first microphone and the second microphone at each test position; A phase difference mapping table is generated based on the phase difference data of each test position. The phase difference mapping table is used to query the phase difference between the first microphone and the second microphone at different positions.

[0056] It should be noted that the frequency sweep test signal can be a standardized audio signal with continuously varying frequencies, covering the range of human hearing (e.g., 20Hz-20kHz). In this embodiment, the frequency sweep test signal is played in the lip area of ​​the test user to simulate the sound source location when the user speaks, providing a standardized sound source input for the testing of the first and second microphones. The second test audio signal can be audio data collected by the second microphone when the frequency sweep test signal is played, including the frequency response of the signal, denoted as... f represents the frequency of the signal.

[0057] It should also be noted that the preset ratio can be the movement step size of the moving part or the position division rule. For example, dividing the range where the moving part is pulled out into 10 equal parts, and when fully retracted... When fully pulled out .Pick .

[0058] At this point, the microphone can be moved to different test positions at percentages of 10%, 20%,...100% to systematically collect the first test audio signal from the first microphone at different distances. The first test audio signal can be the audio data collected by the first microphone when the frequency sweep test signal is played, including the signal's frequency response, denoted as... .

[0059] Understandably, phase difference data can be the phase angle difference between the signals acquired by the first and second microphones at the same frequency point, reflecting the impact of positional changes on the pickup phase of the first microphone. For example, at a frequency of 1kHz, if the phase of the first microphone signal is 30° and that of the second microphone is 10°, then the phase difference is 20°. The phase difference mapping table can be a lookup table indexed by the position of the first microphone, storing the phase difference data corresponding to each position. By querying this table using the real-time position, the phase compensation value can be quickly obtained.

[0060] For example, the phase difference data between the first test audio signal and the second test audio signal at different positions can be expressed as: ; in, This indicates the phase of the second test audio signal; This indicates the phase of the first test audio signal at different positions; This indicates the phase difference between the first test audio signal and the second test audio signal at different test positions.

[0061] In this embodiment, as Figure 3 As shown, before using the smart earphones, the manufacturer can play a sweep frequency test signal in the user's mouth area to simulate the sound source location of the user's speech. At this time, the sweep frequency test signal is collected by the second microphone, and the second test audio signal of the second microphone is recorded. Then, when the moving part is tested at different positions, the first microphone records the first test audio signal, and the position data for each position is recorded accordingly. Finally, based on the second and first test audio signals, the phase difference data between the two sets of microphones at each test position is determined, and the phase difference data corresponding to each position is stored using the test position as an index, generating a phase difference mapping table. Thus, through the phase difference mapping table, the phase difference data of the two sets of microphones at different distances can be accurately obtained, providing a quantitative basis for subsequent phase correction.

[0062] Furthermore, considering the lack of calibration between the first and second audio signals, in this embodiment, the step of adjusting the respective weights of the first and second audio signals based on the current position may include: Based on the current position, determine the current phase difference data corresponding to the current position from a preset phase difference mapping table. The phase difference mapping table includes the phase difference data between the audio signal of the first microphone and the audio signal of the second microphone at different positions. Based on the current phase difference data, the first audio signal is phase-calibrated to obtain the calibrated first audio signal; Based on the current position, the respective weights of the second audio signal and the calibrated first audio signal are adjusted.

[0063] It should be noted that the current phase difference data can be the phase difference value corresponding to the current position, obtained from the phase difference mapping table. For example, if the current position is when the moving part is pulled out to 70%, the phase difference data of each frequency point at that position is extracted from the table for subsequent calibration.

[0064] During phase calibration, the phase of the first audio signal can be adjusted in the frequency domain. For example, if the phase difference data at a certain frequency point in the first audio signal is -10°, the phase of the first audio signal at that frequency is increased by 10° to align it with the phase of the second audio signal, thus eliminating phase distortion caused by changes in microphone position.

[0065] For example, to facilitate understanding of the phase calibration process, an example is given below. Based on microphone position parameters... Frequency domain phase compensation is performed on the first audio signal: ; in, ; It is the first audio signal, which is a time-domain signal; FFT stands for Fast Fourier Transform. The transformed first audio signal is a frequency domain signal; j is a fixed coefficient. The first audio signal after phase calibration is a frequency domain signal.

[0066] The first audio signal after phase calibration is as follows: ; Wherein, IFFT stands for Inverse Fast Fourier Transform; The first audio signal after calibration is a time-domain signal.

[0067] Therefore, after calibrating the first audio signal, the target audio signal is obtained by weighting the first and second audio signals calibrated at the current position as follows: ; in, The calibrated target audio signal; It is the first audio signal after calibration, which is a time-domain signal; It is the first proportion weight of the first audio signal; It is the second proportion weight of the second audio signal; It is the second audio signal, which is a time-domain signal.

[0068] Since this embodiment determines the current phase difference data of the first audio signal by the current position and uses the current phase difference data to perform phase calibration on the first audio signal, it can accurately compensate for the phase inconsistency caused by the position change of the first microphone, thereby ensuring that the calibrated first audio signal is aligned with the second audio signal.

[0069] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. On this basis, in order to improve the sound quality of the target audio signal, in this embodiment, after the step of fusing the first audio signal and the second audio signal according to the adjusted weighting to obtain the target audio signal, the method further includes: Step S41: Obtain the first frequency response data of the first microphone and the second frequency response data of the second microphone at the current location.

[0070] It should be noted that the first frequency response data (denoted as...) The frequency response curve of the first microphone, obtained through the above tests, can be represented by frequency on the horizontal axis and gain (i.e., decibel value) on the vertical axis, as described above. This reflects the first microphone's ability to pick up sounds of different frequencies. Similarly, the second frequency response data (denoted as...) The frequency response characteristic curve of the second microphone can be obtained through the above tests.

[0071] Step S42: Subtract the first frequency response data and the second frequency response data to obtain the gain deviation of the first microphone relative to the second microphone at the current position.

[0072] Understandably, gain deviation can be the difference in gain between the first and second frequency response data at the same frequency point. It reflects the difference in sound pickup intensity between the first and second microphones at various frequency points and is a core parameter for subsequent compensation. For example, at 500Hz, if the first microphone gain is +3dB and the second microphone gain is +1dB, then the gain deviation is +2dB.

[0073] Step S43: Compensate the target audio signal based on the gain deviation to obtain a compensated target audio signal, which is used for audio playback.

[0074] Understandably, the compensated target audio signal can be an audio signal that has undergone gain deviation correction. By adjusting each frequency point of the target signal in the frequency domain according to the opposite value of the gain deviation (e.g., if the deviation is +2dB, then compensate by -2dB), its frequency response characteristics are made consistent with the second microphone, ensuring the naturalness and consistency of the timbre during the final audio playback.

[0075] Furthermore, in order to compensate the target audio signal, the step of compensating the target audio signal based on the gain deviation to obtain the compensated target audio signal includes: The target audio signal is subjected to a fast Fourier transform to obtain a frequency domain signal, and the corresponding gain value is determined based on the gain deviation. Multiply each frequency point of the frequency domain signal by the gain value to obtain the compensated frequency domain signal; The compensated frequency domain signal is subjected to inverse fast Fourier transform to obtain the compensated target audio signal.

[0076] It should be noted that the Fast Fourier Transform (FFT) can be an algorithm that converts a time-domain signal into a frequency-domain signal. The frequency-domain signal can be a complex array of the target audio signal after FFT transformation, containing amplitude (gain) and phase information at each frequency point. This signal can be used for direct frequency domain compensation.

[0077] It should also be noted that the gain value can be a compensation coefficient determined based on the gain deviation, usually the negative of the gain deviation (e.g., if the deviation is +2dB, the gain value is -2dB). This value is used to adjust the amplitude of the frequency domain signal to eliminate frequency response differences.

[0078] Understandably, the Inverse Fast Fourier Transform (IFFT) is the inverse operation of the FFT, which can convert the compensated frequency domain signal back into a time domain signal for subsequent audio playback.

[0079] In this embodiment, the signal is converted to the frequency domain by FFT, which allows for precise adjustment of the gain at each frequency point, eliminating the local frequency response problem that is difficult to handle by traditional time-domain filtering.

[0080] For example, to facilitate understanding of the above compensation process, an example is given below. After obtaining the target audio signal, the target audio signal can be compensated as follows: ; in, That is, for the aforementioned target audio signal conduct After transformation, the resulting target audio signal is a frequency domain signal; That is, the first microphone of the moving part is in position ,frequency At this point, the gain deviation relative to the second microphone is considered; based on the gain deviation, each frequency point of the transformed target audio signal is multiplied by the corresponding gain value, and finally... The compensated target audio signal representing the time domain can then be obtained. .

[0081] In the embodiments of this application, such as Figure 4 As shown, Figure 4 This is an overall flowchart of the audio processing procedure provided in Embodiment 3 of this application. After obtaining the current position of the first microphone, the headphones can first calibrate the phase according to the current position; then, the two sets of microphones are adjusted and fused to obtain the target audio signal. Finally, by utilizing the gain deviation of the first microphone relative to the second microphone at the current position, targeted dynamic compensation can be performed to eliminate frequency response imbalances caused by microphone position, individual differences, or environmental factors. This not only reduces the abrupt changes in sound during the switching process of the moving parts, but also further improves the sound quality of the played audio, enhancing the audio experience for headphone users when using the moving parts.

[0082] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the audio processing method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0083] This application also provides an audio processing device, please refer to... Figure 5 , Figure 5 This is a block diagram of the module structure of the audio processing device according to an embodiment of this application; in this embodiment, the audio processing device includes: The microphone detection module 501 is used to acquire the first audio signal collected by the first microphone and acquire the current position of the first microphone when the first microphone is detected to be enabled. The signal adjustment module 502 is used to adjust the first audio signal according to the current position to obtain a target audio signal, which is used for audio playback; Specifically, when the current position gradually moves closer to the user, the first audio signal is attenuated; when the current position gradually moves away from the user, the first audio signal is amplified.

[0084] In this embodiment, when the moving component moves the first microphone closer to or further away from the user, the microphone detection module collects the current position of the first microphone; then, the signal adjustment module adjusts the collected first audio signal according to the current position of the first microphone to reduce sound jumps, thereby achieving sound stability during the entire adjustment process of the moving component and improving the comfort of the call.

[0085] In one implementation, the signal adjustment module 502 is further configured to acquire a second audio signal via the second microphone; adjust the weighting of the first audio signal and the second audio signal according to the current position; wherein the adjustment operation includes: increasing the weighting of the first audio signal and decreasing the weighting of the second audio signal when the current position gradually moves closer to the user; decreasing the weighting of the first audio signal and increasing the weighting of the second audio signal when the current position gradually moves away from the user; and fusing the first audio signal and the second audio signal according to the adjusted weighting to obtain a target audio signal.

[0086] In one implementation, the signal adjustment module 502 is further configured to obtain a preset transition coefficient and use the current position as an independent variable parameter; determine a first proportion weight of the first audio signal based on the preset transition coefficient and the independent variable parameter, and obtain a second proportion weight of the second audio signal according to the first proportion weight; and add the product between the first audio signal and the first proportion weight and the product between the second audio signal and the second proportion weight to obtain the target audio signal.

[0087] In one implementation, the signal adjustment module 502 is further configured to determine the current phase difference data corresponding to the current position from a preset phase difference mapping table, the phase difference mapping table including phase difference data between the audio signals of the first microphone and the audio signals of the second microphone at different positions; perform phase calibration on the first audio signal based on the current phase difference data to obtain a calibrated first audio signal; and adjust the respective weights of the second audio signal and the calibrated first audio signal according to the current position.

[0088] In one implementation, the microphone detection module 501 is further configured to: acquire a frequency sweep test signal through the second microphone to obtain a second test audio signal when the smart earphone is worn by the test user; play the frequency sweep test signal in the lip area of ​​the test user; move the moving component to different test positions according to a preset ratio and acquire the frequency sweep test signal through the first microphone to obtain a first test audio signal for each test position; determine the phase difference data between the first microphone and the second microphone at each test position based on the second test audio signal and the first test audio signal; and generate a phase difference mapping table according to the phase difference data of each test position, the phase difference mapping table being used to query the phase difference between the first microphone and the second microphone at different positions.

[0089] In one implementation, the signal adjustment module 502 is further configured to acquire first frequency response data of the first microphone and second frequency response data of the second microphone at the current position; subtract the first frequency response data and the second frequency response data to obtain the gain deviation of the first microphone relative to the second microphone at the current position; compensate the target audio signal based on the gain deviation to obtain a compensated target audio signal, which is used for audio playback.

[0090] In one implementation, the signal adjustment module 502 is further configured to perform a fast Fourier transform on the target audio signal to obtain a frequency domain signal, and determine the corresponding gain value based on the gain deviation; multiply each frequency point of the frequency domain signal by the gain value to obtain a compensated frequency domain signal; and perform an inverse fast Fourier transform on the compensated frequency domain signal to obtain a compensated target audio signal.

[0091] Other embodiments or specific implementations of the audio processing device of this application can be found in the above-described method embodiments, and will not be repeated here.

[0092] The audio processing apparatus provided in this application, employing the audio processing method described in the above embodiments, can solve the technical problem of difficulty in maintaining consistent speech loudness during microphone switching. Compared with the prior art, the beneficial effects of the audio processing apparatus provided in this application are the same as those of the audio processing method provided in the above embodiments, and other technical features in the audio processing apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0093] This application provides a smart earphone, which includes an earphone body and a movable component disposed on the earphone body. The movable component is provided with a first microphone, and the movable component moves to move the first microphone closer to or away from the wearer.

[0094] The smart headset further includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the audio processing methods in the above embodiments.

[0095] The following is for reference. Figure 6 , Figure 6 This is a schematic diagram of the hardware operating environment involved in the smart headphones in the embodiments of this application, showing a structural schematic diagram suitable for implementing the smart headphones in the embodiments of this application. The smart headphones in the embodiments of this application may include, but are not limited to, over-ear headphones, gaming headphones, in-ear headphones, etc. Figure 6The smart earphones shown are merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0096] like Figure 6 As shown, a smart headset may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the smart headset. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a microphone; an output device 1008 including, for example, a speaker; a storage device 1003 including, for example, a hard disk; and a communication device 1009. The communication device 1009 allows the smart headset to communicate wirelessly or wiredly with other devices to exchange data. Although a smart headset with various systems is shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented or possessed alternatively.

[0097] The smart earphone provided in this application, employing the audio processing method described in the above embodiments, can solve the technical problem of difficulty in maintaining consistent voice loudness during microphone switching. Compared with the prior art, the beneficial effects of the smart earphone provided in this application are the same as those of the audio processing method provided in the above embodiments, and other technical features of the smart earphone are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0098] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0099] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0100] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the audio processing method in the above embodiments.

[0101] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0102] The aforementioned computer-readable storage medium may be included in the smart headphones; or it may exist independently and not assembled into the smart headphones.

[0103] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the smart earphones, cause the smart earphones to: acquire a first audio signal collected by the first microphone and acquire the current position of the first microphone when the first microphone is detected to be enabled; adjust the first audio signal according to the current position to obtain a target audio signal, the target audio signal being used for audio playback; wherein, when the current position gradually moves closer to the user, the first audio signal is attenuated; and when the current position gradually moves away from the user, the first audio signal is amplified.

[0104] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0106] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0107] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for performing the above-described audio processing method, which can solve the technical problem of difficulty in maintaining consistent voice loudness during microphone switching. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the audio processing method provided in the above embodiments, and will not be repeated here.

[0108] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.

Claims

1. An audio processing method, characterized in that, The audio processing method is applied to smart headphones; the smart headphones include: a headphone body and a movable component disposed on the headphone body, the movable component is provided with a first microphone, and the movable component moves to move the first microphone closer to or away from the user; The method includes: If the first microphone is detected to be enabled, the first audio signal collected by the first microphone is acquired, and the current position of the first microphone is acquired. The first audio signal is adjusted according to the current position to obtain a target audio signal, which is used for audio playback; Specifically, when the current position gradually moves closer to the user, the first audio signal is attenuated; when the current position gradually moves away from the user, the first audio signal is amplified.

2. The method as described in claim 1, characterized in that, The smart earphone is also equipped with a second microphone, which is located on the earphone body; The step of adjusting the first audio signal according to the current position to obtain the target audio signal includes: The second audio signal is acquired through the second microphone; The weighting of the first audio signal and the second audio signal is adjusted according to the current position. The adjustment operation includes: increasing the weight of the first audio signal and decreasing the weight of the second audio signal when the current position gradually moves closer to the user; and decreasing the weight of the first audio signal and increasing the weight of the second audio signal when the current position gradually moves away from the user. Based on the adjusted weighting, the first audio signal and the second audio signal are fused to obtain the target audio signal.

3. The method as described in claim 2, characterized in that, The step of adjusting the weighting of the first audio signal and the second audio signal according to the current position includes: Obtain the preset transition coefficient and use the current position as the independent variable parameter; Based on the preset transition coefficient and the independent variable parameter, the first proportion weight of the first audio signal is determined, and the second proportion weight of the second audio signal is obtained according to the first proportion weight. The step of fusing the first audio signal and the second audio signal according to the adjusted weighting to obtain the target audio signal includes: The target audio signal is obtained by adding the product of the first audio signal and the first weighted average and the product of the second audio signal and the second weighted average.

4. The method as described in claim 2, characterized in that, The step of adjusting the weighting of the first audio signal and the second audio signal according to the current position includes: Based on the current position, determine the current phase difference data corresponding to the current position from a preset phase difference mapping table. The phase difference mapping table includes phase difference data between the audio signals of the first microphone and the audio signals of the second microphone at different positions. Based on the current phase difference data, the first audio signal is phase-calibrated to obtain the calibrated first audio signal; Based on the current position, the respective weights of the second audio signal and the calibrated first audio signal are adjusted.

5. The method as described in claim 4, characterized in that, Before the step of acquiring the first audio signal collected by the first microphone when the first microphone is detected to be enabled, the method further includes: When the smart earphone is worn by the test user, the frequency sweep test signal is collected through the second microphone to obtain the second test audio signal, and the frequency sweep test signal is played in the lip area of ​​the test user; The moving component is moved to different test positions according to a preset ratio, and the frequency sweep test signal is collected through the first microphone to obtain the first test audio signal for each test position; Based on the second test audio signal and the first test audio signal, determine the phase difference data between the first microphone and the second microphone at each test position; A phase difference mapping table is generated based on the phase difference data of each test position. The phase difference mapping table is used to query the phase difference between the first microphone and the second microphone at different positions.

6. The method according to any one of claims 2 to 5, characterized in that, After the step of fusing the first audio signal and the second audio signal according to the adjusted weighting to obtain the target audio signal, the method further includes: Acquire the first frequency response data of the first microphone, and acquire the second frequency response data of the second microphone at the current location; The difference between the first frequency response data and the second frequency response data is used to obtain the gain deviation of the first microphone relative to the second microphone at the current position; The target audio signal is compensated based on the gain deviation to obtain a compensated target audio signal, which is then used for audio playback.

7. The method as described in claim 6, characterized in that, The step of compensating the target audio signal based on the gain deviation to obtain the compensated target audio signal includes: The target audio signal is subjected to a fast Fourier transform to obtain a frequency domain signal, and the corresponding gain value is determined based on the gain deviation. Multiply each frequency point of the frequency domain signal by the gain value to obtain the compensated frequency domain signal; The compensated frequency domain signal is subjected to inverse fast Fourier transform to obtain the compensated target audio signal.

8. An audio processing apparatus, characterized in that, The device includes: The microphone detection module is used to acquire the first audio signal collected by the first microphone and acquire the current position of the first microphone when the first microphone is detected to be enabled. A signal adjustment module is used to adjust the first audio signal according to the current position to obtain a target audio signal, which is used for audio playback; Specifically, when the current position gradually moves closer to the user, the first audio signal is attenuated; when the current position gradually moves away from the user, the first audio signal is amplified.

9. A smart earphone, characterized in that, The smart earphone includes an earphone body and a movable component disposed on the earphone body. The movable component is provided with a first microphone. The movable component moves to move the first microphone closer to or away from the wearer. The smart earphone further includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the audio processing method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the audio processing method as described in any one of claims 1 to 7.