Dynamic pickup method and device, electronic equipment and storage medium
Through the dynamic sound pickup method, alternative sound pickup audio from multiple microphones are used for screening and synthesis, which solves the problem that headphones are difficult to capture sound from all directions, achieves more accurate sound source positioning and information extraction, and improves the sound pickup effect of headphones.
Patent Information
- Application Number
- CN202510164146.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-03
AI Technical Summary
Existing headphones are difficult to accurately capture effective sound information from all directions when worn when using them, resulting in the wearer being unable to effectively perceive important sounds from outside.
The dynamic sound pickup method is adopted to obtain the alternative sound pickup audio of all the alternative microphones on the headset, perform audio filtering and synthesis, dynamically determine the sound pickup direction, accurately locate the sound source, and extract and synthesize the effective information from the secondary sound pickup audio into the main sound pickup audio.
It improves the sound pick-up effect of the headphones, allowing the wearer to more accurately perceive important sound information from all directions, and enhances the wearer's external perception ability.
Smart Images

Figure CN120091250A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of earphones, and particularly to a dynamic sound pickup method and device, an electronic device, and a storage medium. Background Art
[0002] When an earphone wearer uses an earphone, the earphone will block the wearer's perception ability of some external sounds. For example, when walking on the street, the wearer may not hear the horn of a vehicle coming from behind, or may not notice the emergency braking sound of a train in the subway, or may not hear someone calling the earphone wearer. Generally, in order to improve the earphone wearer's perception ability of effective external information, a directional sound pickup technology is adopted. However, since sound sources come from all directions in the real environment, it is difficult for directional sound pickup to accurately capture effective information. Therefore, how to improve the sound pickup effect of earphones has become an urgent problem to be solved. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose a dynamic sound pickup method and device, an electronic device, and a storage medium, aiming to improve the sound pickup effect of earphones.
[0004] To achieve the above purpose, in the first aspect of the embodiments of this application, a dynamic sound pickup method is proposed, which is applied to an earphone. The earphone includes at least one alternative microphone. The method includes:
[0005] Obtain the alternative sound pickup audio of all the alternative microphones;
[0006] Perform audio screening on the alternative sound pickup audio to obtain a main sound pickup audio and a secondary sound pickup audio;
[0007] Perform audio synthesis on the main sound pickup audio and the secondary sound pickup audio to obtain a target playback audio;
[0008] Play the target playback audio based on the earphone.
[0009] In some embodiments, the performing audio screening on the alternative sound pickup audio to obtain a main sound pickup audio and a secondary sound pickup audio includes:
[0010] Perform environmental audio evaluation on the alternative sound pickup audio to obtain environmental audio evaluation data;
[0011] Perform human voice target detection on the alternative sound pickup audio to obtain human voice target detection data;
[0012] Perform audio screening on the alternative sound pickup audio according to the environmental audio evaluation data and the human voice target detection data to obtain the main sound pickup audio;
[0013] Audio screening is performed on the alternative pick-up audio according to the main pick-up audio to obtain the secondary pick-up audio.
[0014] In some embodiments, performing voice target detection on the alternative pick-up audio to obtain voice target detection data includes:
[0015] Performing semantic conversion on the alternative pick-up audio to obtain alternative semantic data;
[0016] Performing intent recognition on the alternative semantic data to obtain alternative intent data;
[0017] Performing voice target detection on the alternative pick-up audio according to the alternative intent data to obtain the voice target detection data.
[0018] In some embodiments, performing audio synthesis on the main pick-up audio and the secondary pick-up audio to obtain the target playback audio includes:
[0019] Performing voice extraction on the secondary pick-up audio according to the main pick-up audio to obtain the target enhanced audio;
[0020] Performing echo cancellation on the target enhanced audio to obtain the selected enhanced audio;
[0021] Performing audio merging on the main pick-up audio according to the selected enhanced audio to obtain the target playback audio.
[0022] In some embodiments, the performing voice extraction on the secondary pick-up audio according to the main pick-up audio to obtain the target enhanced audio includes:
[0023] Performing frequency domain conversion on the main pick-up audio to obtain main frequency domain data, and performing frequency domain conversion on the secondary pick-up audio to obtain secondary frequency domain data;
[0024] Performing voice frequency domain screening on the secondary frequency domain data according to the main frequency domain data to obtain target voice frequency domain data;
[0025] Performing inverse frequency domain conversion on the target voice frequency domain data to obtain the target enhanced audio.
[0026] In some embodiments, the performing voice frequency domain screening on the secondary frequency domain data according to the main frequency domain data to obtain target voice frequency domain data includes:
[0027] Performing common frequency domain feature extraction on the secondary frequency domain data according to the main frequency domain data to obtain voice common frequency domain features;
[0028] Performing beamforming processing on the voice common frequency domain features to obtain beam frequency domain features;
[0029] Perform coherent analysis processing on the common frequency domain features of the human voice to obtain coherent frequency domain features;
[0030] Perform adaptive filtering on the common frequency domain features of the human voice to obtain filtered frequency domain features;
[0031] Perform aggregation calculation based on the beam frequency domain features, the coherent frequency domain features, and the filtered frequency domain features to obtain the target human voice frequency domain data.
[0032] In some embodiments, the earphone includes a three-axis accelerometer and a three-axis gyroscope. After playing the target playback audio based on the earphone, the method further includes:
[0033] Obtain the sound source direction of the main pick-up audio to obtain main sound source direction data;
[0034] Perform head pose calculation based on the three-axis accelerometer and the three-axis gyroscope to obtain head pose data;
[0035] Perform direction adjustment based on the main sound source direction data and the head pose data to obtain target sound source direction data;
[0036] Perform beamforming direction adjustment on all the alternative microphones according to the target sound source direction data.
[0037] To achieve the above object, a second aspect of the embodiments of the present application provides a dynamic sound pickup device, which is applied to an earphone. The earphone includes at least one alternative microphone, and the device includes:
[0038] A data acquisition module, configured to acquire alternative pick-up audio of all the alternative microphones;
[0039] An audio screening module, configured to perform audio screening on the alternative pick-up audio to obtain main pick-up audio and secondary pick-up audio;
[0040] An audio synthesis module, configured to perform audio synthesis on the main pick-up audio and the secondary pick-up audio to obtain target playback audio;
[0041] An audio playback module, configured to perform audio playback on the target playback audio based on the earphone.
[0042] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0043] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method described in the first aspect above.
[0044] A dynamic sound pickup method and device, an electronic device, and a storage medium provided by the present application first obtain the alternative sound pickup audio of all alternative microphones on the earphone, and then perform audio screening on the alternative sound pickup audio to obtain the main sound pickup audio and the secondary sound pickup audio, so as to dynamically determine the sound pickup direction of the earphone when the wearer is wearing the earphone, thereby accurately positioning the sound source and improving the sound pickup effect of the earphone; further, perform audio synthesis on the main sound pickup audio and the secondary sound pickup audio, so as to extract the effective information in the secondary sound pickup audio and synthesize it into the main sound pickup audio to achieve the extraction of effective information; finally, play the target audio based on the earphone to achieve dynamic directional sound pickup of sound sources from all directions in the real environment, effectively screen the information from all directions, and ultimately improve the sound pickup effect of the earphone. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flowchart of the dynamic sound pickup method provided by the embodiments of the present application;
[0046] Figure 2 is Figure 1 a flowchart of step S102 in
[0047] Figure 3 is Figure 2 a flowchart of step S202 in
[0048] Figure 4 is Figure 1 a flowchart of step S103 in
[0049] Figure 5 is Figure 4 a flowchart of step S401 in
[0050] Figure 6 is Figure 5 a flowchart of step S502 in
[0051] Figure 7 is a flowchart of the dynamic sound pickup method provided by another embodiment of the present application;
[0052] Figure 8 is a schematic structural diagram of the dynamic sound pickup device provided by the embodiments of the present application;
[0053] Figure 9 is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0055] It should be noted that although functional module division is carried out in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0057] First, several nouns involved in this application are analyzed:
[0058] Beamforming: Beamforming is a signal processing technology used to control the directivity of signals in an array antenna to optimize the reception or transmission of signals in a specific direction. It is a core component of wireless communication technology and belongs to the field of signal processing. Beamforming technology enhances signals in a specific direction and suppresses signals in other directions by adjusting the phase and amplitude of each element on the array antenna. This technology is widely used in radar, sonar, wireless networks, and mobile communications, and is particularly important in large-scale MIMO (multiple input multiple output) systems that support 5G and future communication systems. Beamforming not only improves the signal reception quality but also enhances the spatial multiplexing ability of the system, effectively improving the overall performance and efficiency of the communication system. In addition, beamforming technology also has important applications in fields such as sound capture, biomedical imaging, and seismic data analysis, helping to improve the signal positioning accuracy and resolution.
[0059] Coherence analysis: Coherence analysis is a signal processing technique used to evaluate the similarity or consistency of two signals in the frequency domain. This technique belongs to the fields of data analysis and signal processing and is widely applied in multiple scientific and engineering domains. By calculating the coherence function between different signals, the degree of phase correlation of these signals at specific frequencies can be determined, thereby evaluating whether they have statistical correlation or are generated from the same or related processes. Coherence analysis is commonly used in neuroscience, seismology, meteorology, mechanical engineering, and communication system analysis to analyze time series data, identify, and understand signal sources and propagation paths. For example, in neuroscience, by analyzing the coherence of electrical signals between different regions of the brain, the functional connectivity between brain regions can be studied. In the engineering field, coherence analysis helps analyze vibration data for fault diagnosis and structural health monitoring. Coherence analysis provides a powerful tool to understand the behavior and dynamics of complex systems by revealing the internal correlations between signals.
[0060] When a headphone wearer uses headphones, the headphones will block the wearer's ability to perceive some external sounds. For example, when walking on the street, the wearer may not be able to hear the horn of an approaching vehicle from behind, or in the subway, may not be able to detect the emergency braking sound of the train, or may not be able to hear someone calling the headphone wearer. Usually, in order to improve the headphone wearer's ability to perceive effective external information, a directional sound pickup technology is adopted. However, since sound sources come from all directions in the real environment, it is difficult for directional sound pickup to accurately capture effective information. Therefore, how to improve the sound pickup effect of headphones has become an urgent problem to be solved.
[0061] Based on this, the embodiments of the present application provide a dynamic sound pickup method, device, electronic device, and storage medium, aiming to improve the sound pickup effect of headphones.
[0062] The dynamic sound pickup method, device, electronic device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the dynamic sound pickup method in the embodiments of the present application is described.
[0063] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0064] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0065] The dynamic sound pickup method provided by the embodiments of the present application relates to the technical field of earphones. The dynamic sound pickup method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the dynamic sound pickup method, etc., but is not limited to the above forms.
[0066] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0067] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0068] Figure 1 is an optional flowchart of the dynamic sound pickup method provided by the embodiments of the present application. Figure 1 The method in is applied to an earphone. The earphone includes at least one alternative microphone. The dynamic sound pickup method may include, but is not limited to, steps S101 to S104.
[0069] Step S101, obtain the alternative sound pickup audio of all alternative microphones.
[0070] Step S102, perform audio screening on the alternative sound pickup audio to obtain the main sound pickup audio and the secondary sound pickup audio.
[0071] Step S103, perform audio synthesis on the main sound pickup audio and the secondary sound pickup audio to obtain the target playback audio.
[0072] Step S104, perform audio playback on the target playback audio based on the earphone.
[0073] Steps S101 to S104 illustrated in the embodiments of the present application first obtain the alternative sound pickup audio of all alternative microphones on the earphone, and then perform audio screening on the alternative sound pickup audio to obtain the main sound pickup audio and the secondary sound pickup audio, thereby dynamically determining the sound pickup direction of the earphone when the wearer is wearing the earphone, so as to accurately locate the sound source and improve the sound pickup effect of the earphone; further, perform audio synthesis on the main sound pickup audio and the secondary sound pickup audio, so as to extract the effective information in the secondary sound pickup audio and synthesize it into the main sound pickup audio to achieve the extraction of effective information; finally, perform audio playback on the target audio based on the earphone to achieve dynamic directional sound pickup of sound sources from all directions in the real environment, effectively screen the information from all directions, and ultimately improve the sound pickup effect of the earphone.
[0074] In step S101 of some embodiments, the alternative microphone is a microphone in the microphone array on the earphone. In one embodiment, the earphone is divided into a left earphone and a right earphone. The earphone includes a three-axis accelerometer and a three-axis gyroscope. When the earphone wearer wears the earphone, the accelerometer data and gyroscope data of the three-axis accelerometer on the earphone are obtained, and then the earphone orientation is calculated based on the accelerometer data and gyroscope data to determine whether the earphone is worn on the left ear or the right ear. The audio of six microphones on the earphone is obtained. There are three microphones on the left earphone and the right earphone respectively, and the orientations of the three microphones are the front of the face, the side of the head, and the back of the face. The pick-up audio of all six microphones on the left earphone and the right earphone is obtained, and the pick-up audio of all alternative microphones is obtained to obtain the alternative pick-up audio.
[0075] Please refer to Figure 2 , in some embodiments, step S102 may include but is not limited to steps S201 to S204:
[0076] Step S201, perform an environmental audio evaluation on the alternative pick-up audio to obtain environmental audio evaluation data;
[0077] Step S202, perform a human voice target detection on the alternative pick-up audio to obtain human voice target detection data;
[0078] Step S203, perform audio screening on the alternative pick-up audio according to the environmental audio evaluation data and the human voice target detection data to obtain the main pick-up audio;
[0079] Step S204, perform audio screening on the alternative pick-up audio according to the main pick-up audio to obtain the secondary pick-up audio.
[0080] Steps S201 to S204 illustrated in the embodiments of the present application, by performing an environmental audio evaluation on the alternative pick-up audio to obtain environmental audio evaluation data, then performing a human voice target detection on the alternative pick-up audio to obtain human voice target detection data, then performing audio screening on the alternative pick-up audio according to the environmental audio evaluation data and the human voice target detection data to obtain the main pick-up audio, and finally performing audio screening on the alternative pick-up audio according to the main pick-up audio to obtain the secondary pick-up audio, thereby improving the perception ability of external sounds. On the one hand, the environmental audio is evaluated to determine the effective information in case of emergency, and on the other hand, the human voice audio is evaluated to detect the human voice audio of others to the earphone wearer, so as to perform directional sound source pick-up from two dimensions and improve the pick-up effect of the earphone.
[0081] In step S201 of some embodiments, the environmental audio evaluation is to evaluate the ambient sound in the alternative picked-up audio to determine whether the sound source includes a sound source in an emergency. Specifically, the alternative picked-up audio is compared with a preset audio in terms of features. For example, the preset audio includes braking sounds, siren sounds, explosion sounds, or sounds emitted in other dangerous situations. After comparing the features of the alternative picked-up audio with the preset audio, environmental audio evaluation data is obtained. The environmental audio evaluation data is a value from 0 to 1, representing the similarity between the alternative picked-up audio and the preset audio.
[0082] Please refer to Figure 3 , in some embodiments, step S202 may include but is not limited to steps S301 to S303:
[0083] Step S301, perform semantic conversion on the alternative picked-up audio to obtain alternative semantic data;
[0084] Step S302, perform intent recognition on the alternative semantic data to obtain alternative intent data;
[0085] Step S303, perform human voice target detection on the alternative picked-up audio according to the alternative intent data to obtain human voice target detection data.
[0086] Steps S301 to S303 illustrated in the embodiments of the present application perform semantic conversion on the alternative picked-up audio to obtain alternative semantic data, then perform intent recognition on the alternative semantic data to obtain alternative intent data, and finally perform human voice target detection on the alternative picked-up audio according to the alternative intent data to obtain human voice target detection data, so as to achieve dynamic directional sound pickup according to semantic intent from human voice sound sources in all directions, screen out human voice sound sources that interact with the headphone wearer, and thus improve the dynamic sound pickup effect of the headphones.
[0087] In step S301 of some embodiments, the semantic conversion is to convert the alternative picked-up audio from audio to text, and use the converted text as the alternative semantic data. In some embodiments, converting audio to text, that is, performing semantic conversion, can use automatic speech recognition (ASR) based on deep learning, or a conversion method combining a hidden Markov model (HMM) and a Gaussian mixture model (GMM) to achieve the conversion of audio to text.
[0088] In step S302 of some embodiments, intent recognition is performed on the alternative semantic data, and it is determined whether the text in the alternative semantic data is directed at a signal sent by the headphone wearer or not. In one embodiment, the alternative semantic data is input into a pre-trained natural language model, and a preset prompt word and the alternative semantic data are input into the pre-trained natural language model, so as to obtain a value from 0 to 1, that is, alternative intent data. For example, the alternative semantic data is "The bus in front is about to arrive at the station. Please prepare to get off, passengers", and the alternative semantic data is input into the pre-trained natural language model, and the natural language model outputs 0.732, that is, the alternative intent data.
[0089] It should be noted that since the microphone positions of the headphones are located at different positions, the audio collected by the microphones is not the same, that is, the alternative intent data of each alternative picked-up audio is different.
[0090] In step S303 of some embodiments, voice target detection is performed on the alternative picked-up audio according to the alternative intent data. Specifically, the alternative intent data is compared with a preset intent threshold. When the alternative intent data is greater than the preset intent threshold, the voice target detection data is determined to be 1. When the alternative intent data is less than the preset intent threshold, the voice detection data is determined to be 0. The voice target detection data being 1 means that the target of the sound source is the headphone wearer, and the voice target detection data being 0 means that the target of the sound source is not the headphone wearer.
[0091] In step S203 of some embodiments, the alternative picked-up audio is screened according to the environmental audio evaluation data and the voice target detection data. Specifically, when the environmental audio evaluation data is greater than a preset emergency situation threshold, the sound source corresponding to the environmental audio evaluation data is determined as the main picked-up audio. When the environmental audio evaluation data is less than the preset emergency situation threshold, the sound source corresponding to 1 in the voice target detection data is obtained as the main picked-up audio.
[0092] In step S204 of some embodiments, the alternative picked-up audio is screened according to the main picked-up audio, and the audio in the alternative picked-up audio that is not the main picked-up audio is obtained to get the secondary picked-up audio.
[0093] Please refer to Figure 4 , in some embodiments, step S103 may include but is not limited to steps S401 to S403:
[0094] Step S401, performing voice extraction on the secondary picked-up audio according to the main picked-up audio to obtain a target enhanced audio;
[0095] Step S402, performing echo cancellation on the target enhanced audio to obtain a selected enhanced audio;
[0096] Step S403: Perform audio merging on the main picked-up audio according to the selected enhanced audio to obtain the target playback audio.
[0097] Steps S401 to S403 illustrated in the embodiments of the present application are as follows. First, perform voice extraction on the secondary picked-up audio according to the content of the main picked-up audio to extract a clearer and more accurate target enhanced audio. Subsequently, perform echo cancellation processing on the target enhanced audio to eliminate possible environmental echoes and interference sounds, generating a cleaner selected enhanced audio. On this basis, perform audio merging on the main picked-up audio using the selected enhanced audio, and achieve content supplementation and quality improvement by integrating the information of the two audio channels. Finally, the generated target playback audio not only retains the core information of the main picked-up audio but also integrates the effective supplementary information in the secondary picked-up audio, ensuring the clarity and integrity of the target playback audio.
[0098] Please refer to Figure 5 , in some embodiments, step S401 includes but is not limited to steps S501 to S503:
[0099] Step S501: Perform frequency-domain conversion on the main picked-up audio to obtain main frequency-domain data, and perform frequency-domain conversion on the secondary picked-up audio to obtain secondary frequency-domain data;
[0100] Step S502: Perform voice frequency-domain screening on the secondary frequency-domain data according to the main frequency-domain data to obtain target voice frequency-domain data;
[0101] Step S503: Perform inverse frequency-domain conversion on the target voice frequency-domain data to obtain the target enhanced audio.
[0102] Steps S501 to S503 illustrated in the embodiments of the present application can achieve precise extraction and significant enhancement of the target voice by performing frequency-domain conversion, screening, and inverse conversion processing on the main picked-up audio and the secondary picked-up audio, effectively removing noise and irrelevant information in the secondary picked-up audio, and at the same time supplementing possible missing details in the main picked-up audio, making the finally generated target enhanced audio clearer, more complete, and of higher quality.
[0103] In step S501 of some embodiments, frequency-domain conversion refers to transforming an audio signal from a time-domain representation to a frequency-domain representation for analyzing its energy distribution at different frequencies. Frequency-domain data refers to the distribution information of an audio signal in the frequency domain, usually including the amplitude and phase of each frequency component, and is used to describe the spectral characteristics of the audio signal. In some embodiments, the frequency-domain conversion can be a fast Fourier transform (FFT), which can quickly decompose a time-domain signal into its frequency-domain representation. It can also be a short-time Fourier transform (STFT), which provides joint distribution information of frequency and time by segmenting the audio signal.
[0104] Please refer to Figure 6, in some embodiments, step S502 includes but is not limited to steps S601 to S605:
[0105] Step S601, extracting common frequency domain features of the secondary frequency domain data based on the main frequency domain data to obtain common frequency domain features of human voices;
[0106] Step S602, performing beamforming processing on the common frequency domain features of human voices to obtain beam frequency domain features;
[0107] Step S603, performing coherence analysis processing on the common frequency domain features of human voices to obtain coherence frequency domain features;
[0108] Step S604, performing adaptive filtering on the common frequency domain features of human voices to obtain filtered frequency domain features;
[0109] Step S605, performing aggregation calculation based on the beam frequency domain features, coherence frequency domain features, and filtered frequency domain features to obtain target human voice frequency domain data.
[0110] Steps S601 to S605 illustrated in the embodiments of the present application, through the extraction of common frequency domain features, effectively identify similar human voice features in the main pick-up audio and the secondary pick-up audio, reduce the influence of environmental noise and other interferences, and then further strengthen the directivity of the human voice signal through beamforming processing, making the extracted human voice clearer and more concentrated. Then, through coherence analysis processing, the coherence characteristics between signals are utilized to enhance the relevant parts of the signals, thereby further improving the quality of the human voice signal. Next, through adaptive filtering, the filtering parameters can be adjusted according to the real-time signal characteristics, accurately removing residual noise and retaining the target human voice information. Finally, through the aggregation calculation of the beam frequency domain features, coherence frequency domain features, and filtered frequency domain features, the advantages of multiple features are comprehensively utilized, making the generated target human voice frequency domain data have higher clarity, integrity, and thus significantly improving the effect of audio processing.
[0111] In step S601 of some embodiments, the common frequency-domain feature extraction refers to separating the frequency information with common features from the secondary frequency-domain data compared to the primary frequency-domain data, which is used to extract the similar vocal frequency components in the two sets of data and exclude irrelevant noise or background signals. The vocal common frequency-domain feature refers to the spectral information shared by the primary and secondary frequency-domain data and related to the human voice, which is the significant signal feature within a certain frequency range. In some embodiments, the common frequency-domain feature extraction can be a method based on similarity measurement, such as identifying the similar frequency components in the primary and secondary audio signals through frequency-domain correlation analysis. In other embodiments, a machine learning model is used, for example, adopting spectral clustering technology to cluster the common features in the frequency-domain data into the target frequency range. For example, in a street scenario, the primary and secondary frequency-domain data can be obtained from the audio recorded by the primary and secondary microphones respectively. Then, through frequency-domain correlation analysis, the common vocal spectrum of the two audio channels can be extracted to obtain the clear vocal common frequency-domain feature.
[0112] In step S602 of some embodiments, the beamforming process refers to adjusting the amplitude and phase of the audio signals collected by multiple microphones to enhance the sound signal in a specific direction while suppressing the interference in other directions, which is used for the vocal signal in the target direction, so as to extract clearer and more accurate audio features. The beam frequency-domain feature is the feature information of the vocal signal in the frequency domain after the directional enhancement and interference suppression processing. In some embodiments, the beamforming process is delay-and-sum beamforming, which enhances the signal from a specific direction by time-aligning and adding the signals of multiple microphones.
[0113] For example, in a remote meeting, beamforming processing is performed on the vocal common frequency-domain feature to focus on the vocal signal in the direction of the speaker, while suppressing the background noise and interference signals in other directions, and finally the beam frequency-domain feature is obtained.
[0114] In step S603 of some embodiments, the coherence analysis process refers to analyzing the coherence between signals, that is, the correlation degree or phase consistency between the frequency-domain data, to identify and strengthen the frequency components with higher correlation, which is used to highlight the parts with higher correlation in the signal, so as to enhance the frequency-domain features related to the target vocal. The coherent frequency-domain feature is a kind of frequency-domain component, which represents the high-correlation frequency components extracted after the coherence analysis in the frequency domain. The coherent frequency-domain feature can more accurately reflect the main features of the target vocal. In some embodiments, the coherence analysis process can be a coherence function analysis, which calculates the coherence function between signals to quantify the phase and amplitude correlation at different frequencies, so as to extract the parts with higher correlation.
[0115] In step S604 of some embodiments, adaptive filtering refers to dynamically adjusting the parameters of a filter to adapt to the changes in signals and noise in real time, thereby optimizing the extraction of the human voice signal and the suppression of noise, enhancing the human voice signal in the target frequency band while reducing interference and background noise. The filtered frequency-domain features are the features of the human voice signal dynamically optimized in the frequency domain. In one embodiment, the adaptive filtering can be the LMS (Least Mean Squares) algorithm, which dynamically adjusts the filter weights by minimizing the mean square value of the error signal to achieve precise processing of the signal.
[0116] In step S605 of some embodiments, the aggregation calculation is performed based on the beam frequency-domain features, the coherence frequency-domain features, and the filtered frequency-domain features to obtain the human voice in the frequency domain form. The aggregation calculation is shown in Equation (1):
[0117] Y(t,f) = λ 1 Y beam (t,f) + λ 2 Y coh (t,f) + λ 3 Y adapt (t,f) (1),
[0118] where Y(t,f) is the target human voice frequency-domain data, Y beam (t,f) is the beam frequency-domain feature, Y coh (t,f) is the coherence frequency-domain feature, Y adapt (t,f) is the filtered frequency-domain feature, λ 1 is the weighted component of the beam frequency-domain feature, λ 2 is the weighted component of the coherence frequency-domain feature, λ 3 is the weighted component of the filtered frequency-domain feature.
[0119] In step S503 of some embodiments, an inverse frequency-domain conversion is performed on the target human voice frequency-domain data, that is, the target human voice frequency-domain data is converted from the frequency-domain space to the time-domain space, corresponding to the frequency-domain conversion of the main picked-up audio. For example, if the main picked-up audio is subjected to a Fourier transform to obtain the main frequency-domain data, then the target human voice frequency-domain data is subjected to an inverse Fourier transform to obtain the target enhanced audio.
[0120] In step S402 of some embodiments, echo cancellation is to remove the echo components caused by the device itself or environmental reflections from the audio signal, thereby improving the clarity and quality of the audio. The selected enhanced audio refers to the audio signal after echo cancellation processing, which has higher clarity and less echo interference. In one embodiment, the echo cancellation is based on frequency-domain echo cancellation, which uses spectral analysis methods to identify and suppress the echo frequency band to achieve echo cancellation.
[0121] In step S403 of some embodiments, audio merging refers to fusing two audio signals into one audio, while simultaneously retaining the valid information of the selected enhanced audio and the main picked-up audio. This is used to optimize the overall audio quality and information integrity while strengthening the target audio signal. The target playback audio is the result of the merging, which is an audio output that contains both the core content of the main picked-up audio and supplements the improved details in the selected enhanced audio. In one embodiment, audio merging is performed on the selected enhanced audio and the main picked-up audio based on the weighted superposition method. By assigning weights to the two audio signals and adding them proportionally, a fused signal is generated, which is the target playback audio.
[0122] In step S104 of some embodiments, an audio playback program is preset in the earphone, and the target playback audio is played based on the audio playback program.
[0123] Please refer to Figure 7 , in some embodiments, the earphone includes a three-axis accelerometer and a three-axis gyroscope. After step S104, the dynamic sound pickup method may further include, but is not limited to, steps S701 to S704:
[0124] Step S701, obtaining the sound source direction of the main picked-up audio to obtain main sound source direction data;
[0125] Step S702, performing head pose calculation based on the three-axis accelerometer and the three-axis gyroscope to obtain head pose data;
[0126] Step S703, performing direction adjustment according to the main sound source direction data and the head pose data to obtain target sound source direction data;
[0127] Step S704, performing beamforming direction adjustment on all alternative microphones according to the target sound source direction data.
[0128] Steps S701 to S704 illustrated in the embodiments of the present application can, through dynamic adjustment by combining the main sound source direction data and the head pose data, track the position of the speaker or the target sound source in real time in a complex environment, ensure the accuracy of the sound pickup direction, and at the same time, the adjustment of the beamforming direction further enhances the clarity of the target sound source and effectively suppresses noise and interference from other directions.
[0129] In step S701 of some embodiments, obtaining the sound source direction of the main picked-up audio refers to determining the specific direction of the sound source. The main sound source direction data is the sound source direction information of the main picked-up audio, which is data presented in the form of three-dimensional coordinates.
[0130] In step S702 of some embodiments, the head pose calculation is to calculate the movement direction of the head of the headphone wearer. After the headphone plays the target playback audio, the head of the headphone wearer may make adjustments in directions such as rotation and swing. Dynamically obtain the accelerometer data of the triaxial accelerometer, and then obtain the gyroscope data of the triaxial gyroscope. Perform head pose calculation based on the accelerometer data and the gyroscope data, and obtain the rotation direction of the head pose in real time, that is, obtain the head pose data.
[0131] In step S703 of some embodiments, perform direction adjustment according to the head pose data and the main sound source direction data to obtain the target sound source direction data. For example, the head pose data is represented as the head turning ninety degrees to the left, and the sound source direction is determined to be directly in front of the body according to the main sound source direction. Finally, determine the target sound source direction data according to the sound source direction and the head pose data.
[0132] In step S704 of some embodiments, the beamforming direction adjustment is to use the direction information of the target sound source to dynamically adjust the beamforming parameters of all alternative microphones in the headphones to optimize the capture effect of sounds in a specific direction. In one embodiment, the beamforming direction adjustment specifically is to calculate the relative position of each microphone with respect to the sound source and the required time delay according to the target sound source direction data, and then apply the corresponding time delay to the audio signal collected by each microphone, so that the sounds from the target direction are phase-aligned at all microphones. Finally, superimpose the delayed signals to enhance the sound signal from the target direction and suppress the noise in other directions. For example, when the target sound source is in front of the headphones, calculate and apply an appropriate delay so that the front sounds received by all microphones are synchronously superimposed to enhance the clarity of the front human voice.
[0133] Please refer to Figure 8 , the embodiment of the present application further provides a dynamic sound pickup device, which is applied to headphones. The headphones include at least one alternative microphone, and the above dynamic sound pickup method can be implemented. The device includes:
[0134] A data acquisition module 801, configured to acquire the alternative pickup audio of all alternative microphones;
[0135] An audio screening module 802, configured to screen the alternative pickup audio to obtain the main pickup audio and the secondary pickup audio;
[0136] An audio synthesis module 803, configured to synthesize the main pickup audio and the secondary pickup audio to obtain the target playback audio;
[0137] An audio playback module 804, configured to perform audio playback of the target playback audio based on the headphones.
[0138] The specific implementation manner of this dynamic sound pickup device is basically the same as the specific embodiment of the above dynamic sound pickup method, and will not be elaborated herein.
[0139] An embodiment of this application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above dynamic sound pickup method is implemented. This electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0140] Please refer to Figure 9 , Figure 9 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0141] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application;
[0142] A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the dynamic sound pickup method of the embodiments of this application;
[0143] An input / output interface 903, which is used to implement information input and output;
[0144] A communication interface 904, which is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0145] A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);
[0146] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.
[0147] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above dynamic sound pickup method is implemented.
[0148] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0149] The dynamic sound pickup method, dynamic sound pickup device, electronic device, and storage medium provided by the embodiments of the present application first obtain the alternative sound pickup audio of all alternative microphones on the earphone, and then perform audio screening on the alternative sound pickup audio to obtain the main sound pickup audio and the secondary sound pickup audio, so as to dynamically determine the sound pickup direction of the earphone when the wearer is wearing the earphone, thereby accurately positioning the sound source and improving the sound pickup effect of the earphone; further, perform audio synthesis on the main sound pickup audio and the secondary sound pickup audio, so as to extract the effective information in the secondary sound pickup audio and synthesize it into the main sound pickup audio to achieve the extraction of effective information; finally, play the target audio based on the earphone to achieve dynamic directional sound pickup of sound sources from all directions in the real environment, effectively screen the information from all directions, and ultimately improve the sound pickup effect of the earphone.
[0150] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0151] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0153] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0154] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0155] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (individual) of the following" or a similar expression means any combination of these items, including any combination of single items (individuals) or plural items (individuals). For example, at least one (individual) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.
[0157] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0158] In addition, the functional units in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0159] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.
[0160] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A dynamic sound pickup method, characterized in that: Applied to a headset, the headset comprising at least one alternative microphone, the method comprising: Acquire alternative pickup audio of all the alternative microphones; Performing audio screening on the candidate sound pickup audio to obtain a main sound pickup audio and a secondary sound pickup audio; Performing audio synthesis on the main sound pickup audio and the auxiliary sound pickup audio to obtain target playback audio; The target playback audio is played based on the earphone.
2. The method according to claim 1, characterized in that The step of performing audio screening on the candidate sound pickup audio to obtain the main sound pickup audio and the auxiliary sound pickup audio comprises: Performing environmental audio evaluation on the candidate picked-up audio to obtain environmental audio evaluation data; Performing human voice target detection on the candidate picked-up audio to obtain human voice target detection data; Performing audio screening on the candidate sound pickup audio according to the environmental audio evaluation data and the human voice target detection data to obtain the main sound pickup audio; The candidate sound pickup audio is audio-filtered according to the main sound pickup audio to obtain the secondary sound pickup audio.
3. The method according to claim 2, characterized in that The performing human voice target detection on the candidate picked-up audio to obtain human voice target detection data includes: Performing semantic conversion on the candidate picked-up audio to obtain candidate semantic data; Performing intent recognition on the candidate semantic data to obtain candidate intent data; Perform human voice target detection on the alternative sound pickup audio according to the alternative intention data to obtain the human voice target detection data.
4. The method according to claim 1, characterized in that The step of synthesizing the main sound pickup audio and the auxiliary sound pickup audio to obtain the target playback audio includes: Extracting human voice from the secondary sound pickup audio according to the primary sound pickup audio to obtain target enhanced audio; Performing echo cancellation on the target enhanced audio to obtain selected enhanced audio; The main sound pickup audio is audio-merged according to the selected enhanced audio to obtain the target playback audio.
5. The method according to claim 4, characterized in that The extracting human voice from the auxiliary sound pickup audio according to the main sound pickup audio to obtain target enhanced audio includes: Performing frequency domain conversion on the main sound pickup audio to obtain main frequency domain data, and performing frequency domain conversion on the secondary sound pickup audio to obtain secondary frequency domain data; Performing human voice frequency domain screening on the secondary frequency domain data according to the primary frequency domain data to obtain target human voice frequency domain data; Perform frequency domain inverse conversion on the target human voice frequency domain data to obtain the target enhanced audio.
6. The method according to claim 5, characterized in that The performing human voice frequency domain screening on the secondary frequency domain data according to the primary frequency domain data to obtain target human voice frequency domain data includes: Extracting common frequency domain features from the secondary frequency domain data according to the primary frequency domain data to obtain common frequency domain features of human voice; Performing beamforming processing on the common frequency domain features of human voice to obtain beam frequency domain features; Performing coherent analysis on the common frequency domain features of human voice to obtain coherent frequency domain features; Adaptively filtering the common frequency domain features of human voice to obtain filtered frequency domain features; Aggregate calculation is performed according to the beam frequency domain features, the coherent frequency domain features and the filter frequency domain features to obtain the target human voice frequency domain data.
7. The method according to claim 1, characterized in that The earphone includes a three-axis accelerometer and a three-axis gyroscope, and after the target playback audio is played based on the earphone, the method further includes: Obtaining the sound source direction of the main picked-up audio to obtain main sound source direction data; Calculate the head posture based on the three-axis accelerometer and the three-axis gyroscope to obtain head posture data; Performing direction adjustment according to the main sound source direction data and the head posture data to obtain target sound source direction data; The beamforming directions of all the candidate microphones are adjusted according to the target sound source direction data.
8. A dynamic sound pickup device, characterized in that: Applied to a headset, the headset comprising at least one optional microphone, the device comprising: A data acquisition module, used for acquiring the alternative sound pickup audio of all the alternative microphones; An audio screening module, used for screening the candidate sound pickup audio to obtain a main sound pickup audio and a secondary sound pickup audio; An audio synthesis module, used for synthesizing the main sound pickup audio and the auxiliary sound pickup audio to obtain a target playback audio; An audio playing module is used to play the target audio based on the earphone.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the dynamic sound pickup method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the dynamic sound pickup method according to any one of claims 1 to 7 is implemented.