Pickup method and device, electronic equipment and program product
By acquiring the raw audio information of the microphone array in the electronic device and combining it with voiceprint information and/or location information for enhancement and suppression processing, the problem of inaccurate audio information extraction in electronic devices is solved, and accurate audio information extraction is achieved.
Patent Information
- Application Number
- CN202410869230.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-12-30
AI Technical Summary
Existing electronic devices are unable to accurately extract the desired audio information during audio data acquisition, resulting in inaccurate audio information.
By acquiring the raw audio information collected by the microphone array, the prior information, including voiceprint information and/or location information, is determined. This information is then used to enhance the raw audio information and suppress irrelevant information. A spatial pickup model is used to enhance and suppress the audio information.
The accuracy of audio information extraction has been improved by combining voiceprint information and/or location information to enhance and suppress audio information, thereby increasing the accuracy of the final extracted audio information.
Smart Images

Figure CN121237116A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic equipment technology, and more specifically, to a sound pickup method, apparatus, electronic equipment, and software product. Background Technology
[0002] With the development of science and technology, electronic devices are becoming increasingly widespread and multifunctional, becoming an essential part of people's daily lives. Currently, electronic devices can be used to collect audio information and extract desired audio information from the collected audio information; however, the extracted audio information often suffers from inaccuracy. Summary of the Invention
[0003] In view of the above problems, this application proposes a sound pickup method, device, electronic device, and program product to solve the above problems.
[0004] In a first aspect, embodiments of this application provide a sound pickup method applied to an electronic device. The method includes: acquiring raw audio information collected by a microphone array of the electronic device; determining prior information for the raw audio information, wherein the prior information includes at least one of voiceprint information and location information; enhancing audio information in the raw audio information that is associated with the prior information, and suppressing audio information in the audio information that is not associated with the prior information, thereby obtaining target audio information.
[0005] Secondly, embodiments of this application provide a sound pickup device applied to an electronic device. The device includes: a raw audio information acquisition module for acquiring raw audio information collected by a microphone array of the electronic device; a priori information determination module for determining priori information for the raw audio information, wherein the priori information includes at least one of voiceprint information and location information; and a target audio information acquisition module for enhancing audio information in the raw audio information that is associated with the priori information and suppressing audio information in the audio information that is not associated with the prior information, thereby obtaining target audio information.
[0006] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory is coupled to the processor, the memory stores instructions, and when the instructions are executed by the processor, the processor performs the above-described method.
[0007] Fourthly, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method.
[0008] The sound pickup method, apparatus, electronic device, and program product provided in this application acquire raw audio information collected by a microphone array of an electronic device, determine prior information for the raw audio information, wherein the prior information includes at least one of voiceprint information and location information, enhance audio information in the raw audio information that is associated with the prior information, and suppress audio information in the audio information that is not associated with the prior information, thereby obtaining target audio information. By combining prior information including voiceprint information and / or location information to enhance a part of the acquired audio information and suppress another part, the accuracy of the finally extracted audio information can be improved. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A schematic flowchart of a sound pickup method provided in an embodiment of this application is shown;
[0011] Figure 2 A schematic flowchart of a sound pickup method provided in an embodiment of this application is shown;
[0012] Figure 3 This illustration shows a flowchart of obtaining target audio information from raw audio information using a spatial pickup model, as provided in an embodiment of this application.
[0013] Figure 4 A schematic flowchart of a sound pickup method provided in an embodiment of this application is shown;
[0014] Figure 5 A schematic flowchart of a sound pickup method provided in an embodiment of this application is shown;
[0015] Figure 6 A schematic flowchart of a sound pickup method provided in an embodiment of this application is shown;
[0016] Figure 7 A schematic flowchart of a sound pickup method provided in an embodiment of this application is shown;
[0017] Figure 8 A schematic diagram of a sound pickup method provided in an embodiment of this application is shown;
[0018] Figure 9 This illustration shows another sound pickup diagram of the sound pickup method provided in the embodiments of this application;
[0019] Figure 10 A block diagram of the pickup device provided in an embodiment of this application is shown;
[0020] Figure 11 A block diagram of an electronic device for performing a sound pickup method according to an embodiment of this application is shown;
[0021] Figure 12 A storage unit for storing or carrying program code implementing the sound pickup method according to an embodiment of the present application is shown. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0023] In current scenarios where electronic devices collect audio information, the desired audio information cannot be accurately extracted from the collected audio information, resulting in inaccurate extracted audio information. Through long-term research, the inventors discovered and proposed a sound pickup method, device, electronic device, and program product according to an embodiment of this application. By combining prior information, including voiceprint information and / or location information, to enhance a portion of the collected audio information and suppress another portion, the accuracy of the finally extracted audio information can be improved. The specific sound pickup method will be described in detail in subsequent embodiments.
[0024] Please see Figure 1 , Figure 1 A schematic flowchart of a sound pickup method according to an embodiment of this application is shown. This sound pickup method enhances a portion of the acquired audio information and suppresses another portion by combining prior information including voiceprint information and / or location information, thereby improving the accuracy of the finally extracted audio information. In a specific embodiment, this sound pickup method is applied to, for example... Figure 10 The illustrated microphone 200 and the electronic device 100 equipped with the microphone 200 are shown. Figure 11 The following will use an electronic device as an example to illustrate the specific process of this embodiment. Of course, it is understood that the electronic device used in this embodiment may include smartphones, tablets, wearable electronic devices, etc., and is not limited thereto. The following will focus on... Figure 1 The process shown will be explained in detail. The sound pickup method may specifically include the following steps:
[0025] Step S110: Acquire raw audio information captured by the microphone array of the electronic device.
[0026] Optionally, the electronic device may be equipped with a microphone array. For example, the electronic device may be equipped with a microphone array consisting of one top microphone and two bottom microphones. Of course, the microphone array may also include other microphones, which is not limited here.
[0027] In this embodiment, when the electronic device has a need for audio acquisition, audio information can be acquired through the microphone array of the electronic device, and the audio information acquired through the microphone array of the electronic device can be used as the original audio information.
[0028] In some implementations, when the electronic device has a recording / video recording requirement, it can be determined that the electronic device has an audio acquisition requirement, and audio information can be acquired through the microphone array of the electronic device; when the electronic device has a call requirement, it can be determined that the electronic device has an audio acquisition requirement, and audio information can be acquired through the microphone array of the electronic device; when the electronic device has a voice transmission requirement, it can be determined that the electronic device has an audio acquisition requirement, and audio information can be acquired through the microphone array of the electronic device, etc., without limitation.
[0029] Step S120: Determine prior information for the original audio information, wherein the prior information includes at least one of voiceprint information and location information.
[0030] Prior information can be used as a reference factor to extract the desired audio information from the original audio information. That is, prior information can be used to extract audio information that matches or is associated with the prior information from the original speech information.
[0031] Optionally, the prior information may include at least one of voiceprint information and location information. That is, the prior information may include only voiceprint information, only location information, both voiceprint information and location information, voiceprint information and other information, location information and other information, or voiceprint information, location information and other information, etc., without limitation.
[0032] In this embodiment, prior information for the original audio information can be determined. Optionally, this prior information may be automatically determined by the electronic device based on the current usage scenario, automatically determined based on historical usage data, or determined based on the user's input selection operation; no limitation is made here.
[0033] As an example, if the prior information is automatically determined by the electronic device based on the current usage scenario, then the prior information can be determined based on the current orientation of the electronic device's screen, such as determining the current orientation of the electronic device's screen as the directional information for the original audio information; it can be determined based on the current direction of the electronic device, such as determining the current direction of the electronic device as the directional information for the original audio information; it can be determined based on the original audio information currently collected by the electronic device, such as determining the voiceprint information with the highest frequency in the original audio information as the voiceprint information for the original audio information, or determining the voiceprint information corresponding to the loudest audio segment in the original audio information as the voiceprint information for the original audio information, etc., without limitation.
[0034] As another example, if the prior information is automatically determined by the electronic device based on historical usage data, then the prior information can be determined based on the historical location information of the electronic device, such as determining the location information of the most recently set location of the electronic device as the location information for the original audio information, or determining the location information of the electronic device most frequently set in a recent period as the location information for the original audio information; or it can be determined based on the historical voiceprint information of the electronic device, such as determining the voiceprint information of the most recently set voiceprint of the electronic device as the voiceprint information for the original audio information, or determining the voiceprint information of the electronic device most frequently set in a recent period as the voiceprint information for the original audio information, etc., without limitation.
[0035] As an feasible approach, in scenarios where the desired audio information is extracted from the original audio information, the prior information can be fixed after being set. Therefore, given the prior information for the original audio information, the desired audio information can always be extracted from the original audio information based on that prior information.
[0036] As another feasible approach, in scenarios where the desired audio information is extracted from the original audio information, the prior information can be dynamically changed after being set. Therefore, given the prior information for the original audio information, the prior information can be updated as needed, and the desired audio information can be extracted from the original audio information based on the updated prior information.
[0037] Step S130: Enhance the audio information in the original audio information that is associated with the prior information, and suppress the audio information in the audio information that is not associated with the prior information to obtain the target audio information.
[0038] In this embodiment, given the original audio information and prior information, the audio information in the original audio information that is related to the prior information can be enhanced, and the audio information in the audio information that is not related to the prior information can be suppressed to obtain the target audio information.
[0039] In some implementations, given the original audio information and prior information, audio information associated with the prior information can be extracted from the original audio information, and other audio information in the original audio information besides that associated with the prior information can be identified as audio information not associated with the prior information. Then, the audio information in the original audio information associated with the prior information can be enhanced, and the audio information in the original audio information not associated with the prior information can be suppressed to obtain the target audio information.
[0040] As an feasible approach, when the prior information is voiceprint information, given the original audio information and voiceprint information, audio information associated with the voiceprint information (such as audio information matching the corresponding voiceprint information) can be extracted from the original audio information. Other audio information in the original audio information, excluding those associated with the voiceprint information, is then identified as audio information not associated with the voiceprint information. Subsequently, the audio information associated with the voiceprint information in the original audio information can be enhanced, while the audio information not associated with the voiceprint information can be suppressed to obtain the target audio information.
[0041] As another feasible approach, when the prior information is directional information, given the original audio information and directional information, audio information related to the directional information (such as the corresponding directional location being within the range indicated by the directional information) can be extracted from the original audio information. Other audio information in the original audio information besides that related to the directional information can be identified as audio information not related to the directional information. Then, the audio information related to the directional information in the original audio information can be enhanced, and the audio information not related to the directional information can be suppressed to obtain the target audio information.
[0042] As another feasible approach, when the prior information consists of voiceprint information and location information, then, given the original audio information and location information, audio information that is simultaneously associated with both voiceprint and location information can be extracted from the original audio information. Other audio information in the original audio information, excluding that associated with both voiceprint and location information, can be identified as audio information not associated with location information. Subsequently, the audio information associated with both voiceprint and location information can be enhanced, while the audio information not associated with either can be suppressed to obtain the target audio information.
[0043] As an example, let's consider prior information including voiceprint and location information. Given the original audio information, voiceprint information, and location information, voiceprint recognition technology can be used to locate and identify the voice of the target speaker in the original audio information. This may involve extracting sound features, such as MFCCs (Melbourne Frequency Cepstral Coefficients), and classifying them using machine learning models (such as GMM, DNN, RNN, etc.). Once the target voiceprint corresponding to the voiceprint information is identified, specific filters or dynamic processors (such as multi-band compressors) can be used to enhance its characteristic frequencies, making it clearer. Beamforming technology can be used to focus the original audio information at the target location corresponding to the location information, while suppressing noise from other directions. Beamforming can be implemented using various algorithms, such as delay and superposition, minimum variance distortion unconstrained (MVDR), LMS adaptive filtering, etc.
[0044] One embodiment of this application provides a sound pickup method that acquires raw audio information collected by a microphone array of an electronic device, determines prior information for the raw audio information, wherein the prior information includes at least one of voiceprint information and location information, enhances audio information in the raw audio information that is associated with the prior information, and suppresses audio information in the audio information that is not associated with the prior information, thereby obtaining target audio information. By combining prior information including voiceprint information and / or location information to enhance a portion of the acquired audio information and suppress another portion, the accuracy of the finally extracted audio information can be improved.
[0045] Please see Figure 2 , Figure 2 A schematic flowchart of a sound pickup method according to an embodiment of this application is shown. This method is applied to the aforementioned electronic device, and will be discussed below. Figure 2 The process shown will be explained in detail. The sound pickup method may specifically include the following steps:
[0046] Step S210: Acquire raw audio information captured by the microphone array of the electronic device.
[0047] Step S220: Determine prior information for the original audio information, wherein the prior information includes at least one of voiceprint information and location information.
[0048] For a detailed description of steps S210-S220, please refer to steps S110-S120, which will not be repeated here.
[0049] Step S230: Perform a short-time Fourier transform on the original audio information to obtain the original time-frequency signal.
[0050] Please see Figure 3 , Figure 3 This illustration shows a flowchart of obtaining target audio information from raw audio information using a spatial pickup model, as provided in an embodiment of this application.
[0051] In this embodiment, given the original audio information, a time-frequency transformation, i.e., a short-time Fourier transform, can be performed on the original audio information to obtain the original time-frequency signal. Since the original audio information is composed of microphone signals collected by multiple microphones in the microphone array of the electronic device, the microphone signals collected by the multiple microphones can be time-frequency transformed to obtain a multi-channel time-frequency signal as the original time-frequency signal.
[0052] The multi-channel time-frequency signal can be represented as follows:
[0053] Y=[Y1(f,t),Y2(f,t),Y3(f,t)]
[0054] The spatial and spectral characteristics of the original time-frequency signal are defined as follows:
[0055]
[0056] Step S240: Input the original time-frequency signal and the prior information into the spatial sound pickup model to obtain the target time-frequency signal output by the spatial sound pickup model. The target time-frequency signal is obtained by enhancing the time-frequency signals in the original time-frequency signal that are related to the prior information and suppressing the time-frequency signals in the original time-frequency signal that are not related to the prior information through the spatial sound pickup model.
[0057] In some implementations, a spatial sound pickup model can be pre-set. This model can enhance time-frequency signals in the time-frequency signal that are associated with predetermined prior information, and suppress time-frequency signals in the video signal that are not associated with the predetermined prior information. Therefore, in this embodiment, when the original time-frequency signal is obtained, the original time-frequency signal and the prior information can be input into the spatial sound pickup model. The spatial sound pickup model enhances the time-frequency signals in the original time-frequency signal that are associated with the prior information and suppresses the time-frequency signals in the original time-frequency signal that are not associated with the prior information to obtain and output the target time-frequency signal. Accordingly, the electronic device can obtain the target time-frequency signal output by the spatial sound pickup model.
[0058] The spatial sound pickup model can be obtained through machine learning. Specifically, a training dataset is first collected, in which one type of data has attributes or features that distinguish it from another type of data. Then, the collected training dataset is used to train a neural network according to a preset algorithm, thereby summarizing patterns based on the training dataset to obtain the spatial sound pickup model. In this embodiment, the input parameters of the training dataset may include multiple first audio information and multiple prior information, and the output parameters may include multiple corresponding second audio information.
[0059] Understandably, the spatial sound pickup model can be pre-trained and stored locally on the electronic device. Based on this, after acquiring the raw time-frequency signal and prior information, the electronic device can directly access the spatial sound pickup model locally. For example, it can directly send instructions to the spatial sound pickup model to instruct it to read the raw time-frequency signal and prior information from the target storage area. Alternatively, the electronic device can directly input the raw time-frequency signal and prior information into the locally stored spatial sound pickup model. This effectively avoids the slowdown in inputting the raw time-frequency signal and prior information into the spatial sound pickup model due to network factors, thereby improving the speed at which the spatial sound pickup model acquires the raw time-frequency signal and prior information and enhancing the user experience.
[0060] Furthermore, the spatial sound pickup model can be pre-trained and stored on a server connected to the electronic device. Based on this, after acquiring the raw time-frequency signal and prior information, the electronic device can send instructions via the network to the spatial sound pickup model stored on the server, instructing the model to read the raw time-frequency signal and prior information obtained by the electronic device through the network. Alternatively, the electronic device can send the raw time-frequency signal and prior information to the spatial sound pickup model stored on the server via the network. By storing the spatial sound pickup model on the server, the storage space occupied by the electronic device is reduced, minimizing the impact on its normal operation.
[0061] Optionally, the spatial sound pickup model includes an encoder, a recurrent neural network (RNN), and a decoder. The encoder can employ a multi-layer stacked structure, with each layer being a Conv2d+PReLU+Batch Norm model operator. The RNN can be a multi-layered GRU queue. The decoder can also employ a multi-layer stacked structure, with each layer being a TransposedConv2d+PReLU+Batch Norm model operator.
[0062] In some implementations, inputting the original time-frequency signal and prior information into a spatial sound pickup model to obtain the target time-frequency signal output by the spatial sound pickup model may include: inputting the original time-frequency signal into an encoder for processing to obtain a semantic vector corresponding to the original time-frequency signal output by the encoder; inputting the prior information and semantic vector into a recurrent neural network for processing to obtain intermediate data output by the recurrent neural network; inputting the intermediate data into a decoder for processing to obtain a time-frequency masking coefficient output by the decoder; enhancing the time-frequency signals in the original time-frequency signal that are related to the prior information based on the time-frequency masking coefficient, and suppressing the time-frequency signals in the original time-frequency signal that are not related to the prior information to obtain the target time-frequency signal.
[0063] The Encoder's input is the microphone signal after time-frequency transformation, i.e., the multi-channel time-frequency signal Y = [Y1(f,t), Y2(f,t), Y3(f,t)], resulting from the short-time Fourier transform. The spatial and spectral characteristics are defined as follows: The decoder outputs multi-channel time-frequency masking coefficients M(f,t), which are applied to the input channels. Obtain the target time-frequency signal.
[0064] The input to the Encoder is a multi-channel time-frequency signal (i.e., the original time-frequency signal). The multi-channel time-frequency signal is encoded into feature vectors (i.e., semantic vectors) suitable for RNN processing. This usually involves some feature extraction steps, such as extracting the amplitude and phase information of the spectrum, or applying some form of feature transformation (such as Mel frequency cepstral coefficients MFCC). The feature vector sequence after encoding by the Encoder, each feature vector corresponds to a time frame.
[0065] The input to an RNN is a sequence of feature vectors output by the encoder and prior information (such as voiceprint feature vectors and / or orientation feature vectors). The RNN can utilize its recurrent structure to capture temporal dependencies and patterns in the feature vector sequence. The RNN can analyze the entire sequence and produce an output at each time step. These outputs can be seen as predictions or representations for the next time step. The output vector at each time step of the RNN can contain contextual information about the entire sequence.
[0066] The decoder takes the RNN's output vector at each time step (i.e., the intermediate data mentioned above) as input and decodes the RNN's output into multi-channel time-frequency masking coefficients. These time-frequency masking coefficients are typically used to adjust the amplitude or phase of the original time-frequency signal to achieve some form of audio enhancement or suppression. Each of the multi-channel time-frequency masking coefficients can correspond to a frequency component and a time frame in the original time-frequency signal. In this embodiment, the time-frequency masking coefficients can be used to strengthen time-frequency signals associated with prior information and suppress background noise and other interference signals. Optionally, the time-frequency masking coefficients can be assigned different weights to each time-frequency point to enhance or suppress the corresponding frequency component. Optionally, when adjusting the original time-frequency signal, the time-frequency masking coefficients can be multiplied element-wise with the original time-frequency signal to adjust each frequency component in the original time-frequency signal according to the time-frequency masking coefficients, thereby emphasizing time-frequency signals associated with prior information and suppressing time-frequency signals not associated with prior information.
[0067] In some implementations, when the prior information includes voiceprint information and azimuth information, inputting the original time-frequency signal and the prior information into the spatial sound pickup model to obtain the target time-frequency signal output by the spatial sound pickup model may include: extracting voiceprint features from the voiceprint information to obtain a voiceprint feature vector; extracting azimuth features from the azimuth information to obtain a azimuth feature vector; concatenating the voiceprint feature vector and the azimuth feature vector to obtain a priori feature vector; and inputting the original time-frequency signal and the priori feature vector into the spatial sound pickup model to obtain the target time-frequency signal output by the spatial sound pickup model.
[0068] Among them, the spatial sound pickup model can be based on personalized audio s(t) or location prior information d∈R. D The output is controlled by this mechanism. Personalized audio can be obtained by pre-collecting user audio or by selecting segments with high signal-to-noise ratios during real-time audio pickup. The personalized audio can then be processed by a pre-trained voiceprint extraction network, such as the ECAPA-TDNN model, to obtain the voiceprint feature vector s∈R. N The azimuth prior information is represented by a D-dimensional one-hot vector, which is transformed by stacked fully connected layers and the PReLU operator to obtain the azimuth feature vector. The voiceprint feature vector and the azimuth feature vector are concatenated and fed into the RNN module along with the output of the Encoder.
[0069] Specifically, this embodiment can also set two special configurations to deal with scenarios where the voiceprint feature vector has insufficient discriminative power or the directional information is biased. That is, when the voiceprint feature vector is a vector of all zeros, speaker differentiation is not performed, and when the prior information is a vector of all ones, sound pickup direction differentiation is not performed. In this case, the spatial sound pickup model is a general noise reduction model, that is, it only suppresses environmental noise and preserves all speech in the environment.
[0070] Step S250: Perform an inverse Fourier transform on the target time-frequency signal to obtain the target audio information.
[0071] In this embodiment, given the target time-frequency signal, an inverse Fourier transform can be performed on the target time-frequency signal to obtain the target audio information. The short-time Fourier transform and the inverse Fourier transform use the same window parameters and overlap to reconstruct the original time-domain signal.
[0072] The sound pickup method provided in one embodiment of this application is compared to... Figure 1 The sound pickup method shown in this embodiment further involves performing a short-time Fourier transform on the original audio information to obtain the original time-frequency signal. The original time-frequency signal and prior information are then input into a spatial sound pickup model to obtain the target time-frequency signal output by the spatial sound pickup model. This target time-frequency signal is obtained by enhancing the time-frequency signals in the original time-frequency signal that are related to the prior information and suppressing the time-frequency signals in the original time-frequency signal that are not related to the prior information through the spatial sound pickup model. The target audio information is obtained by performing an inverse Fourier transform on the target time-frequency signal. Thus, the target audio information can be extracted from the original audio information through the spatial sound pickup model, which can improve the convenience and accuracy of the extracted target audio information.
[0073] Please see Figure 4 , Figure 4 A schematic flowchart of a sound pickup method according to an embodiment of this application is shown. This method is applied to the aforementioned electronic device. In this embodiment, the prior information includes voiceprint information. The following will focus on... Figure 4 The process shown will be explained in detail. The sound pickup method may specifically include the following steps:
[0074] Step S310: Acquire raw audio information captured by the microphone array of the electronic device.
[0075] For a detailed description of step S310, please refer to step S110, which will not be repeated here.
[0076] Step S320: Determine multiple voiceprint information to be selected corresponding to the original audio information.
[0077] Optionally, prior information includes voiceprint information.
[0078] In this embodiment, when the original audio information is obtained, multiple voiceprint information corresponding to the original audio information can be determined as multiple voiceprint information to be selected.
[0079] In some implementations, when the original audio information is obtained, it can be segmented into shorter speech segments (frames) to facilitate further analysis. Subsequently, features representing the speaker's voiceprint are extracted from each speech segment as candidate voiceprint information. These features may include fundamental frequency, formants, intensity, duration, etc. Based on this, multiple candidate voiceprint information corresponding to the original audio information can be determined.
[0080] As an feasible approach, after extracting features representing the speaker's voiceprint from each speech segment, the extracted voiceprint features can be clustered to identify different voiceprint information. Finally, based on the clustering results, multiple candidate voiceprint information corresponding to the original audio information can be determined.
[0081] Step S330: Determine the voiceprint information with the highest frequency of occurrence from the plurality of voiceprint information to be selected, and determine the voiceprint information with the highest frequency of occurrence as the voiceprint information for the original audio information.
[0082] In this embodiment, when multiple candidate voiceprint information pieces are determined corresponding to the original audio information, the frequency of occurrence of each candidate voiceprint information piece in the original audio information can be determined. Based on the frequency of occurrence of each candidate voiceprint information piece in the original audio information, the candidate voiceprint information with the highest frequency is determined from the multiple candidate voiceprint information pieces, and this candidate voiceprint information with the highest frequency is determined as the voiceprint information for the original audio information. It can be understood that the voiceprint information with the highest frequency in the original audio information represents the speaker whose audio has the highest proportion, and is likely the target of the audio collection; therefore, its corresponding voiceprint information can be determined as the voiceprint information for the original audio information.
[0083] In some implementations, the current sound pickup scenario of the electronic device can be determined. If the current sound pickup scenario indicates that the electronic device needs to collect a person's voice, the voiceprint information with the highest frequency can be determined from multiple candidate voiceprint information, and the voiceprint information with the highest frequency is determined as the voiceprint information for the original audio information. Optionally, if the current sound pickup scenario includes a speech recording scenario, a lecture recording scenario, a concert recording scenario, etc., it can be determined that the current sound pickup scenario indicates that the electronic device needs to collect a person's voice.
[0084] In some implementations, if the current audio pickup scenario indicates that the electronic device needs to capture the voices of more than one person, the number of speakers input by the user can be obtained. Then, based on the frequency of occurrence of multiple candidate voiceprint information from high to low, candidate voiceprint information matching the number of speakers can be determined from the multiple candidate voiceprint information. This candidate voiceprint information matching the number of speakers is then identified as the voiceprint information for the original audio information. For example, if there are two speakers, the two candidate voiceprint information with the highest frequency can be determined from the multiple candidate voiceprint information. Optionally, if the current audio pickup scenario includes a conference recording scenario, it can be determined that the current audio pickup scenario indicates that the electronic device needs to capture the voices of more than one person.
[0085] Step S340: Enhance the audio information in the original audio information that is associated with the prior information, and suppress the audio information in the audio information that is not associated with the prior information to obtain the target audio information.
[0086] For a detailed description of step S340, please refer to step S130, which will not be repeated here.
[0087] The sound pickup method provided in one embodiment of this application is compared to... Figure 1 The sound pickup method shown in this embodiment further determines multiple candidate voiceprint information corresponding to the original audio information, identifies the candidate voiceprint information with the highest frequency from the multiple candidate voiceprint information, and identifies the candidate voiceprint information with the highest frequency as the voiceprint information for the original audio information. Thus, the corresponding voiceprint information can be determined from the original audio information according to the frequency of voiceprint occurrence, improving the convenience and accuracy of determining voiceprint information.
[0088] Please see Figure 5 , Figure 5 A schematic flowchart of a sound pickup method according to an embodiment of this application is shown. This method is applied to the aforementioned electronic device. In this embodiment, the prior information includes directional information, and the electronic device includes a screen. The following will focus on... Figure 5 The process shown will be explained in detail. The sound pickup method may specifically include the following steps:
[0089] Step S410: Acquire raw audio information captured by the microphone array of the electronic device.
[0090] For a detailed description of step S410, please refer to step S110, which will not be repeated here.
[0091] Step S420: Determine the orientation of the screen of the electronic device.
[0092] Optionally, prior information may include orientation information.
[0093] In this embodiment, the electronic device may include a screen. The orientation of the screen of the electronic device can be determined. Optionally, the orientation of the screen of the electronic device may include directly in front, diagonally behind, or to the left front, etc., and is not limited here.
[0094] Step S430: Determine the orientation of the screen of the electronic device as the directional information for the original audio information.
[0095] In this embodiment, when the orientation of the electronic device's screen is determined, the orientation of the electronic device's screen can be defined as the directional information relative to the original audio information. Optionally, if the orientation of the electronic device's screen is determined to be "directly forward," then "directly forward" can be defined as the directional information relative to the original audio information.
[0096] In some implementations, when the orientation of the electronic device's screen is determined as the directional information for the original audio information, it is possible to detect whether the orientation of the electronic device's screen has changed. Specifically, if a change in the orientation of the electronic device's screen is detected, the angle of change can be detected. If the angle of change does not reach an angle threshold, the previously determined directional information can be maintained. If the angle of change reaches an angle threshold, the duration for which the electronic device maintains the changed angle can be determined. If the duration does not reach a duration threshold, the previously determined directional information can be maintained. If the duration reaches a duration threshold, the changed orientation can be determined as the directional information for the original audio information.
[0097] Step S440: Enhance the audio information in the original audio information that is associated with the prior information, and suppress the audio information in the audio information that is not associated with the prior information to obtain the target audio information.
[0098] For a detailed description of step S440, please refer to step S130, which will not be repeated here.
[0099] The sound pickup method provided in one embodiment of this application is compared to... Figure 1 The sound pickup method shown in this embodiment also determines the orientation of the screen of the electronic device. The orientation of the screen of the electronic device is determined as the directional information for the original audio information, thereby realizing the automatic determination of directional information, and the determined directional information meets the usage requirements.
[0100] Please see Figure 6 , Figure 6 A schematic flowchart of a sound pickup method according to an embodiment of this application is shown. This method is applied to electronic devices, and will be discussed below. Figure 6 The process shown will be explained in detail. The sound pickup method may specifically include the following steps:
[0101] Step S510: Acquire raw audio information captured by the microphone array of the electronic device.
[0102] For a detailed description of step S510, please refer to step S110, which will not be repeated here.
[0103] Step S520: Display the interactive interface.
[0104] Optionally, the prior information may include at least one of voiceprint information and location information.
[0105] In this embodiment, the electronic device can display an interactive interface, which can be used to select and determine prior information.
[0106] In some implementations, the electronic device may display an interactive interface in response to a selection instruction based on prior information. Optionally, the electronic device may determine that a selection instruction based on prior information has been received upon receiving target voice information; it may determine that a selection instruction based on prior information has been received upon receiving a target touch operation (such as a target touch operation on a target icon, a target touch operation on a target button, etc.); it may determine that a selection instruction based on prior information has been received upon receiving a target shaking operation (such as shaking angle meeting a preset angle, shaking number meeting a preset number, etc.), etc., without limitation.
[0107] Step S530: In response to the selection operation applied to the interactive interface, determine prior information for the original audio information.
[0108] In this embodiment, the electronic device can detect operations performed on the interactive interface during the display process. Specifically, if a selection operation is detected on the interactive interface, the device can respond to the selection operation and determine prior information regarding the original audio information.
[0109] As an example, the interactive interface may include multiple prior information to be selected (such as multiple voiceprint information to be selected, multiple location information to be selected). Then, the user can select and determine the prior information for the original audio information from the multiple prior information to be selected based on the selection operation performed on the interactive interface.
[0110] As another example, the interactive interface may include an input box, through which the user can input prior information about the original audio information.
[0111] Step S540: Enhance the audio information in the original audio information that is associated with the prior information, and suppress the audio information in the audio information that is not associated with the prior information to obtain the target audio information.
[0112] For a detailed description of step S540, please refer to step S130, which will not be repeated here.
[0113] The sound pickup method provided in one embodiment of this application is compared to... Figure 1 The sound pickup method shown in this embodiment also displays an interactive interface. In response to the selection operation on the interactive interface, prior information for the original audio information is determined. Thus, prior information can be determined based on the user's interactive operation, thereby improving the user's interactive experience.
[0114] Please see Figure 7 , Figure 7 A schematic flowchart of a sound pickup method according to an embodiment of this application is shown. This method is applied to electronic devices, and will be discussed below. Figure 7 The process shown will be explained in detail. The sound pickup method may specifically include the following steps:
[0115] Step S610: Acquire raw audio information captured by the microphone array of the electronic device.
[0116] Step S620: Determine prior information for the original audio information, wherein the prior information includes at least one of voiceprint information and location information.
[0117] For a detailed description of steps S610-S620, please refer to steps S110-S120, which will not be repeated here.
[0118] Step S630: Enhance the audio information within the range corresponding to the directional information in the original audio information, and suppress the audio information outside the range corresponding to the directional information in the original audio information to obtain the audio information to be verified.
[0119] Optionally, the prior information includes at least location information, and may also include voiceprint information in addition to location information.
[0120] In this embodiment, when the original audio information is obtained, the audio information within the range corresponding to the directional information in the original audio information can be enhanced, and the audio information outside the range corresponding to the directional information in the original audio information can be suppressed, so as to obtain the audio information to be verified.
[0121] In some implementations, given the original audio information and location information, audio information within the range corresponding to the location information and audio information outside the range corresponding to the location information can be extracted from the original audio information. Then, the audio information within the range corresponding to the location information in the original audio information can be enhanced, and the audio information outside the range corresponding to the location information can be suppressed to obtain the audio information to be verified.
[0122] Step S640: If the audio information to be verified includes audio information corresponding to different voiceprints, then the audio information to be verified is filtered based on the voiceprint information to obtain the target audio information.
[0123] In this embodiment, upon obtaining the audio information to be verified, it is possible to detect whether the audio information to be verified includes audio information corresponding to different voiceprints. If it is detected that the audio information to be verified includes audio information corresponding to different voiceprints, it can be considered that the audio information extracted through location information contains interference. In this case, the audio information to be verified can be filtered based on voiceprint information to obtain the target audio information, thereby improving the accuracy of the obtained target audio information. If it is detected that the audio information to be verified does not contain audio information corresponding to different voiceprints, it can be considered that the audio information extracted through location information does not contain interference, and the audio information to be verified can be determined as the target audio information.
[0124] The sound pickup method provided in one embodiment of this application is compared to... Figure 1 The sound pickup method shown in this embodiment first enhances the audio information within the range corresponding to the directional information in the original audio information, and suppresses the audio information outside the range corresponding to the directional information in the original audio information to obtain the audio information to be verified. Then, if the audio information to be verified includes audio information corresponding to different voiceprints, the audio information to be verified is filtered based on the voiceprint information to obtain the target audio information. Thus, the accuracy of audio information pickup can be improved by performing secondary filtering through directional information and voiceprint information.
[0125] Please see Figure 8 , Figure 8 A schematic diagram of a sound pickup method provided in an embodiment of this application is shown. For example... Figure 8 As shown, the speaker A's directional information is input into the spatial sound pickup model, and the pickup area of the spatial sound pickup model is as follows. Figure 8 As shown, the pickup range is approximately 60°. Sounds outside the pickup area, such as speaker B and ambient noise, are suppressed by more than 30 dB, while speaker A's speech loss or gain is less than 3 dB. The spatial pickup model can adjust different pickup areas by changing prior information, for example, adjusting the pickup direction to the direction of speaker B.
[0126] Please see Figure 9 , Figure 9 This illustration shows another sound pickup diagram of the sound pickup method provided in the embodiments of this application. For example... Figure 9 As shown, when the interfering speaker C and speaker A are close to each other, the two cannot be effectively distinguished by the location information alone. At this time, the voiceprint information of speaker A is input into the spatial sound pickup model. The spatial sound pickup model can extract only the voice of A, while speaker C and environmental noise are suppressed, with a suppression amount greater than 30dB.
[0127] Please see Figure 10 , Figure 10 A block diagram of a microphone pickup device according to an embodiment of this application is shown. This microphone pickup device 200 is applied to the aforementioned electronic device, and will be discussed below. Figure 10 The block diagram shown illustrates that the sound pickup device 200 includes: a raw audio information acquisition module 210, a priori information determination module 220, and a target audio information acquisition module 230, wherein:
[0128] The raw audio information acquisition module 210 is used to acquire raw audio information collected by the microphone array of the electronic device.
[0129] The prior information determination module 220 is used to determine prior information for the original audio information, wherein the prior information includes at least one of voiceprint information and location information.
[0130] Furthermore, when the prior information includes voiceprint information, the prior information determination module 220 includes: a candidate voiceprint information determination submodule and a first prior information determination submodule, wherein:
[0131] The candidate voiceprint information determination submodule is used to determine multiple candidate voiceprint information corresponding to the original audio information.
[0132] The first prior information determination submodule is used to determine the candidate voiceprint information with the highest frequency from the plurality of candidate voiceprint information, and to determine the candidate voiceprint information with the highest frequency as the voiceprint information for the original audio information.
[0133] Furthermore, when the prior information includes orientation information, the prior information determination module 220 includes: an orientation determination submodule and a second prior information determination submodule, wherein:
[0134] An orientation determination submodule is used to determine the orientation of the screen of the electronic device.
[0135] The second prior information determination submodule is used to determine the orientation of the screen of the electronic device as directional information for the original audio information.
[0136] Furthermore, the prior information determination module 220 includes: an interactive interface display submodule and a third prior information determination submodule, wherein:
[0137] The interactive interface display submodule is used to display the interactive interface.
[0138] The third prior information determination submodule is used to determine prior information for the original audio information in response to the selection operation applied to the interactive interface.
[0139] The target audio information acquisition module 230 is used to enhance the audio information in the original audio information that is associated with the prior information, and to suppress the audio information in the audio information that is not associated with the prior information, so as to obtain the target audio information.
[0140] Further, the target audio information acquisition module 230 includes: an original time-frequency signal acquisition submodule, a target time-frequency signal acquisition submodule, and a first target audio information acquisition submodule, wherein:
[0141] The original time-frequency signal acquisition submodule is used to perform a short-time Fourier transform on the original audio information to obtain the original time-frequency signal.
[0142] The target time-frequency signal acquisition submodule is used to input the original time-frequency signal and the prior information into the spatial sound pickup model to obtain the target time-frequency signal output by the spatial sound pickup model. The target time-frequency signal is obtained by the spatial sound pickup model through enhancement processing of time-frequency signals in the original time-frequency signal that are related to the prior information, and suppression processing of time-frequency signals in the original time-frequency signal that are not related to the prior information.
[0143] Furthermore, the spatial sound pickup model includes an encoder, a recurrent neural network, and a decoder, and the target time-frequency signal acquisition submodule includes: a semantic vector acquisition unit, an intermediate data acquisition unit, and a first target time-frequency signal acquisition unit, wherein:
[0144] The semantic vector acquisition unit is used to input the original time-frequency signal into the encoder for processing, and obtain the semantic vector corresponding to the original time-frequency signal output by the encoder.
[0145] An intermediate data acquisition unit is used to input the prior information and the semantic vector into the recurrent neural network for processing, and obtain the intermediate data output by the recurrent neural network.
[0146] The first target time-frequency signal acquisition unit is used to input the intermediate data into the decoder for processing, obtain the time-frequency masking coefficient output by the decoder, enhance the time-frequency signals in the original time-frequency signal that are related to the prior information based on the time-frequency masking coefficient, and suppress the time-frequency signals in the original time-frequency signal that are not related to the prior information to obtain the target time-frequency signal.
[0147] Furthermore, when the prior information includes voiceprint information and azimuth information, the target time-frequency signal acquisition submodule includes: a voiceprint feature vector acquisition unit, an azimuth feature vector acquisition unit, a prior feature vector acquisition unit, and a second target time-frequency signal acquisition unit, wherein:
[0148] The voiceprint feature vector acquisition unit is used to extract voiceprint features from the voiceprint information to obtain a voiceprint feature vector.
[0149] The azimuth feature vector acquisition unit is used to extract azimuth features from the azimuth information to obtain azimuth feature vectors.
[0150] The prior feature vector acquisition unit is used to concatenate the voiceprint feature vector and the orientation feature vector to obtain the prior feature vector.
[0151] The second target time-frequency signal acquisition unit is used to input the original time-frequency signal and the prior feature vector into the spatial sound pickup model to obtain the target time-frequency signal output by the spatial sound pickup model.
[0152] The first target audio information acquisition submodule is used to perform inverse Fourier transform on the target time-frequency signal to obtain the target audio information.
[0153] Further, the target audio information acquisition module 230 includes: a submodule for acquiring audio information to be verified and a second target audio information acquisition submodule, wherein:
[0154] The submodule for obtaining audio information to be verified is used to enhance the audio information within the range corresponding to the directional information in the original audio information, and to suppress the audio information outside the range corresponding to the directional information in the original audio information, so as to obtain the audio information to be verified.
[0155] The second target audio information acquisition submodule is used to filter the audio information to be verified based on the voiceprint information if the audio information to be verified includes audio information corresponding to different voiceprints, and obtain the target audio information.
[0156] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0157] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0158] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0159] Please see Figure 11 This document illustrates a structural block diagram of an electronic device 100 provided in an embodiment of this application. The electronic device 100 can be a smartphone, tablet computer, e-reader, or other electronic device capable of running applications. The electronic device 100 in this application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications can be stored in the memory 120 and configured to be executed by one or more processors 110, and the one or more applications are configured to perform the methods described in the foregoing method embodiments.
[0160] The processor 110 may include one or more processing cores. The processor 110 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120. Optionally, the processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 110 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 110 and may be implemented separately using a communication chip.
[0161] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing functions (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the electronic device 100 during use (such as phonebook data, audio and video data, chat log data, etc.).
[0162] Please see Figure 12 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 300 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0163] The computer-readable storage medium 300 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 300 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 300 has storage space for program code 310 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 310 may be compressed, for example, in a suitable form.
[0164] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described sound pickup method.
[0165] In summary, the sound pickup method, apparatus, electronic device, and program product provided in this application acquire raw audio information collected by a microphone array of an electronic device, determine prior information for the raw audio information, wherein the prior information includes at least one of voiceprint information and location information, enhance audio information in the raw audio information that is associated with the prior information, and suppress audio information in the audio information that is not associated with the prior information to obtain target audio information. Thus, by combining prior information including voiceprint information and / or location information to enhance a part of the acquired audio information and suppress another part, the accuracy of the finally extracted audio information can be improved.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method of picking up sound, characterized by, The method is applied to an electronic device, and the method comprises: obtaining original audio information collected by a microphone array of the electronic device; determining prior information for the original audio information, wherein the prior information comprises at least one of voiceprint information and orientation information; performing enhancement processing on audio information in the original audio information that is associated with the prior information and performing suppression processing on audio information in the original audio information that is not associated with the prior information to obtain target audio information.
2. The method of claim 1, wherein, The method comprises: performing short-time Fourier transform on the original audio information to obtain original time-frequency signals; inputting the original time-frequency signals and the prior information into a spatial pickup model to obtain target time-frequency signals output by the spatial pickup model, wherein the target time-frequency signals are obtained by performing enhancement processing on time-frequency signals in the original time-frequency signals that are associated with the prior information and performing suppression processing on time-frequency signals in the original time-frequency signals that are not associated with the prior information by the spatial pickup model; performing inverse Fourier transform on the target time-frequency signals to obtain the target audio information.
3. The method of claim 2, wherein, The spatial pickup model comprises an encoder, a recurrent neural network, and a decoder, and the method comprises: inputting the original time-frequency signals into the encoder to obtain semantic vectors corresponding to the original time-frequency signals output by the encoder; inputting the prior information and the semantic vectors into the recurrent neural network to obtain intermediate data output by the recurrent neural network; inputting the intermediate data into the decoder to obtain time-frequency masking coefficients output by the decoder, and performing enhancement processing on time-frequency signals in the original time-frequency signals that are associated with the prior information and performing suppression processing on time-frequency signals in the original time-frequency signals that are not associated with the prior information based on the time-frequency masking coefficients to obtain the target time-frequency signals.
4. The method of claim 2, wherein, When the prior information comprises voiceprint information and orientation information, the method comprises: performing voiceprint feature extraction on the voiceprint information to obtain a voiceprint feature vector; performing orientation feature extraction on the orientation information to obtain an orientation feature vector; performing splicing processing on the voiceprint feature vector and the orientation feature vector to obtain a prior feature vector; inputting the original time-frequency signals and the prior feature vector into the spatial pickup model to obtain the target time-frequency signals output by the spatial pickup model.
5. The method of claim 1, wherein, When the prior information comprises voiceprint information, the method comprises: determining a plurality of to-be-selected voiceprint information corresponding to the original audio information; determine the voiceprint information with the highest occurrence frequency from the plurality of to-be-selected voiceprint information as the voiceprint information for the original audio information.
6. The method of claim 1, wherein, When the prior information includes orientation information and the electronic device includes a screen, the determining the prior information for the original audio information includes: determining an orientation of a screen of the electronic device; determining the orientation of the screen of the electronic device as the orientation information for the original audio information.
7. The method of claim 1, wherein, The determining the prior information for the original audio information includes: displaying an interactive interface; determining the prior information for the original audio information in response to a selection operation acting on the interactive interface.
8. The method of claim 1, wherein, The enhancing processing on the audio information in the original audio information that is associated with the prior information and the suppressing processing on the audio information in the original audio information that is not associated with the prior information to obtain the target audio information includes: enhancing processing on the audio information in the original audio information within a range corresponding to the orientation information and suppressing processing on the audio information in the original audio information outside the range corresponding to the orientation information to obtain to-be-verified audio information; if the to-be-verified audio information includes audio information corresponding to different voiceprints, performing screening on the to-be-verified audio information based on voiceprint information to obtain the target audio information.
9. A pickup device, characterized in that An apparatus applied to an electronic device includes: an original audio information acquisition module configured to acquire original audio information collected by a microphone array of the electronic device; a prior information determination module configured to determine prior information for the original audio information, wherein the prior information includes at least one of voiceprint information and orientation information; a target audio information obtaining module configured to perform enhancing processing on audio information in the original audio information that is associated with the prior information and suppressing processing on audio information in the original audio information that is not associated with the prior information to obtain target audio information.
10. An electronic device, comprising: A computer program product includes a computer program, and the computer program is executed by a processor to implement the method in any one of claims 1-8.
11. A computer program product, characterised in that, The computer program product includes a computer program, and the computer program is executed by a processor to implement the method in any one of claims 1-8.