Processing method, device and electronic equipment
By setting up multiple pickups on handheld calling devices to collect noise data, and using AI or sound algorithms to determine and eliminate noise, the problem of poor voice quality in noisy environments is solved, achieving better call clarity and user experience.
Patent Information
- Application Number
- CN202210917288.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-01
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-08-01
AI Technical Summary
In noisy environments, it is difficult to maintain a friendly experience in the voice quality of handheld calling devices, and the prior art is difficult to effectively reduce noise interference.
Multiple pickups are used to collect noise data at different locations and orientations, and the target noise data is determined through AI noise reduction algorithm or acoustic noise reduction algorithm, and the voice data is reduced based on this data.
Improves voice signal-to-noise ratio, optimizes user's call experience, and enhances voice clarity in noisy environments.
Smart Images

Figure CN115294998B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to data processing technology, and more specifically, to a processing method, device and electronic device. Background Art
[0002] The usage scenarios of handheld communication devices are complex and changeable. How to better improve the voice quality so that it still has a friendly user experience in noisy environments has been a long-term research topic for technicians in this field. Summary of the Invention
[0003] In view of this, this application provides the following technical solutions:
[0004] A processing method, comprising:
[0005] Obtaining first voice data collected by a first microphone, wherein the first voice data includes noise data;
[0006] Obtaining second noise data collected by a second microphone and third noise data collected by a third microphone, wherein the first voice data, the second noise data, and the third noise data are data collected during the same period; the first microphone, the second microphone, and the third microphone are located at different positions of the electronic device, are not in a straight line, and have different sound pickup directions;
[0007] determining target noise data based on the second noise data and the third noise data, wherein the target noise data is used to at least partially offset the noise data of the first speech data;
[0008] Noise reduction processing is performed on the first speech data based on the target noise data to obtain second speech data.
[0009] Optionally, determining target noise data based on the second noise data and the third noise data includes:
[0010] determining a spatial location of a noise source based on the second noise data and the third noise data;
[0011] determining a phase of the noise source relative to the first microphone based on the spatial position of the noise source;
[0012] Target noise data in the first speech data is determined based on the phase.
[0013] Optionally, determining target noise data based on the second noise data and the third noise data includes:
[0014] The first speech data is compared with a noise data set consisting of the second noise data and the third noise data to determine target noise data in the first speech data.
[0015] Optionally, performing noise reduction processing on the first speech data based on the target noise data to obtain second speech data includes:
[0016] The target noise data in the first speech data is attenuated or eliminated.
[0017] Optionally, determining target noise data based on the second noise data and the third noise data includes:
[0018] Based on the second noise data and the third noise data, an AI noise reduction algorithm or a sound wave noise reduction algorithm is used to determine target noise data in the first voice data.
[0019] Optionally, obtaining the second noise data collected by the second microphone and the third noise data collected by the third microphone includes:
[0020] Obtaining second audio data collected by the second microphone and third audio data collected by the third microphone;
[0021] Determining a first object voiceprint feature based on a relative positional relationship among the first microphone, the second microphone, and the third microphone, in combination with the first voice data, the second audio data, and the third audio data;
[0022] The sound signal corresponding to the first object voiceprint feature in the second audio data is deleted to obtain second noise data, and the signal corresponding to the first object voiceprint feature in the third audio data is deleted to obtain third noise data.
[0023] Optionally, the electronic device further includes a fourth pickup, wherein the first pickup and the second pickup are arranged on opposite sides of the electronic device in a first direction, the third pickup and the fourth pickup are arranged on opposite sides of the electronic device in a second direction, and the first direction is perpendicular to the direction, and further includes:
[0024] Obtaining fourth noise data collected by the fourth microphone;
[0025] The determining target noise data based on the second noise data and the third noise data includes:
[0026] Target noise data is determined based on the second noise data, the third noise data, and the fourth noise data.
[0027] Optionally, before determining the target noise data based on the second noise data and the third noise data, the method further includes:
[0028] A first processing is performed on the first voice data, the second noise data, and the third noise data, wherein the first processing includes echo cancellation, analysis, and non-stationary audio signal processing.
[0029] The present application also discloses a processing device, comprising:
[0030] A voice acquisition module, configured to acquire first voice data collected by a first microphone, wherein the first voice data includes noise data;
[0031] a noise acquisition module, configured to obtain second noise data collected by a second microphone and third noise data collected by a third microphone, wherein the first voice data, the second noise data, and the third noise data are data collected during the same period; the first microphone, the second microphone, and the third microphone are located at different positions of the electronic device, are not in a straight line, and have different sound pickup directions;
[0032] a noise determination module, configured to determine target noise data based on the second noise data and the third noise data, wherein the target noise data is used to at least partially offset the noise data of the first speech data;
[0033] The noise reduction processing module is used to perform noise reduction processing on the first speech data based on the target noise data to obtain second speech data.
[0034] Furthermore, the present application also discloses an electronic device, comprising:
[0035] processor;
[0036] a memory for storing executable program instructions of the processor;
[0037] The executable program instructions include: obtaining first voice data collected by a first microphone, wherein the first voice data includes noise data; obtaining second noise data collected by a second microphone and third noise data collected by a third microphone, wherein the first voice data, the second noise data and the third noise data are data collected during the same period of time; the first microphone, the second microphone and the third microphone are located at different positions of the electronic device, the first microphone, the second microphone and the third microphone are not in a straight line and have different sound pickup directions; determining target noise data based on the second noise data and the third noise data, wherein the target noise data is used to at least partially offset the noise data of the first voice data; and performing noise reduction processing on the first voice data based on the target noise data to obtain second voice data.
[0038] The embodiments of the present application disclose a processing method, device and electronic device, the method comprising: obtaining first voice data collected by a first microphone, the first voice data including noise data; obtaining second noise data collected by a second microphone and third noise data collected by a third microphone; the first microphone, the second microphone and the third microphone are not in a straight line and have different sound pickup orientations; determining target noise data based on the second noise data and the third noise data, the target noise data being used to at least partially offset the noise data of the first voice data; and performing noise reduction processing on the first voice data based on the target noise data to obtain second voice data. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0040] Figure 1 A flowchart of a processing method disclosed in an embodiment of the present application;
[0041] Figure 2 This is a schematic diagram of a user using a mobile phone to make a call according to an embodiment of the present application;
[0042] Figure 3 A schematic plan view of the microphone installation position on a smartphone disclosed in an embodiment of the present application;
[0043] Figure 4 This is a schematic diagram of the noise pickup range disclosed in the embodiments of this application;
[0044] Figure 5 A flowchart of determining target noise data disclosed in an embodiment of the present application;
[0045] Figure 6 This is a flow chart of obtaining noise data collected by a microphone disclosed in an embodiment of the present application;
[0046] Figure 7 A flowchart of an implementation architecture of a processing method disclosed in an embodiment of the present application;
[0047] Figure 8 A schematic structural diagram of a processing device disclosed in an embodiment of the present application;
[0048] Figure 9 This is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0050] The embodiments of the present application can be applied to electronic devices. The present application does not limit the product form of the electronic device, and can include but is not limited to smartphones, tablet computers, wearable devices, personal computers (PCs), netbooks, etc., which can be selected according to application requirements. To facilitate a better understanding of the specific implementation of the present application, the following embodiments will be described using smartphones as an example.
[0051] Figure 1 This is a flowchart of a processing method disclosed in an embodiment of this application. Figure 1 As shown, the processing method may include:
[0052] Step 101: obtaining first voice data collected by a first microphone, wherein the first voice data includes noise data.
[0053] The first microphone can be a microphone on an electronic device that mainly collects the user's voice. It is usually set at the position closest to the user's mouth when the user uses the electronic device. For example, the microphone set at the bottom of a smartphone. When the user makes a call, the user will put the phone close to the face and place the bottom of the phone close to the mouth. At this time, the microphone at the bottom of the phone can collect the user's complete voice data, such as Figure 2 Position relationship shown.
[0054] When a user makes or receives a call on a smartphone, there may be some noise in the environment, such as car horns, animals barking, and music. These noises and the user's voice will be collected by the first microphone. If this noisy voice data is directly transmitted to the other terminal, the other party will not be able to hear the user's content clearly. Therefore, before the collected voice data is output to the other terminal, it is necessary to reduce the noise of the voice data.
[0055] Step 102: Obtain second noise data collected by the second microphone and third noise data collected by the third microphone.
[0056] The first voice data, the second noise data, and the third noise data are data collected during the same period. In addition, the first microphone, the second microphone, and the third microphone are located at different positions of the electronic device, are not in a straight line, and have different sound pickup directions.
[0057] Those skilled in the art will appreciate that a microphone has a certain effective pickup range, which can be specifically understood as a cone-shaped area with a certain angle and centered at the microphone. In this embodiment, the electronic device is provided with two microphones (a second microphone and a third microphone) installed at different locations and with different pickup orientations to collect reference noise. The reference noise is used to perform noise reduction processing on the first voice data collected by the first microphone.
[0058] Figure 3 This is a schematic plan view of the microphone installation position on the smartphone disclosed in the embodiment of this application. Figure 3 As shown, the first pickup ① is installed at the bottom of the mobile phone, the second pickup ② is installed at the top of the mobile phone, and the third pickup ③ is installed on the side of the mobile phone.
[0059] Figure 4 This is a schematic diagram of the noise pickup range disclosed in the embodiment of this application. The black structures represent the second pickup and the third pickup respectively, and the area within the dotted line extending from the pickup is the main pickup range of the pickup. Figure 4 As shown in the figure, the second pickup range of the second microphone on the top of the phone is mainly in the range of the top direction of the phone, and the third pickup range of the third microphone on the side of the phone is mainly in the range of the side of the phone; under the joint action of these two microphones, the overall effective pickup range is wider, making the final collected reference noise signal stronger and the angle information richer, which helps the noise reduction algorithm better identify the relevant information of the noise source. Of course, Figure 4The pickup range of the displayed pickup is the main pickup range of the pickup, and it is a plane diagram. In fact, it is a three-dimensional conical range area. If the sound source is outside the main pickup range, the pickup can also collect the sound of the sound source, but the signal of the sound source collected will be relatively weak.
[0060] Step 103: Determine target noise data based on the second noise data and the third noise data, wherein the target noise data is used to at least partially offset the noise data of the first speech data.
[0061] In this embodiment, the target noise data in the first voice data can be determined by using, but not limited to, an AI noise reduction algorithm or a sound wave noise reduction algorithm.
[0062] As previously mentioned, the use of two microphones with different sound pickup orientations provides a wider effective noise pickup range, making the collected noise data more valuable for reference. In this step, the target noise data determined based on the second noise data and the third noise data can be data related to the noise emitted by the noise source, or at least a portion of the noise data in the first speech data. The implementation of determining the target noise data will be described in detail in later embodiments and will not be further elaborated here.
[0063] Step 104: Perform noise reduction processing on the first speech data based on the target noise data to obtain second speech data.
[0064] After the target noise data is determined, noise reduction processing can be performed on the first voice data based on the target noise data to at least partially eliminate the noise data in the first voice data and obtain second voice data with a higher signal-to-noise ratio.
[0065] The processing method described in this embodiment collects noise data through two microphones with different pickup directions, picks up noise in a wide range, and can obtain reference noise with stronger signals and more comprehensive angle information, thereby providing support for more accurate voice noise reduction and optimizing the user experience.
[0066] Figure 5 This is a flow chart of determining target noise data disclosed in an embodiment of the present application. Figure 5 As shown, determining the target noise data based on the second noise data and the third noise data may include:
[0067] Step 501: Determine the spatial position of a noise source based on the second noise data and the third noise data.
[0068] Since the pickup ranges of the second and third microphones that pick up noise are fixed relative to the electronic device itself, the spatial position of the noise source can be determined through corresponding analysis of the noise data picked up by the two microphones.
[0069] Step 502: Determine the phase of the noise source relative to the first microphone based on the spatial position of the noise source.
[0070] After the spatial position of the noise source is determined, the orientation of the noise source relative to the first microphone, that is, the phase of the noise source relative to the first microphone, can be determined.
[0071] Step 503: Determine target noise data in the first speech data based on the phase.
[0072] Once the phase of the noise source relative to the first microphone is determined, the sound wave signal from this phase can be determined as the target noise data. The target noise data from this phase can then be processed in a targeted manner to achieve the purpose of noise reduction for the first voice data.
[0073] In this implementation, based on the noise collected by two microphones with different positions and pickup directions, the spatial position of the noise source can be analyzed and determined, and then the target noise data in the first voice data can be gradually determined according to the spatial position of the noise source, so as to facilitate subsequent targeted noise reduction processing of the first voice data.
[0074] In another implementation, determining the target noise data based on the second noise data and the third noise data may include: comparing the first speech data with a noise data set consisting of the second noise data and the third noise data to determine the target noise data in the first speech data.
[0075] The first voice data includes both the user's voice and noise data, while the second and third noise data only include noise data. It is understood that the noise data in the first voice data collected by the first microphone is identical to the second noise data collected by the second microphone and the third noise data collected by the third microphone. However, due to the different pickup ranges of the microphones, the noise data in the first voice data differs from the second and third noise data in signal strength.
[0076] In this implementation, part of the audio data in the first speech data that has the same waveform and / or voiceprint features as the second noise data and / or the third noise data can be determined as the target noise data.
[0077] After determining the target noise data, noise reduction processing is performed on the first speech data based on the target noise data to obtain the second speech data. Specifically, this can be achieved by attenuating or eliminating the target noise data in the first speech data. Whether attenuating or eliminating the target noise data, the signal-to-noise ratio of the second speech data obtained after processing is superior to that of the first speech data before processing, thereby achieving the purpose of speech noise reduction.
[0078] In the previous embodiments, the second and third microphones were described as collecting second and third noise data. However, in reality, the second and third microphones can collect all sounds in the environment, including the user's voice. This implementation involves identifying the user's voice through recognition after the second and third microphones collect audio data containing the user's voice. The user's voice portion is then deleted, leaving only the noise data as the reference noise input to the noise reduction algorithm.
[0079] The process of obtaining noise data collected by the microphone can be found in Figure 6 .like Figure 6 As shown, obtaining the second noise data collected by the second microphone and the third noise data collected by the third microphone may include:
[0080] Step 601: Obtain second audio data collected by the second microphone and third audio data collected by the third microphone.
[0081] The second audio data and the third audio data here are both audio data including user voice, that is, original audio data collected by the second microphone and the third microphone without any processing.
[0082] Step 602: Determine a first object voiceprint feature based on the relative position relationship among the first microphone, the second microphone, and the third microphone, in combination with the first voice data, the second audio data, and the third audio data.
[0083] It is clear that when the user is making or receiving a call, all three microphones will collect the user's voice data; since the user's mouth is closest to the first microphone, after determining which voiceprint or voiceprints correspond to the person's voice signal based on the characteristics of human speech, the voice signal of the person with the strongest signal in the audio data collected by the first microphone can be determined as the user's voice, and the voiceprint characteristics of the user (first object) can be determined.
[0084] Step 603: Delete the sound signal corresponding to the first object voiceprint feature in the second audio data to obtain second noise data, and delete the signal corresponding to the first object voiceprint feature in the third audio data to obtain third noise data.
[0085] After determining the user's voiceprint characteristics, the corresponding sound signals in the second and third audio data are deleted, thereby deleting the voice data used in the second and third audio data, and obtaining pure noise data. The pure noise data here can also include the voices of people other than the user.
[0086] In other implementations, the electronic device may further include a fourth pickup. The first pickup and the second pickup are arranged on opposite sides of the electronic device in the first direction, Figure 3 As shown, that is, it is set at the bottom and top of the mobile phone; the third pickup and the fourth pickup are set at two opposite sides of the second direction of the electronic device, and the first direction is perpendicular to the direction; Figure 3 , the third microphone and the fourth microphone can be set on the left and right side of the mobile phone respectively.
[0087] Based on this, and in accordance with the disclosure of the aforementioned embodiment, the noise reduction method may further include: obtaining fourth noise data collected by the fourth microphone. Determining target noise data based on the second noise data and the third noise data may include: determining the target noise data based on the second noise data, the third noise data, and the fourth noise data.
[0088] In this implementation, the phone is equipped with four microphones: mic2, mic3, and mic4, which collect noise data. Compared to the previous solution that does not include a fourth microphone, this implementation has a wider noise pickup range and can obtain richer signal and angle information for the noise data, which helps to more accurately perform noise reduction on the first voice data.
[0089] Figure 7 This is a flowchart of the implementation architecture of a processing method disclosed in an embodiment of this application. Figure 7As shown in the figure, in a specific implementation, the mobile phone includes three microphones. In the handheld call mode, three-way voice signals are input through the three microphones, firstly, the linear echo cancellation filter is used for echo cancellation, and then the short-time Fourier transformer is used for spectrum analysis and non-steady-state audio signal processing; microphone 1 is the microphone closest to the user's mouth in the physical space in the handheld call mode, and will be locked to pick up only the main voice signal, that is, only microphone 1 will pass through VAD (Voice ActivityDetection, voice endpoint detection), microphone 2 and microphone 3 will be locked as noise-only microphones, that is, microphones 2 and 3 will pass through the noise reference module. Microphone 2 is designed at the top of the phone and can receive more noise from the top, back, and part of the side of the phone. Microphone 3 is designed at the upper position of the phone frame and can receive more noise from the side, front, and part of the top and back of the phone. Microphone 2 and microphone 3 form a physical space complement in phase. Microphone 2 + microphone 3 and microphone 1 form a facial array microphone. Compared with the traditional microphone 1 + microphone 2 linear array microphone, it forms a beamforming spatial filter with a larger noise reduction range. Compared with the design of realizing noise reduction with two microphones, this system will input a more accurate noise reference into the noise algorithm for noise comparison, and then send it to the noise post-processing module for further suppression and post-processing of the frequency band. Finally, it will output a voice signal that does not contain noise or has very little noise component.
[0090] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0091] The above embodiments disclosed in the present application describe the method in detail. The method of the present application can be implemented using various devices. Therefore, the present application also discloses a device, and a specific embodiment is given below for detailed description.
[0092] Figure 8 This is a schematic diagram of the structure of a processing device disclosed in an embodiment of this application. Figure 9 As shown, the processing device 80 may include:
[0093] The voice acquisition module 801 is configured to acquire first voice data collected by a first microphone, wherein the first voice data includes noise data.
[0094] The noise acquisition module 802 is used to obtain second noise data collected by the second microphone and third noise data collected by the third microphone, where the first voice data, the second noise data, and the third noise data are data collected during the same period of time; the first microphone, the second microphone, and the third microphone are located at different positions of the electronic device, are not in a straight line, and have different sound pickup directions.
[0095] The noise determination module 803 is configured to determine target noise data based on the second noise data and the third noise data, wherein the target noise data is used to at least partially offset the noise data of the first speech data.
[0096] The noise reduction processing module 804 is configured to perform noise reduction processing on the first speech data based on the target noise data to obtain second speech data.
[0097] The processing device described in this embodiment collects noise data through two microphones with different pickup directions, picks up noise in a wide range, and can obtain reference noise with stronger signals and more comprehensive angle information, thereby providing support for more accurate voice noise reduction and optimizing the user experience.
[0098] In one implementation, the noise determination module can be specifically used to: determine the spatial position of the noise source based on the second noise data and the third noise data; determine the phase of the noise source relative to the first microphone based on the spatial position of the noise source; and determine the target noise data in the first speech data based on the phase.
[0099] In one implementation, the noise determination module may be specifically configured to compare the first speech data with a noise data set consisting of the second noise data and the third noise data to determine target noise data in the first speech data.
[0100] In one implementation, the noise reduction processing module may be specifically configured to attenuate or eliminate the target noise data in the first speech data.
[0101] In one implementation, the noise determination module may be specifically configured to: determine target noise data in the first speech data using an AI noise reduction algorithm or a sound wave noise reduction algorithm based on the second noise data and the third noise data.
[0102] In one implementation, the noise acquisition module can be specifically used to: obtain second audio data collected by the second microphone and third audio data collected by the third microphone; determine the first object voiceprint feature based on the relative position relationship of the first microphone, the second microphone and the third microphone, in combination with the first voice data, the second audio data and the third audio data; delete the sound signal corresponding to the first object voiceprint feature in the second audio data to obtain second noise data, and delete the signal corresponding to the first object voiceprint feature in the third audio data to obtain third noise data.
[0103] In one implementation, the electronic device further includes a fourth pickup, wherein the first pickup and the second pickup are arranged on opposite sides of the electronic device in a first direction, the third pickup and the fourth pickup are arranged on opposite sides of the electronic device in a second direction, and the first direction is perpendicular to the direction, then the noise acquisition module is further used to: obtain fourth noise data collected by the fourth pickup; the noise determination module is further used to: determine target noise data based on the second noise data, the third noise data and the fourth noise data.
[0104] In one implementation, the pre-processing module is used to perform a first processing on the first speech data, the second noise data and the third noise data before the noise determination module determines the target noise data based on the second noise data and the third noise data, wherein the first processing includes echo cancellation, analysis and non-stationary audio signal processing.
[0105] Any of the processing devices described in the above embodiments includes a processor and a memory. The speech acquisition module, noise acquisition module, noise determination module, noise reduction processing module, etc. in the above embodiments are all stored in the memory as program modules, and the processor executes the above program modules stored in the memory to realize corresponding functions.
[0106] The processor contains a kernel, which retrieves the corresponding program module from the memory. There can be one or more kernels, and the kernel parameters can be adjusted to process the returned data.
[0107] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0108] In an exemplary embodiment, a computer-readable storage medium is also provided, which can be directly loaded into the internal memory of a computer and contains software code. After being loaded and executed by a computer, the computer program can implement the steps shown in any embodiment of the above processing method.
[0109] In an exemplary embodiment, a computer program product is also provided, which can be directly loaded into the internal memory of a computer and contains software code. After being loaded and executed by the computer, the computer program can implement the steps shown in any embodiment of the processing method described above.
[0110] Furthermore, an embodiment of the present application provides an electronic device. Figure 9 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application. Figure 9 As shown, the electronic device 90 includes at least one processor 901, and at least one memory 902 and a bus 903 connected to the processor; wherein the processor and the memory communicate with each other through the bus; the processor is used to call the executable program instructions in the memory to execute the above-mentioned processing method.
[0111] The executable program instructions include: obtaining first voice data collected by a first microphone, wherein the first voice data includes noise data; obtaining second noise data collected by a second microphone and third noise data collected by a third microphone, wherein the first voice data, the second noise data and the third noise data are data collected during the same period of time; the first microphone, the second microphone and the third microphone are located at different positions of the electronic device, the first microphone, the second microphone and the third microphone are not in a straight line and have different sound pickup directions; determining target noise data based on the second noise data and the third noise data, wherein the target noise data is used to at least partially offset the noise data of the first voice data; and performing noise reduction processing on the first voice data based on the target noise data to obtain second voice data.
[0112] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0113] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0114] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0115] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A treatment method, comprising: Obtaining first voice data collected by a first microphone, wherein the first voice data includes noise data; Obtaining second noise data collected by a second microphone and third noise data collected by a third microphone, wherein the first voice data, the second noise data, and the third noise data are data collected during the same period; the first microphone, the second microphone, and the third microphone are located at different positions of the electronic device, are not in a straight line, and have different sound pickup directions; determining a spatial location of a noise source based on the second noise data and the third noise data; determining a phase of the noise source relative to the first microphone based on the spatial position of the noise source; determining target noise data in the first speech data based on the phase; Noise reduction processing is performed on the first speech data based on the target noise data to obtain second speech data.
2. The processing method according to claim 1, wherein determining the target noise data in the first speech data based on the phase comprises: The acoustic wave signal from the phase is determined as target noise data.
3. The processing method according to claim 1, wherein the performing noise reduction processing on the first speech data based on the target noise data to obtain the second speech data comprises: The target noise data in the first speech data is attenuated or eliminated.
4. The processing method according to claim 1, further comprising: A first processing is performed on the first voice data, the second noise data, and the third noise data, wherein the first processing includes echo cancellation, analysis, and non-stationary audio signal processing.
5. A processing device, comprising: A voice acquisition module, configured to acquire first voice data collected by a first microphone, wherein the first voice data includes noise data; a noise acquisition module, configured to obtain second noise data collected by a second microphone and third noise data collected by a third microphone, wherein the first voice data, the second noise data, and the third noise data are data collected during the same period; the first microphone, the second microphone, and the third microphone are located at different positions of the electronic device, are not in a straight line, and have different sound pickup directions; a noise determination module, configured to determine a spatial position of a noise source based on the second noise data and the third noise data; and determine a phase of the noise source relative to the first microphone based on the spatial position of the noise source; determining target noise data in the first speech data based on the phase; The noise reduction processing module is used to perform noise reduction processing on the first speech data based on the target noise data to obtain second speech data.
6. An electronic device comprising: processor; a memory for storing executable program instructions of the processor; The executable program instructions include: obtaining first voice data collected by a first microphone, wherein the first voice data includes noise data; obtaining second noise data collected by a second microphone and third noise data collected by a third microphone, wherein the first voice data, the second noise data, and the third noise data are data collected during the same period of time; the first microphone, the second microphone, and the third microphone are located at different positions of the electronic device, and the first microphone, the second microphone, and the third microphone are not in a straight line and have different sound pickup directions; determining a spatial position of a noise source based on the second noise data and the third noise data; determining a phase of the noise source relative to the first microphone based on the spatial position of the noise source; determining target noise data in the first voice data based on the phase; and performing noise reduction processing on the first voice data based on the target noise data to obtain second voice data.
Citation Information
Patent Citations
Audio capture with multiple microphones
CN108777732A
Voice recognition method and system
CN110931007A